Senior AI Engineer building production Agent and RAG systems for NetBrain’s no-code network automation platform. Improving LLM reliability, observability, and automation at enterprise scale.
Responsibilities
Design and implement core capabilities for an enterprise-grade Agent platform, including orchestration patterns, tool execution, context and memory management, and safety guardrails
Design Agent execution and governance mechanisms, including Human-in-the-Loop approval workflows, multi-tenant permission isolation, policy enforcement, and secure execution controls
Build reusable Agent Skills, standardized tool interfaces, and a scalable tool ecosystem integrated with NetBrain platform capabilities and business workflows
Design and implement LLM post-training strategies, preference alignment, and parameter-efficient fine-tuning techniques
Build self-learning feedback loops using production traces, user feedback, and evaluation results
Analyze and optimize LLM behavior, including instruction following, tool calling, structured outputs, contextual understanding, reasoning stability, and hallucination mitigation
Build LLM and Agent evaluation frameworks, automated regression pipelines, quality gates, and hallucination-detection mechanisms
Build AI system observability capabilities including distributed tracing, structured logging, metrics, dashboards, and alerting
Diagnose and resolve production AI failures and unexpected model behavior changes
Design reliable backend services for AI and Agent workloads with asynchronous processing, retries, timeouts, caching, rate limiting, and fault isolation
Optimize latency, throughput, token consumption, and infrastructure cost for large-scale production workloads
Lead technical design for critical modules and system-level capabilities
Collaborate with Engineering, Product, QA, and other teams to deliver production solutions
Prototype, benchmark, and productionize GraphRAG, Knowledge Graphs, MCP, LLM post-training, and Agent self-learning technologies
Evaluate Agent frameworks and infrastructure and provide recommendations for platform architecture and product technology strategy
Requirements
Bachelor's degree or higher in Computer Science, Artificial Intelligence, Electrical Engineering, or a related technical field; equivalent practical experience will also be considered
3+ years of experience in software engineering, machine learning, or applied AI
2+ years building, deploying, and operating production-grade LLM or Agent applications
Delivered at least one LLM-powered feature end-to-end and owned its ongoing operation and improvement after production launch
Deep understanding of Agent architectures and LLM behavioral characteristics
Hands-on experience building multi-step workflows involving reasoning, tool execution, state management, structured outputs, validation, and error recovery
Ability to diagnose and resolve production LLM/Agent failures
Strong Python and distributed backend engineering skills
Experience with API and service development, asynchronous and concurrent programming, retries, timeouts, caching, rate limiting, testing, logging, and cross-service performance debugging
Hands-on experience designing evaluation systems for LLM applications
Strong understanding of security risks associated with LLM and Agent applications
Ability to independently design, implement, debug, deploy, and operate complex production systems
Preferred: experience with RAG, advanced retrieval systems, Knowledge Graphs, GraphRAG, LangGraph, LangChain, AutoGen, LlamaIndex, MCP, LangSmith, LLM fine-tuning, and applying LLM technologies to complex technical domains
Fluent in both English and Chinese
Benefits
Bonus
RRSP
Medical/dental coverage
Comprehensive benefits package
Reasonable accommodation in the application process
AI Engineer building production AI, Generative AI, and predictive pricing solutions for RBC Capital Markets. Developing scalable desk - facing tools for Sales & Trading.
Full - stack AI engineer developing agentic applications for Magna’s automotive mobility technology business. Building cloud - deployed workflows, data integrations, and intuitive interfaces for R&D teams.
AI and Automation Leader building secure enterprise AI and automation for Top Aces’ defense training and aircraft sustainment services. Improving aviation readiness, MRO, supply chain, compliance, and corporate productivity.
AI developer building machine - learning models and image datasets for GenAIz's life science and pharmaceutical software. Collaborating with developers, designers, product, business, and architecture teams.
Lead AI Engineer building secure AWS AI, ML, and GenAI solutions for Sun Life’s financial - services business. Delivering Bedrock applications, RAG pipelines, intelligent agents, and enterprise automation.
AI Engineer building AWS - based GenAI applications, RAG solutions, and intelligent agents for Sun Life’s financial security and health services. Collaborating across technology and business teams in a hybrid environment.
Senior AI Engineer building secure, production - grade AI, ML, and GenAI solutions on AWS for Sun Life’s financial security and health business. Delivering RAG, intelligent agents, automation, and reusable enterprise platforms.
AI/ML Software Engineer shaping AI strategy and building scalable AI/ML components. Developing solutions atop Smile Digital Health’s FHIR - based healthcare data platform.
Senior Principal Engineer leading AI - native software architecture and delivery at Revvity, a health - technology solutions provider. Mentoring engineers and driving scalable production systems.