Staff AI Builder designing and shipping production GenAI systems for Robots and Pencils’ enterprise clients. Leading agentic AI, RAG, AWS architecture, reliability, and technical mentoring.
Responsibilities
Lead the design and delivery of AI/ML systems
Define AI architecture and lead model development and optimization
Design and build agentic workflows, including reasoning loops, tool/function calling, and single- and multi-agent orchestration
Build and maintain RAG pipelines with chunking, embeddings, OpenSearch vector search, re-ranking, and refresh handling
Integrate AWS Bedrock and Agent Core, including MCP-based tool design
Write and iterate on production system prompts
Build evaluation and observability into agents using golden datasets, RAGAS-style metrics, LLM-as-judge, and tracing
Design for LLM failures using retries, backoff, circuit breakers, fallback models, and user-facing degradation handling
Build backend services in Python and Node.js, including AWS Lambda and API Gateway serverless architectures
Design DynamoDB single-table schemas for conversation state, agent memory, and session history
Support event-driven orchestration with Step Functions, SQS, and EventBridge
Contribute to frontend integration points and write clean, tested code across the stack
Support deployment, monitoring, and production troubleshooting in AWS and Docker environments
Partner with product, design, and delivery leads across global teams to scope and ship features end-to-end
Mentor engineers and improve agentic engineering practices, including use of Claude Code
Own features and releases end-to-end, including debugging, hardening, and production reliability
Requirements
6+ years of professional software engineering experience, including meaningful time shipping GenAI/LLM-powered systems in production
Hands-on depth in agentic AI, including reasoning loops, tool/function calling, and multi-agent orchestration
Practical RAG expertise, including chunking strategies, embeddings, vector databases such as OpenSearch or similar, cosine similarity search, and re-ranking
Experience building evaluation and observability for LLM systems, including golden datasets, LLM-as-judge, RAGAS or comparable metrics, and LangFuse/LangSmith or similar tracing tools
Strong prompt engineering skills
Hands-on experience with the AWS GenAI stack: Bedrock, Agent Core, Lambda, DynamoDB single-table design, S3, SQS, EventBridge, and Step Functions
Strong Python and Node.js skills
Experience building full-stack applications and RESTful APIs
Understanding of production reliability patterns for LLM-backed systems, including retries/backoff, circuit breakers, and fallback models
Experience with Docker and cloud-native deployment
Ability to own ambiguous, integration-heavy problems and provide technically deep answers
Helpful extras: direct Bedrock Agent Core experience; regulated or high-stakes domain experience; workflow orchestration familiarity; production LLM observability experience
AI Harness Engineer building reliable AI tools, agent harnesses, and P2P systems for Tether’s local AI platform. Developing QVAC features that connect users to scalable local AI.
Senior Director leading AI enablement and enterprise systems transformation for Cineflix, a global TV creator and distributor. Modernising systems, deploying internal LLM capabilities and driving organisation - wide adoption.
Director leading enterprise data governance, catalog, quality, and AI readiness for CBC/Radio - Canada’s public service media organization. Hybrid role guiding governance programs, councils, stakeholders, and data teams.
Director leading AI innovation pipelines, accelerators, governance, and enablement at CBC/Radio - Canada, Canada’s public broadcaster. Driving scalable AI initiatives across media, technology, and corporate teams.
Expert AI Governance Advisor defining responsible AI governance, risk controls, and regulatory practices at Beneva. Supporting insurance and financial services through ethical, compliant AI adoption.
AI personalization rater evaluating Gemini responses using connected Google applications. Providing structured feedback to improve accuracy, relevance, helpfulness, and personalization.
AI model risk governance leader developing enterprise frameworks, standards, and training for responsible AI at global financial services provider Manulife. Hybrid role partnering with regulated - industry stakeholders.
AI model risk oversight manager at Manulife, an international financial services provider. Challenging governance, validation, and controls for high - risk AI initiatives.
AI risk oversight manager at Manulife, a global financial services provider. Challenging AI model and operational risks across regulated enterprise portfolios.
AI Data Center Infrastructure Engineer designing racks, power, cabling, and network layouts for Cerebras’ large - scale AI chip systems. Automating deployment workflows and integrating compute infrastructure.