Technical lead building safe, evaluable LLM agents and AI infrastructure for OpenLoop’s telehealth platform. Setting architecture, observability, retrieval, and model operations direction.
Responsibilities
Set the technical direction for agent runtime and orchestration
Build evaluation systems, including regression tests and AI-graded evaluations
Define model access and operations, including provider routing, version upgrades, and outage planning
Own observability and cost attribution for AI systems
Build retrieval and grounding over company data
Translate human workflows into agent workflows and determine which work must remain human-led
Protect patient data and ensure security and privacy in AI systems
Make cloud and data infrastructure decisions on GCP and partner with the Data Platform team
Conduct code and design reviews and mentor engineers new to AI
Explain AI trade-offs to product, operations, and clinical stakeholders
Own AI architecture and technical quality while partnering with the AI Engineering Manager on people and delivery
Carry long-range technical direction for AI across the company until an AI Principal Engineer is hired
Requirements
12 or more years of software engineering experience, with real technical leadership experience (tech lead, staff, or principal-level scope)
Two or more years building and running LLM-based systems in production, not just demos or personal projects
Deep, hands-on expertise in agent runtime or evaluation, plus working knowledge across the rest of the AI engineering stack
A track record of measuring AI quality with real evaluation methods and data
Experience leading a small team's technical direction from zero, or close to it
Healthcare or other regulated-industry experience (PHI, HIPAA) preferred
GCP experience preferred
Experience building an internal platform used by other engineering teams preferred
Background in classic machine learning as well as LLMs preferred
Experience scaling a team from a single pod into multiple teams preferred
Authorization to work in the location of the job posting
Senior software engineer evaluating AI coding agents such as Codex and Claude Code. Providing rigorous written and video feedback on engineering quality.
Senior software engineer evaluating AI coding agents for G2i. Assessing engineering judgment, explanations, and trustworthiness across Codex, Claude Code, and Cursor.
Senior engineer evaluating AI coding agents such as Codex, Claude Code, and Cursor. Providing rigorous written and video feedback on engineering judgment, reasoning, and interaction quality.
Forward Deployed Engineer building full - stack AI solutions on AWS and Azure for Huron’s consulting clients. Partnering daily with business users to deliver tested features.
Staff AI Security Engineer securing EQ Bank’s enterprise AI, machine learning, and cloud platforms. Building guardrails, controls, automation, and detection capabilities for Canada’s Challenger Bank.
Staff Software Engineer owning OAuth, authorization, and agent delegation systems. Building governed identity infrastructure for Redpanda’s enterprise AI data platform.
Senior software engineer evaluating AI coding agents for G2i’s engineering team. Assessing reasoning, explanations, and engineering judgment in Codex, Claude Code, and Cursor interactions.
Senior software engineer evaluating AI coding agents such as Codex, Claude Code, and Cursor. Providing rigorous written and video feedback on engineering quality.
Senior software engineer evaluating Codex, Claude Code, and Cursor interactions. Providing rigorous written and video feedback on AI - generated coding quality for G2i.
Senior engineer evaluating Codex, Claude Code, and Cursor interactions for G2i. Providing rigorous written and video feedback on AI - generated engineering work.