Applied AI Engineer building evaluation pipelines, production agents, and reliable clinical AI for Tali’s healthcare platform.
Owning models, retrieval, routing, and observability across ambient scribe, billing, search, and clinical workflows.
Responsibilities
Own AI systems end to end, including models, prompts, retrieval, data, evaluations, and debugging
Build evaluation pipelines, automated judges, regression suites, human review processes, and supporting datasets
Measure audio-path quality using word and speaker error rates and audio-quality metrics
Diagnose and fix failure modes across prompts, retrieval, routing, models, audio capture, and fine-tuning
Build production agents with tool use, orchestration, guardrails, and failure recovery
Build supporting search, vector storage, and agent harness infrastructure
Debug visits end to end from audio to delivered clinical note and identify systemic issues
Analyze warehouse data to size problems and validate fixes
Manage model routing, staged rollouts, and attribution
Own AI systems in production, respond to quality alerts, identify causes, and determine fixes
Translate clinical complaints into problem statements, metrics, and actionable plans
Establish applied AI evaluation standards across Tali
Build evaluation tooling, debug traces, engage clinicians about failures, and improve infrastructure
Drive AI capability improvements and productionize the next generation of agents
Requirements
5+ years in production ML, applied AI, or research engineering
Experience owning production systems used by real users
Deep evaluation experience, including graders, regression suites, or judge pipelines
Experience with agentic systems involving multiple models, tool calls, retrieval, and failure recovery
Strong systems engineering skills in backend services, data pipelines, and observability
Data-centric approach to improving AI systems and feedback loops
Proficiency in Python and modern ML tooling
Ability to provide and receive candid technical feedback
Ability to make and communicate engineering and business cases and own outcomes
Ability to improve systems through code reviews and mentoring
Speech recognition or real-time audio experience is a bonus
Experience in a regulated domain such as healthcare or finance is a bonus
Clinical experience is a bonus
Experience with Python and TypeScript
Experience with GCP and Cloud Run
Experience with Vertex AI, Claude, and other frontier LLM and ASR providers
Benefits
Flexible work hours
Comprehensive health and wellness coverage from day one
Unmetered wellness days
Competitive PTO
Winter shutdown Dec 25 - Jan 1
Birthdays and Taliversaries
'Extra long' long weekends
$2000 annually in "Knowledge Dollars" to learn, grow, and level up
Technical lead building safe, evaluable LLM agents and AI infrastructure for OpenLoop’s telehealth platform. Setting architecture, observability, retrieval, and model operations direction.
Senior software engineer evaluating AI coding agents such as Codex and Claude Code. Providing rigorous written and video feedback on engineering quality.
Senior software engineer evaluating AI coding agents for G2i. Assessing engineering judgment, explanations, and trustworthiness across Codex, Claude Code, and Cursor.
Senior engineer evaluating AI coding agents such as Codex, Claude Code, and Cursor. Providing rigorous written and video feedback on engineering judgment, reasoning, and interaction quality.
Forward Deployed Engineer building full - stack AI solutions on AWS and Azure for Huron’s consulting clients. Partnering daily with business users to deliver tested features.
Staff AI Security Engineer securing EQ Bank’s enterprise AI, machine learning, and cloud platforms. Building guardrails, controls, automation, and detection capabilities for Canada’s Challenger Bank.
Staff Software Engineer owning OAuth, authorization, and agent delegation systems. Building governed identity infrastructure for Redpanda’s enterprise AI data platform.
Senior software engineer evaluating AI coding agents for G2i’s engineering team. Assessing reasoning, explanations, and engineering judgment in Codex, Claude Code, and Cursor interactions.
Senior software engineer evaluating AI coding agents such as Codex, Claude Code, and Cursor. Providing rigorous written and video feedback on engineering quality.
Senior engineer evaluating Codex, Claude Code, and Cursor interactions for G2i. Providing rigorous written and video feedback on AI - generated engineering work.