Staff ML engineer building models, evaluations, and agentic systems for Sourcegraph’s code-understanding products. Improving enterprise code search quality, latency, cost, and reliability.
Responsibilities
Set the Code Understanding team's direction for models, evaluations, and agentic systems
Design and harden multi-step, tool-using agent loops for reliable, observable, affordable enterprise-scale products
Determine when and how to use evaluations, smoke tests, metrics, and qualitative review
Select, upgrade, fine-tune, and train models where appropriate
Improve retrieval, ranking, context windows, and citations for code-grounded answers
Optimize model cost and latency through profiling, distillation, caching, and right-sizing
Own meaningful agentic product slices end-to-end from problem framing through rollout and measurement
Establish evaluations, dashboards, and guardrails for responsible model and prompt changes
Mentor and up-level teammates in agent engineering
Engage with customers and translate feedback into requirements, scopes, and milestones
Contribute across the codebase and influence technical direction beyond the immediate team
Participate in the on-call support rotation
Drive roadmap direction and measurable improvements in answer quality, cost, latency, and agentic capabilities
Requirements
Staff engineer and technical leader with production machine learning, evaluation, and agent systems expertise
Personally owned a production model lifecycle from dataset construction through evaluation, production rollout, and monitoring
Trained or fine-tuned at least one model
Experience designing reliable, observable, and cost-bounded multi-step agentic systems
Strong evaluation judgment, including representative datasets, baselines, error taxonomies, and release criteria
Ability to make quality, latency, and cost tradeoffs using model selection, prompting, retrieval, caching, distillation, and fine-tuning
Ability to operate autonomously on ambiguous, high-technical-risk problems
Strong software engineering skills and ability to ship production services
Comfortable across Go, TypeScript, GraphQL, Postgres, and Docker, or clearly able and eager to learn them
Fluent with agentic coding tools and able to understand and own every line they submit
Comfortable in an async-first, multi-service, fast-paced remote environment
Ability to mentor engineers through pairing, design reviews, and code reviews
Customer- and product-driven approach
Working hours must overlap with EST for at least 20 hours per week
Benefits
Meaningful equity
Competitive cash compensation
Generous perks and benefits
Open and transparent compensation philosophy
Pay bands designed for competitive and equitable compensation
Globally distributed remote work arrangement
Flexible location options in almost any part of the world
Applied Machine Learning Scientist developing Generative AI and predictive ML solutions at TD, a major North American bank. Supporting model evaluation, deployment, monitoring, and responsible AI governance.
Applied Machine Learning Scientist developing Generative AI and predictive ML solutions for TD banking. Evaluating models, managing AI lifecycles, and supporting responsible implementation.
Senior Machine Learning Engineer building conversational AI agents and production ML systems. Helping Numa automate automotive dealership service and sales through evaluation - first tooling and infrastructure.
Senior ML Engineer building production ML, RAG, and agentic AI capabilities for SailPoint’s cloud identity security platform. Driving scalable, customer - focused AI solutions from research to production.
Senior ML Engineer developing debiased pCTR and conversion models for Instacart’s grocery advertising ecosystem. Advancing ranking, retrieval, and sequence modeling across ads surfaces.
Senior AI/ML Engineer building Generative AI, RAG, and agentic solutions for pharmaceutical Statistical Programming. Deploying secure, validated, production - ready AI applications with Python and AWS.
Senior Machine Learning Engineer building LLM - powered lab interpretation and clinical decision - support tools for Fullscript’s healthcare platform. Owning AI systems from prototyping through production.
Graduate co - op scientist at TD, a global financial institution, developing machine learning models and analytics solutions. Querying data, engineering features, and translating insights into banking decisions.