Senior Data Scientist – AI Evaluation

Posted 5 hours ago

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Senior Data Scientist defining rigorous evaluation systems for AI models and agents. Improving quality and safe rollout at Alpaca, a global brokerage-infrastructure company.

Responsibilities

  • Design AI evaluations by defining ground truth, metrics, and scoring methods for models and agents
  • Build repeatable evaluation loops to track quality over time and catch regressions before release
  • Partner with engineering and analytics engineering to operationalize evaluation harnesses
  • Translate evaluation results into actionable recommendations for system improvements
  • Establish evaluation guidelines, documentation, and review practices
  • Mentor and align teams on evaluation best practices and measurable AI quality
  • Partner with Product, Engineering, Analytics Engineering, and business stakeholders to define quality standards and drive iteration
  • Own the quality bar independently of the teams that build and optimize the systems

Requirements

  • Track record of quantitative measurement rigor (e.g., LLM/model evaluation, metric validation, or experimentation)
  • Strong statistical and ML foundation, including sample sizing, confidence intervals, significance, handling non-determinism, and validating automated graders against human ground truth
  • Proficiency in Python and SQL, with experience evaluating models in production environments
  • Strong judgment in defining quality metrics and ground truth for ambiguous outputs
  • Excellent communication and cross-functional collaboration skills
  • Strong problem-solving ability in fast-paced, greenfield environments
  • 6–10 years in quantitative data science or ML, with focused experience in measurement or evaluation
  • Quantitative degree is a plus; equivalent industry experience is equally welcome
  • Nice to have: hands-on LLM/agent evaluation in production, eval harnesses, LLM-as-judge calibration, and CI regression gates
  • Nice to have: experience evaluating text-to-SQL, analytics agents, or systems where correctness is verifiable against data
  • Nice to have: background in fintech, brokerage, or domains with business or risk consequences
  • Nice to have: fluency with AI tools in research and engineering workflows

Benefits

  • Competitive Salary & Stock Options
  • Health Benefits
  • New Hire Home-Office Setup: One-time USD $500
  • Monthly Stipend: USD $150 per month via a Brex Card

Job type

Full Time

Experience level

Senior

Salary

Not specified

Degree requirement

No Education Requirement

Tech skills

PythonSQL

Location requirements

RemoteNorth America

Report this job

Found something wrong with the page? Please let us know by submitting a report below.