Senior AI Quality Engineer defining evaluation standards and automation for LLM and agentic workflows. Testing AI-powered safety and operational risk-management software at Novara.
Responsibilities
Design and maintain evaluation frameworks for LLM outputs and agentic workflows, including regression suites, golden datasets, and scoring rubrics
Build test harnesses that catch hallucinations, tool-calling failures, prompt regressions, and unsafe or off-policy behavior before they reach production
Define measurable quality criteria for agent reliability: task completion, factual grounding, latency, cost, and reasoning quality
Integrate evaluation runs into CI/CD so model, prompt, and agent changes are gated the same way code changes are
Partner with engineers on observability and tracing for agent runs, so failures are diagnosable rather than mysterious
Contribute to conventional API and end-to-end automation where AI features sit inside larger product flows
Help shape the team's shared playbook for testing AI features, and coach other QA engineers as agentic work spreads across scrum teams
Work hands-on as an individual contributor reporting to the QA Manager
Partner with product, engineering, and the broader QA team on LLM-driven and agentic workflows
Requirements
6+ years in QA, SDET, or test automation, with real production automation shipping
Hands-on experience testing LLM-based or agentic systems: building evals, working with LLM-as-judge patterns, prompt regression testing, or agent trajectory analysis
Prior experience in a shift-left, embedded QA model
Comfort with at least one modern automation stack (Playwright, Cypress, or similar) and a typed language, TypeScript preferred, Python fine
Deep experience with test frameworks such as vitest, jest, or pytest, and comfort building custom test harnesses rather than only running off-the-shelf suites
API-first testing mindset, including REST and Postman or equivalent
Fluency in HTTP-level API testing, including recording proxies and observing service-to-service traffic
Working knowledge of CI/CD pipelines, GitHub Actions a plus, and how to plug evals into them
Familiarity with cloud secret managers and disciplined handling of sensitive test data in restore-from-prod environments
Ability to reason clearly about probabilistic systems: variance, sample sizes, confidence, and when a flaky result is signal rather than noise
QA Automation Engineer advancing AI - first testing for Quest’s On Demand Migration SaaS on Azure. Building automated tests and cloud QA environments with C#/.NET.
SQA Engineer automating testing and risk - based validation for EDC, Canada’s export - credit and trade - finance Crown corporation. Improving reliability across integrated digital systems.
QA Tester validating SAP, CRM, and AWS data integration for Circular Materials, a national recycling producer responsibility organization. Executing integration tests, reconciliation, and defect resolution.
QA Team Lead overseeing laboratory testing, food safety, and quality systems at Nestlé Canada’s Toronto factory. Supervising technical teams and supporting audits, traceability, and product release.
Quality Assurance Analyst testing Ava Industries’ cloud - based EMR software. Validating clinic workflows, releases, defects, and user experience for Canadian healthcare providers.
Ingénieur qualité logicielle chez EDC, société canadienne aidant les entreprises à réussir à l’étranger. Automatisation des tests, validation assistée par l’IA et intégration de bout en bout.
QA Analyst testing Creator.co’s influencer marketing platform across web and mobile. Applying AI tools to manual testing, bug reporting, and emerging test automation.
FCRM Quality Assurance Analyst reviewing high - risk customers and EDD compliance for TD, a major North American bank. Supporting AML controls, regulatory reporting and financial crime risk programs.
QA Audio Rater auditing production audio ratings and correcting calibration issues. Supporting Welo Data’s AI services program remotely from Canada for 10 hours weekly.