QA Engineer, AI Systems

Posted 2 days ago

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • QA Engineer building evaluation infrastructure for Nexxa’s autonomous AI agents in heavy industries. Testing reasoning, tool use, safety, and reliability across industrial workflows.

Responsibilities

  • Own quality for Nexxa's AI agent systems used in industrial environments
  • Design and build evaluation harnesses and regression suites for LLM-based agents
  • Develop golden datasets and labeled test sets covering edge cases, ambiguous inputs, and adversarial prompts
  • Define and track quality metrics including groundedness, hallucination rate, task success rate, latency/cost tradeoffs, and safety violations
  • Build automated evaluation pipelines for model, prompt, and tool-integration changes and integrate them into CI/CD
  • Conduct structured red-teaming and adversarial testing, including prompt injection, jailbreaks, tool misuse, and unsafe actions
  • Test the full agent action loop: planning, tool selection, tool execution, error recovery, and final output
  • Investigate and triage failures involving models, prompts, tools/APIs, or orchestration logic
  • Partner with ML and backend engineers to produce actionable, reproducible bug reports
  • Establish quality bars and sign-off criteria for new agent capabilities
  • Mentor engineers on testing strategies for probabilistic, LLM-driven systems
  • Advocate for testability and observability in agent architecture
  • Collaborate with ML engineers, backend engineers, and Forward Deployed Engineers

Requirements

  • 5+ years in QA/SDET roles, with demonstrated ownership of test strategy for complex systems
  • Hands-on experience testing LLM-based products, chatbots, or AI agents
  • Practical experience with eval frameworks or tooling such as promptfoo, DeepEval, RAGAS, or LangSmith, or a track record of building your own
  • Strong scripting/programming ability, Python preferred
  • Understanding of LLM agents, including prompting, tool/function calling, context management, RAG, memory, and orchestration frameworks
  • Experience designing test data and labeled datasets, including sourcing, sampling, and managing dataset drift
  • Familiarity with LLM-specific failure modes, including hallucination, prompt injection, context poisoning, tool misuse, goal drift, and non-determinism
  • Strong written communication skills
  • Experience with human-in-the-loop evaluation workflows
  • Background in ML/data science sufficient to read model evals and statistical significance
  • Experience red-teaming or conducting adversarial/security testing on ML systems
  • Familiarity with LLM observability/tracing tools such as LangSmith, Arize, Langfuse, or Weights & Biases
  • Experience testing AI systems in industrial, IoT, or operational technology environments
  • Prior experience setting up eval infrastructure from scratch at a startup or fast-moving team

Benefits

  • Equity package
  • Significant opportunities for career development and advancement
  • Collaborative culture valuing innovation, discipline, and continuous improvement
  • Competitive compensation

Job title

Job type

Full Time

Experience level

Mid levelSenior

Salary

Not specified

Degree requirement

No Education Requirement

Tech skills

IoTPython

Location requirements

HybridTorontoCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.