Senior software engineer evaluating AI coding agents for G2i. Assessing engineering judgment, explanations, and trustworthiness across Codex, Claude Code, and Cursor.
Responsibilities
Evaluate AI-generated coding interactions end to end
Judge whether outputs are useful, correct at a high level, and aligned with how a strong engineer would think
Assess the quality of explanations and reasoning, not just the code
Distinguish between levels of response quality
Give clear, opinionated written feedback on what worked, what did not, and what felt off or misleading
Help define what great looks like when working with Cursor, Codex and Claude Code
Record short video explanations
Projects run from about two weeks to a few months each
Requirements
Senior, Staff or Principal-level engineer (or equivalent experience)
Strong background in TypeScript/JavaScript or Python
Hands-on experience with at least one of OpenAI Codex, Claude Code or Cursor
Deep familiarity with modern AI-assisted development workflows
Able to evaluate code without executing it or reviewing every line
Strong written and spoken English (B2 or above)
Comfortable giving direct, opinionated feedback
High bar for what good engineering looks like
Prior exposure to prompt design or evaluation workflows is nice to have
Experience mentoring senior engineers or defining engineering standards is nice to have
Must complete a take-home evaluation exercise with a recorded Loom walkthrough
Must agree to a simple background check for the project
Benefits
Flexible schedule, 10-20 hours/week
Ongoing project opportunities for evaluators who perform well
Start as soon as the take-home is cleared and a project seat is available
Technical lead building safe, evaluable LLM agents and AI infrastructure for OpenLoop’s telehealth platform. Setting architecture, observability, retrieval, and model operations direction.
Senior software engineer evaluating AI coding agents such as Codex and Claude Code. Providing rigorous written and video feedback on engineering quality.
Senior engineer evaluating AI coding agents such as Codex, Claude Code, and Cursor. Providing rigorous written and video feedback on engineering judgment, reasoning, and interaction quality.
Forward Deployed Engineer building full - stack AI solutions on AWS and Azure for Huron’s consulting clients. Partnering daily with business users to deliver tested features.
Staff AI Security Engineer securing EQ Bank’s enterprise AI, machine learning, and cloud platforms. Building guardrails, controls, automation, and detection capabilities for Canada’s Challenger Bank.
Staff Software Engineer owning OAuth, authorization, and agent delegation systems. Building governed identity infrastructure for Redpanda’s enterprise AI data platform.
Senior software engineer evaluating AI coding agents for G2i’s engineering team. Assessing reasoning, explanations, and engineering judgment in Codex, Claude Code, and Cursor interactions.
Senior software engineer evaluating AI coding agents such as Codex, Claude Code, and Cursor. Providing rigorous written and video feedback on engineering quality.
Senior software engineer evaluating Codex, Claude Code, and Cursor interactions. Providing rigorous written and video feedback on AI - generated coding quality for G2i.
Senior engineer evaluating Codex, Claude Code, and Cursor interactions for G2i. Providing rigorous written and video feedback on AI - generated engineering work.