Senior software engineer evaluating AI coding agents such as Codex, Claude Code, and Cursor. Providing rigorous written and video feedback on engineering quality.
Responsibilities
Evaluate AI-generated coding interactions end to end
Judge whether outputs are useful, correct at a high level, and aligned with strong engineering judgment
Assess explanations and reasoning, not just code
Distinguish between response quality levels
Provide clear, opinionated written feedback on what worked, failed, or felt misleading
Help define what excellent interactions look like for Cursor, Codex, and Claude Code
Record short video explanations
Work on projects lasting approximately two weeks to a few months, with new projects offered based on performance
Requirements
Senior, Staff, or Principal-level software engineer (or equivalent experience)
Strong background in TypeScript/JavaScript or Python
Hands-on experience with at least one of OpenAI Codex, Claude Code, or Cursor
Deep familiarity with modern AI-assisted development workflows
Ability to evaluate code without executing it or reviewing every line
Strong written and spoken English (B2 or above)
Comfortable giving direct, opinionated feedback
High standards for engineering quality
Prior exposure to prompt design or evaluation workflows (nice to have)
Experience mentoring senior engineers or defining engineering standards (nice to have)
Must complete and pass a take-home evaluation exercise with a recorded Loom walkthrough
Must agree to a simple background check
Benefits
Flexible hours, 10–20 hours/week
Ongoing project opportunities for evaluators who perform well
Staff Software Engineer owning OAuth, authorization, and agent delegation systems. Building governed identity infrastructure for Redpanda’s enterprise AI data platform.
Senior software engineer evaluating AI coding agents for G2i’s engineering team. Assessing reasoning, explanations, and engineering judgment in Codex, Claude Code, and Cursor interactions.
Senior software engineer evaluating Codex, Claude Code, and Cursor interactions. Providing rigorous written and video feedback on AI - generated coding quality for G2i.
Senior software engineer evaluating Codex, Claude Code, and Cursor interactions for G2i’s engineering team. Providing rigorous feedback on AI reasoning, code quality, and developer trust.
Senior engineer evaluating Codex, Claude Code, and Cursor interactions for G2i. Providing rigorous written and video feedback on AI - generated engineering work.
Applied AI Lead deploying Heidi’s healthcare AI for strategic healthcare accounts. Driving adoption, expansion, renewal, and measurable clinical and operational value.
Software Engineer, AI building secure, production - ready LLM and cloud - native solutions for Softchoice, an IT solutions provider. Delivering RAG, APIs, automation, and enterprise AI integrations.
Full - stack developer building TypeScript AI customer - support agents for gaiia’s telco operating system. Developing agent runtimes, omnichannel integrations, evaluation tooling, and safety guardrails.
Senior AI Engineer embedding with clinical, product, and operations teams. Shipping validated, monitored AI features for Prenuvo’s proactive whole - body healthcare platform.
Backend Software Engineer creating realistic coding challenges, bug fixes, and deterministic verifiers for Gramian’s AI training projects. Evaluating backend code quality, reliability, testing, and performance.