AI benchmark engineer building French-Canadian Terminal-Bench coding evaluations for LILT. Creating multilingual datasets, prompts, reference implementations, and verifier scripts for AI models.
Responsibilities
Design, build, and validate Terminal-Bench tasks for multilingual software challenges
Evaluate coding agents
Build realistic task environments using datasets and files in French
Create prompts and translations in the native language to identify AI failure points
Develop robust reference implementations
Write reliable, deterministic verifier scripts
Analyze execution logs and calibrate task difficulty from Easy to Very Hard
Run standard Terminal-Bench configurations against Haiku, Sonnet, and Opus model tiers
Participate in a four-layer human quality-control process: creation, human review, calibration review, and audit
Work alongside automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity
Requirements
5+ years of industry experience in software engineering
Proven track record at leading technology companies and/or graduation from top-tier engineering universities
Native or near-native fluency in French, with deep understanding of grammar, register, and phrasing rules
High English proficiency
Strong proficiency in Python
Strong proficiency in standard shell scripting
Strong proficiency in data processing
Extensive experience with Terminal/CLI-based development workflows
Working familiarity with coding agents
Deep technical understanding of multilingual text processing pitfalls
Knowledge of encoding/decoding robustness and Unicode normalization
Knowledge of locale-dependent conventions, including collation, casing, and non-Gregorian dates
Knowledge of text I/O, toolchain interoperability, and safe string operations
For specific languages, knowledge of bidirectional/RTL handling, font fallbacks, and rendering/typography in UI or artifacts
Reliable availability and commitment to on-time delivery
Updated CV in English
Successful completion of a GenAI assessment
Benefits
Flexible schedule with no fixed hours, check-ins, or micromanaging
Competitive rates
Prompt payments
Access to diverse, innovative projects
Portfolio and skills development opportunities
Global community of linguists, subject matter experts, and language professionals
No health insurance, paid time off, or retirement contributions are provided
Hours are not guaranteed
Most tasks require a minimum of 2 hours per day or 10 hours per week
Head of AI Operations building Supabase’s AI - native operating system for its Postgres development platform. Establishing standards, assessments, automated reporting, and company - wide AI - enabled operations.
Administrateur IA assurant la gouvernance, la sécurité et les opérations de Microsoft 365 Copilot chez EDC. Soutien de l’IA d’entreprise pour aider les entreprises canadiennes à réussir à l’étranger.
AI Administrator managing Microsoft 365 Copilot, Entra ID, and AI governance for Export Development Canada. Supporting secure enterprise AI adoption across Ottawa, Toronto, or Montreal.
Generative AI Artist creating AI - driven visuals, sequences, and previs for Rodeo FX’s visual - effects and advertising productions. Building ComfyUI workflows and collaborating with compositing, 3D, and art - direction teams.
AI transformation leader selling and delivering enterprise change consulting at Prosci. Building AI prototypes, guiding organizational change, and growing executive client accounts.
Mila AI Safety research intern developing public - safety guardrails, benchmarking, and model evaluations. Collaborating with researchers and compute teams on applied machine - learning research in Montréal.
AI products lead building GenAI - powered communications surveillance and compliance products at RBC Capital Markets. Leading full - stack engineering, AI agents, and production delivery.
UNESCO consultant analyzing learning assessment data with AI - assisted methods. Developing policy insights, reproducible protocols, and national analyst capacity.
AI Enablement Lead building training, curriculum and metrics for Valtech’s AI - native digital experience consultancy. Enabling practitioners, new hires and teams through AI - native SDLC transformation.
Marketing Specialist advancing AI - powered content, automation, SEO, and analytics for EverCommerce’s EMHware and GoodTherapy healthcare platforms. Improving campaign performance and adoption across Canada.