Independent contractor designing complex evaluation frameworks for AI audio models. Focusing on role-play scenarios, auditing AI models, and generating training datasets.
Responsibilities
Operate autonomously to design complex evaluation frameworks and provide structured training data.
Role-Play Scenario Execution: Creating and executing complex, role-play-based evaluation scenarios that simulate realistic customer service interactions.
Model Performance Auditing: Evaluating AI model performance across standardized qualitative and quantitative metrics.
Technical Metric Evaluation: Assessing the model's basic computer programming literacy.
Representative Dataset Generation: Contributing to the development of diverse, high-quality audio datasets.
Requirements
Demonstrable professional expertise in complex customer support, technical troubleshooting, or conversational AI evaluation.
Native or bilingual proficiency in the target language, including fluency across all language skills (reading, listening, writing, and speaking).
Strong analytical and verbal communication skills to confidently conduct simulated customer support role-plays.
Basic computer programming literacy, specifically a comfortable understanding of JSON structures, functions, methods, and simple logic.
A meticulous, detail-oriented approach to working with structured prompts, complex evaluation rubrics, and technical guidelines.
Required Equipment: Access to a high-quality microphone to ensure clean, reliable audio input during voice evaluations.
Benefits
Company-sponsored benefits such as health insurance and PTO do not apply
Business Logic Expert reviewing AI - generated business content for professional accuracy and etiquette. Helping TELUS Digital improve AI systems through remote, flexible freelance work.
Forward Deployed Engineer building production - oriented agentic AI solutions for enterprise clients. Translating executive business needs into rapid proofs of value and scalable architectures.
Hebrew language expert evaluating audio for nativeness, pronunciation, and authenticity. Providing English feedback to improve AI language understanding and generation.
AI solutions strategy director scaling Instacart’s grocery - retailer technology platform. Developing commercialization, pricing, and go - to - market plans for retailer - focused AI products.
Experienced Vietnamese - speaking Psychiatrist developing psychiatric cases and evaluating AI mental - health content. Gramian Consultancy connects engineering and professional talent with organizations.
Psychology PhD evaluating AI - generated psychological content for Gramian Consultancy, an IT professional - services and engineering talent consultancy. Developing culturally sensitive Vietnamese - English scenarios and expert evaluations.
VP leading platform - agnostic ERP, CRM, and AI delivery teams at Appficiency, a technology and professional services firm. Driving client strategy, sales alignment, and AI - enabled delivery efficiency.
Technology, data and AI advisory leader building Leo Berwick’s power, utilities and renewables practice. Leading diligence, AI value creation and M&A technology work for PE and infrastructure clients.
Freelance AI Linguistic Evaluator assessing Malayalam AI responses for RWS. Translating English prompts and evaluating linguistic accuracy, localization, and cultural relevance.