MLOps Engineer creating and evaluating training data for a leading AI lab's GenAI systems. Focusing on GPU kernels, profiling, distributed debugging, and high-throughput LLM serving.
Responsibilities
Design challenging, domain-relevant tasks across GPU kernels, performance profiling, debugging, and inference serving, and write accurate, well-structured solutions
Guide research and engineering teams to close knowledge gaps and improve AI model performance on ML systems, training infrastructure, and framework-level topics
Evaluate MLOps and ML systems tasks and solutions and provide clear, written technical feedback
Develop guidelines and detailed rubrics or evaluation frameworks covering kernel-level optimization, profiler output interpretation, distributed systems reasoning, and serving throughput and latency trade-offs
Collaborate with subject matter experts to keep training data consistent and accurate
Contribute to AI model training and evaluation work by writing and assessing MLOps and ML systems tasks and solutions for frontier AI training data
Requirements
2+ years of hands-on professional experience in ML systems, ML infrastructure, model serving, or GPU and accelerator performance engineering
Practical experience in at least one of: writing or optimizing custom GPU kernels (CUDA, Triton, Pallas); performance profiling and trace analysis (Kineto, torch.profiler, Nsight, XLA or JAX profiler); debugging distributed or accelerator-bound workloads; serving large language models at scale (vLLM, SGLang, TensorRT-LLM, Ray Serve, KV cache, paged attention, continuous batching)
Working production experience with JAX and/or PyTorch
Familiarity with modern accelerators such as A100, H100, B200 or TPU
Ability to reason about throughput, latency and memory trade-offs
Demonstrable career progression
Ability to engage reliably for at least 40 hours/week during weekdays
Strong written communication skills and ability to explain complex technical decisions clearly
Engineer joining Cerebras to rapidly deploy AI models on proprietary CSX systems. Focused on debugging and enhancing model performance in a dynamic, innovative environment.
Staff Generative AI Engineer developing production - grade AI applications for business value across the organization. Collaborating with engineers, scientists, and product teams to deliver scalable solutions.
Senior NLP/LLM Engineer exploring and analyzing LLM capabilities while collaborating with cross - functional teams at Social Discovery Group. Enhancing AI models' effectiveness and optimizing their performance.
Design and implement AI - powered applications for wealth management using GenAI and LLMs, collaborating with business teams to enhance client engagement and operational efficiency.
Senior Gen AI Developer (LLM) role in Toronto (Hybrid) for 12 months. Requires 6 - 8 yrs experience with LLMs, GenAI, Java, Spring, AI Agents, MLOps, CI/CD.
Lead AI solution design & development, integrate LLMs into Java/Spring apps for NLP & predictive analytics. Collaborate cross - functionally on AI strategy & mentor junior developers.
AI Developer for full lifecycle LLM solutions. Build production - ready AI systems, develop data pipelines, deploy models in cloud environments, and collaborate with stakeholders.
Senior Generative AI Engineer role requiring strong GenAI, Python & LLM experience. Onsite position in Mississauga, Ontario with contract/full - time options.
Machine Learning Engineer specializing in large language models for John Snow Labs. Working on AI model training and optimization for healthcare applications in a fully remote setting.