Lead Inference Platform Engineer focused on optimizing ML models for high-performance inference at Thomson Reuters. Collaborating with engineering teams and deploying AI workloads efficiently.
Responsibilities
Optimize LLMs and ML models for high-performance inference using techniques such as quantization, pruning, distillation, and hardware specific tuning
Deploy and scale inference workloads on GPUs across AWS, Azure, GCP and internal Kubernetes clusters, ensuring predictable performance during peak traffic hours, especially during business hours
Implement routing and failover strategies for OpenAI/Anthropic/Vertex AI traffic
Integrate models into production grade APIs supporting TR products and enterprise workflows
Develop highly optimized environment and eliminate performance bottlenecks to reduce latency
Collaborate with Platform Engineering teams (Landing Zones, Network, Storage, Compute, AI) to ensure inference workloads align with TR’s cloud native patterns (AWS, Azure, GCP, OCI)
Build and optimize containerized inference pipelines using Kubernetes for large‑scale distributed workloads
Ensure compliance with TR’s AI standards for deployment, monitoring, governance, and drift detection
Profile inference performance, identify GPU/CPU bottlenecks, and optimize compute utilization across heterogeneous hardware
Implement observability and health monitoring for inference pipelines, ensuring reliability of enterprise AI services
Collaborate with platform teams to enhance capacity forecasting for AI workloads
Work with Product, Data Science, Architecture, and Enterprise AI teams to onboard new research models into production
Collaborates closely with AI engineers to invent new quantization techniques, improve numerical precision, and explore non‑standard architectures
Partner with Cloud Engineers (Azure, AWS, GCP) to develop guardrails and automation that support inference workload
Support the scale out of AI infrastructure during critical releases and global product rollouts.
Requirements
Strong understanding of ML/LLM fundamentals and inference optimization techniques
Hands-on experience with GPU programming (CUDA preferred), inference runtimes (TensorRT, ONNX Runtime), and deep learning frameworks (PyTorch/TensorFlow)
Proficiency in Python and at least one systems language (C++ strongly preferred for performance critical inference paths)
Experience deploying AI workloads to AWS/GCP/Azure and Kubernetes
Familiarity with vector search systems (OpenSearch vectors) and retrieval augmented generation pipelines
Knowledge of distributed systems, microservices, CI/CD, and cloud native architecture
Experience with AI networks, such as CNNs, transformers, and diffusion model architectures, and their performance characteristics
Understanding of GPU, Multithreading and/or other accelerators with vectorized instructions
Specialized experience in one or more of the following machine learning/deep learning domains: Model compression, hardware aware model optimizations, hardware accelerators architecture, GPU/ASIC architecture, machine learning compilers, high performance computing, performance optimizations, numerics and SW/HW co-design.
Benefits
Flexible vacation
Two company-wide Mental Health Days off
Access to the Headspace app
Retirement savings
Tuition reimbursement
Employee incentive programs
Resources for mental, physical, and financial wellbeing
Staff Platform Engineer defining cloud and DevSecOps strategy for Robots & Pencils’ enterprise AI systems. Leading Kubernetes, AI/ML infrastructure, migrations, reliability, security, and platform standards remotely in Canada.
Principal platform developer designing AWS - native integrations, CI/CD pipelines, and developer tooling for Autodesk’s design software. Leading architecture, reliability, and cross - team engineering initiatives.
Senior MLOps Developer operationalizing machine learning models and scalable AI/ML infrastructure for Autodesk’s design and entertainment software. Building deployment, monitoring, governance, and recovery systems.
Senior SRE/Platform Engineer needed for global company. 6+ years SRE experience, AWS/Azure, Kubernetes, Terraform, observability tools. Contract - to - hire in Mississauga.
Senior Full Stack Engineer building GraphQL, React, and TypeScript platforms for PENN Entertainment’s online gaming and sports media products. Improving shared client tooling, server - driven UI, performance, observability, and release workflows.
Senior platform engineering lead shaping compute and virtualization strategy for BMO, a major bank. Driving modernization, architecture standards, automation and hybrid - cloud infrastructure transformation.
Software Engineer building backend AI platform systems for DraftKings’ sports entertainment and gaming technology. Developing retrieval, vector database, automation, and agent infrastructure for scalable AI applications.
Software Engineer building AI platform backend systems, integrations, and retrieval infrastructure for DraftKings’ digital sports entertainment and gaming products. Developing scalable services, automation, and LLM - powered applications.