DevOps Engineer designing, building, and optimizing cloud infrastructure for machine learning operations at a gaming company. Scaling AI models for production and ensuring system reliability and performance.
Responsibilities
Manage, configure, and automate cloud infrastructure using tools such as Terraform and Ansible.
Implement CI/CD pipelines for ML models and data workflows, focusing on automation, versioning, rollback, and monitoring with tools like Vertex AI, Jenkins, and DataDog.
Build and maintain scalable data and feature pipelines for both real-time and batch processing using BigQuery, BigTable, Dataflow, Composer, Pub/Sub, and Cloud Run.
Set up infrastructure for model monitoring and observability — detecting drift, bias, and performance issues using Vertex AI Model Monitoring and custom dashboards.
Optimize inference performance, improving latency and cost-efficiency of AI workloads.
Ensure overall system reliability, scalability, and performance across the ML/Data platform.
Define and implement infrastructure best practices for deployment, monitoring, logging, and security.
Troubleshoot complex issues affecting ML/Data pipelines and production systems.
Ensure compliance with data governance, security, and regulatory standards, especially for real-money gaming environments.
Requirements
3+ years of experience as a DevOps Engineer, ideally with a focus on ML and Data infrastructure.
Strong hands-on experience with Google Cloud Platform (GCP) — especially BigQuery, Dataflow, Vertex AI, Cloud Run, and Pub/Sub.
Proficiency with Terraform (and bonus points for Ansible).
Solid grasp of containerization (Docker, Kubernetes) and orchestration platforms like GKE.
Experience building and maintaining CI/CD pipelines, preferably with Jenkins.
Strong understanding of monitoring and logging best practices for cloud and data systems.
Scripting experience with Python, Groovy, or Shell.
Familiarity with AI orchestration frameworks (LangGraph or LangChain) is a plus.
Bonus points if you’ve worked in gaming, real-time fraud detection, or AI-driven personalization systems.
DevOps Intern supporting CI/CD, cloud infrastructure, and automation for Ludia’s mobile game studio. Improving reliability and developer tools in production game environments.
Manager leading global SRE teams for Akamai's distributed Cloud IAM services. Improving reliability, scalability, security, and usability through cloud - native tooling and software.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.