Senior MLOps Engineer at Deep Genomics, maintaining ML infrastructure for drug discovery. Enjoy collaborating with scientists to ensure reliable and scalable ML systems in an innovative environment.
Responsibilities
Maintain and improve cloud infrastructure (GCP) using Infrastructure-as-Code tools (Terraform).
Manage IAM, RBAC, and permission policies across cloud environments.
Own and evolve CI/CD pipelines (CircleCI, GitHub Actions) and ensure best practices are followed across the engineering and ML teams.
Administer and support workflow orchestration platforms (e.g., Seqera/Nextflow, Argo, Kubeflow).
Operate and configure ML experiment tracking and registry tooling (e.g., W&B, MLflow).
Build and maintain containerized environments (Docker) and manage Kubernetes clusters.
Manage GPU resources – provisioning, scheduling, and debugging hardware and driver issues.
Write and maintain Python tooling, scripts, and integrations that support ML infrastructure.
Help deploy ML models to production environments and monitor their performance.
Requirements
4+ years of experience operating production infrastructure.
Proficiency with cloud platforms (GCP preferred; AWS/Azure acceptable) and Infrastructure-as-Code (Terraform).
Extensive Hands-on experience with Kubernetes and containerization (Docker).
Solid background in CI/CD systems (CircleCI, GitHub Actions, or similar).
Familiarity with Python package and environment management (e.g., pip, conda, pixi).
Strong Python programming skills.
Self-motivated problem solver with excellent communication skills.
Benefits
Highly competitive compensation, including meaningful stock ownership.
Comprehensive benefits - including health, vision, and dental coverage for employees and families, employee and family assistance program.
Flexible work environment - including flexible hours, extended long weekends, holiday shutdown, unlimited personal days.
Maternity and parental leave top-up coverage, as well as new parent paid time off.
Focus on learning and growth for all employees - learning and development budget & lunch and learns.
Facilities located in the heart of Toronto - the epicenter of machine learning and AI research and development, and in Kendall Square, Cambridge, Mass. - a global center of biotechnology and life sciences.
Machine Learning Specialist developing AI, models, and analytical data products for the Government of Alberta. Applying machine learning to improve public services and policymaking.
Machine Learning Engineer developing production AI models and scalable MLOps platforms for Wave, which helps small businesses thrive. Collaborating on financial - risk applications, governance, observability, and reliable deployment.
Senior Geospatial ML Engineer building satellite - imagery vegetation intelligence for a climate - tech company. Improving production models and pipelines to help utilities prevent wildfires and outages.
Senior Machine Learning Engineer building scalable recommender systems for Thomson Reuters’ legal, tax, compliance, government, and media platforms. Implementing secure ML products and infrastructure.
ML Engineer scaling Torc Robotics’ simulation platform for autonomous trucks. Embedding with Autonomy teams to operationalize replay, recompute, metrics, visualization, and model integrations.
Senior ML Engineer building production AI services and MLOps platforms for Hyatt, a global hospitality company. Optimizing cloud - based inference, infrastructure, and model operations.
Machine learning engineer developing real - time underwriting models for Affirm’s buy - now - pay - later platform. Productionizing risk systems, feature pipelines, monitoring, and experimentation workflows.
Senior Machine Learning Engineer productizing scalable deep learning models for Movable Ink’s content personalization and AI decisioning platform. Advancing ML infrastructure, inference performance, monitoring, and automated testing.
Adversarial ML Engineer red - teaming foundation models and AI systems for client security teams. Investigating vulnerabilities, designing attacks and defenses, and validating remediation across model and agentic layers.
ML Performance Benchmarking Engineer at Cerebras Systems, transforming AI inference with cutting - edge technology. Collaborating with teams to optimize performance across the software stack.