Senior Machine Learning Operations Developer, Inference, AI/ML Platform

Posted last week

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Senior MLOps Developer operationalizing machine learning models and scalable AI/ML infrastructure for Autodesk’s design and entertainment software. Building deployment, monitoring, governance, and recovery systems.

Responsibilities

  • Drive operational excellence of Autodesk’s AI/ML Platform by implementing and optimizing MLOps practices
  • Design and implement automated deployment pipelines for machine learning models from development to production
  • Design, implement, and maintain scalable infrastructure for model training, inference, and data processing
  • Develop and maintain monitoring and logging systems for model performance, system health, and platform efficiency
  • Work with data developers on efficient training and validation data pipelines
  • Implement model version control and contribute to model governance and compliance practices
  • Uphold data privacy, ethical considerations, security best practices, and platform security
  • Identify process automation and optimization opportunities across the MLOps lifecycle
  • Identify and resolve operational issues and contribute to incident response and system recovery
  • Collaborate with research, product engineering, data developers, software developers, and researchers across Autodesk’s design, construction, manufacturing, and media and entertainment domains

Requirements

  • BS or MS in Computer Science or a related field
  • 5+ years of hands-on DevOps and MLOps experience deploying and managing machine learning models in production
  • Proficiency with Infrastructure as Code using Terraform or Ansible
  • Strong expertise with Docker and Kubernetes for orchestrating and scaling machine learning workloads
  • Experience setting up and managing CI/CD pipelines for machine learning projects
  • Strong scripting skills in Python, Bash, or similar languages
  • Familiarity with Prometheus, Grafana, ELK Stack, or similar monitoring and logging tools
  • Understanding of MLOps security best practices, including data encryption, access controls, and compliance standards
  • Collaboration and communication skills for working with cross-functional teams
  • Ability to troubleshoot and resolve complex operational issues promptly
  • Preferred: experience with AWS or Azure
  • Preferred: familiarity with SQL, NoSQL, or data lakes
  • Preferred: exposure to TensorFlow or PyTorch
  • Preferred: Git and Jira experience
  • Preferred: Agile development methodologies

Benefits

  • Starting base salary between $123,000 and $180,400 for Canada-based roles
  • Annual cash bonuses may be included
  • Stock grants may be included
  • Comprehensive benefits package
  • In-person onboarding and/or in-person ID verification may be required

Job type

Full Time

Experience level

Senior

Salary

CA$123,000 - CA$180,400 per year

Degree requirement

Bachelor's Degree

Tech skills

AnsibleAWSAzureDockerGrafanaKubernetesNoSQLPrometheusPythonPyTorchSQLTensorflowTerraform

Location requirements

OnsiteTorontoCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.