Staff ML Engineer – AWS Trainium, SageMaker

Posted last week

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Staff ML Engineer building production PyTorch training pipelines on AWS Trainium and SageMaker. Delivering cost-aware, hardware-optimized AI systems for enterprise clients.

Responsibilities

  • Train and operate models on Amazon SageMaker with AWS Trainium as the underlying compute
  • Write and optimize PyTorch training code for Trainium, including reasoning about NeuronCore architecture, compiler behavior, memory, and throughput tradeoffs
  • Diagnose hardware-specific training run issues and distinguish data or code problems from compiler- or device-level problems
  • Translate Trainium job requests into working, cost-aware end-to-end training pipelines
  • Tune distributed training runs for throughput and cost on SageMaker training infrastructure
  • Work directly with client and internal engineering teams to scope and deliver production training workloads

Requirements

  • Strong, hands-on PyTorch experience, ideally including distributed or multi-device training
  • Production experience with Amazon SageMaker for training and/or inference
  • Comfort working close to the hardware layer, including device-specific compilation and accelerator-level debugging
  • AWS Trainium or Inferentia (Neuron SDK) experience is a strong plus
  • Deep PyTorch experience and a track record of picking up new hardware targets quickly may substitute for Trainium experience
  • Solid Python fundamentals
  • Comfort operating in a client-facing, production engineering environment

Job type

Full Time

Experience level

Lead

Salary

Not specified

Degree requirement

No Education Requirement

Tech skills

AWSPythonPyTorch

Location requirements

RemoteCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.