Staff DevOps Engineer

Posted 2 days ago

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Staff DevOps Engineer owning reliable, scalable infrastructure for Nexxa’s AI systems. Supporting machine learning workloads across heavy-industry operations.

Responsibilities

  • Own and evolve Nexxa's core infrastructure, including compute, networking, storage, and deployment systems, end-to-end
  • Design and operate CI/CD pipelines supporting safe iteration across AI, data, and product engineering teams
  • Build and maintain infrastructure-as-code for reproducible, auditable cloud and on-prem/edge environments
  • Architect and manage Kubernetes platforms for training, inference, and application workloads, including GPU scheduling and autoscaling
  • Support data warehouses, lakehouse architectures, feature stores, embedding indices, retrieval pipelines, and model training, evaluation, and serving infrastructure
  • Define and drive observability practices covering metrics, logging, tracing, and alerting across distributed systems
  • Establish reliability practices including SLOs/SLIs, incident response, postmortems, and on-call rotations
  • Design security and compliance across cloud infrastructure, secrets management, and access control
  • Make pragmatic tradeoffs among cost, latency, reliability, and developer velocity
  • Collaborate with engineering leadership on infrastructure roadmap and platform strategy
  • Mentor engineers on infrastructure best practices and promote operational excellence

Requirements

  • 6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles
  • Deep hands-on experience with AWS, GCP, or Azure at production scale
  • Kubernetes production experience, including GPU workload scheduling
  • Infrastructure-as-code experience with Terraform, Pulumi, or equivalent
  • CI/CD systems experience, including GitHub Actions, GitLab CI, CircleCI, Jenkins, or ArgoCD
  • Experience designing and operating observability stacks such as Prometheus, Grafana, Datadog, or OpenTelemetry
  • Experience supporting ML/AI infrastructure is a strong plus
  • Excellent scripting/programming skills in Python, Go, or Bash
  • Proven ability to independently scope and lead infrastructure projects from design through production rollout
  • Strong incident management instincts and ability to lead through outages and drive root-cause analysis
  • Experience with cloud and edge/on-prem infrastructure, especially industrial or manufacturing contexts
  • Familiarity with Snowflake, BigQuery, Redshift, or Databricks
  • Experience with service mesh, zero-trust networking, or industrial/critical-infrastructure compliance frameworks such as SOC 2 or IEC 62443
  • History of building internal developer platforms or self-service infrastructure tooling
  • Experience scaling infrastructure teams or setting technical direction at Staff level

Benefits

  • Significant opportunities for career development and advancement
  • Comprehensive salary and equity package
  • Innovative environment focused on transforming heavy industries through AI and automation
  • Collaborative culture valuing innovation, discipline, and continuous improvement

Job type

Full Time

Experience level

Lead

Salary

Not specified

Degree requirement

No Education Requirement

Tech skills

Amazon RedshiftAWSAzureBigQueryCloudDistributed SystemsGoogle Cloud PlatformGrafanaJenkinsKubernetesPrometheusPythonTerraformGo

Location requirements

RemoteCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.