Staff DevOps Engineer owning reliable, scalable infrastructure for Nexxa’s AI systems. Supporting machine learning workloads across heavy-industry operations.
Responsibilities
Own and evolve Nexxa's core infrastructure, including compute, networking, storage, and deployment systems, end-to-end
Design and operate CI/CD pipelines supporting safe iteration across AI, data, and product engineering teams
Build and maintain infrastructure-as-code for reproducible, auditable cloud and on-prem/edge environments
Architect and manage Kubernetes platforms for training, inference, and application workloads, including GPU scheduling and autoscaling
Support data warehouses, lakehouse architectures, feature stores, embedding indices, retrieval pipelines, and model training, evaluation, and serving infrastructure
Define and drive observability practices covering metrics, logging, tracing, and alerting across distributed systems
Establish reliability practices including SLOs/SLIs, incident response, postmortems, and on-call rotations
Design security and compliance across cloud infrastructure, secrets management, and access control
Make pragmatic tradeoffs among cost, latency, reliability, and developer velocity
Collaborate with engineering leadership on infrastructure roadmap and platform strategy
Mentor engineers on infrastructure best practices and promote operational excellence
Requirements
6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles
Deep hands-on experience with AWS, GCP, or Azure at production scale
Kubernetes production experience, including GPU workload scheduling
Infrastructure-as-code experience with Terraform, Pulumi, or equivalent
CI/CD systems experience, including GitHub Actions, GitLab CI, CircleCI, Jenkins, or ArgoCD
Experience designing and operating observability stacks such as Prometheus, Grafana, Datadog, or OpenTelemetry
Experience supporting ML/AI infrastructure is a strong plus
Excellent scripting/programming skills in Python, Go, or Bash
Proven ability to independently scope and lead infrastructure projects from design through production rollout
Strong incident management instincts and ability to lead through outages and drive root-cause analysis
Experience with cloud and edge/on-prem infrastructure, especially industrial or manufacturing contexts
Familiarity with Snowflake, BigQuery, Redshift, or Databricks
Experience with service mesh, zero-trust networking, or industrial/critical-infrastructure compliance frameworks such as SOC 2 or IEC 62443
History of building internal developer platforms or self-service infrastructure tooling
Experience scaling infrastructure teams or setting technical direction at Staff level
Benefits
Significant opportunities for career development and advancement
Comprehensive salary and equity package
Innovative environment focused on transforming heavy industries through AI and automation
Collaborative culture valuing innovation, discipline, and continuous improvement
Senior DevOps Engineer building secure, scalable Azure platforms for CARET’s legal and accounting practice - management software. Leading infrastructure, Kubernetes, CI/CD, security, observability, and reliability initiatives.
Senior SRE operating Kubernetes and cloud infrastructure for Penn Entertainment’s sports betting and media platforms. Leading migrations, automation, observability, and incident response across regulated production services.
DevOps Engineer operating multi - cloud Kubernetes infrastructure for InfluxData’s time - series platform. Automating operations and supporting highly available distributed services.
Senior Reliability Engineer improving embedded protection, control, and software products for utility grids. Leading reliability testing, failure analysis, and modernization initiatives for resilient energy systems.
Senior Reliability Engineer improving embedded grid automation reliability for utility - scale energy systems. Leading testing, failure analysis, KPIs, and modernization initiatives with utilities.
Site Reliability Expert leading observability and SRE for Valtech, an experience innovation company. Improving reliability across cloud - native, microservices - based environments.
Staff SRE securing and scaling IAM systems at RBC, a Canadian bank. Designing resilient infrastructure, automating operations, and leading incident response.
DevOps Engineer designing production - style CI/CD, cloud, and infrastructure tasks for Your Software Supplier. Reviewing AI - generated solutions and ensuring correctness and reproducibility.