Principal SRE leading Cerence’s Site Reliability Engineering function for cloud-native automotive AI systems. Owning reliability strategy, incident escalation, observability, automation, and SLO governance.
Responsibilities
Lead Cerence's Site Reliability Engineering team and own the reliability, availability, and operational health of its cloud-native automotive AI platform
Help select, mentor, and technically develop the team across multiple locations
Set technical direction and priorities and contribute performance and growth input to team managers
Design and maintain a sustainable on-call rotation and monitor page load and team health
Own and drive the reliability roadmap across a 2–3 quarter horizon
Define and govern SLI/SLO/SLA frameworks for customer program availability targets up to 99.95%
Serve as Tier 2 technical escalation point for major incidents in partnership with the Global Operations Center
Champion blameless postmortem culture and ensure actionable outcomes
Lead and improve Production Readiness / NFR reviews with development teams
Contribute to root cause analysis and own systemic improvements
Approve high-risk and out-of-window production changes
Set strategic direction for metrics, dashboards, alerting, escalation, and automation
Drive CI/CD automation pipelines for service deployments, rollbacks, and operational tasks
Partner with DevOps and platform teams to evolve shared infrastructure
Embed reliability into the SDLC through collaboration with development managers and architects
Participate in service reliability consulting and architectural reviews
Communicate reliability posture and risk to technical and non-technical stakeholders
Requirements
8+ years of hands-on experience in site reliability, DevOps, or cloud platform roles, including time leading a team or owning a function
A track record of setting technical direction and holding standards across a team — with or without formal authority
Hands-on experience with container orchestration frameworks (Kubernetes, Docker, Istio)
Experience with public cloud platforms (Azure primarily; AWS and Google Cloud)
Familiarity with observability tooling — metrics pipelines, dashboarding, and alerting (e.g., Zabbix, Prometheus, Grafana)
Experience with CI/CD pipelines and infrastructure-as-code practices (e.g., Terraform, Flux)
Proficiency in at least one scripting or programming language (Python, Go, Shell, etc.)
Strong UNIX/Linux background, including system configuration, performance debugging, and network fundamentals (Layer 4/5, DNS, HTTP/S, TLS)
Excellent written and verbal communication skills in English
Previous site reliability leadership experience
Experience leading distributed or multi-site technical teams
Background in high-availability service design (redundancy, failover, blast radius)
Experience with log aggregation and analytics platforms (Loki, Thanos)
Familiarity with ITSM and project tooling (Jira, Confluence)
Experience in automotive, embedded, or latency-sensitive production environments
Benefits
Annual bonus opportunity
Insurance coverage (medical, dental, vision, life, and disability)
Paid time off
Paid holidays
Company contribution to the RRSP (Registered Retirement Savings Plan)
Equity awards for certain positions and levels
Remote and/or hybrid work available depending on the position
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.
Senior Site Reliability Engineer securing AuthZed’s cloud infrastructure and authorization platform, including SpiceDB. Building Kubernetes guardrails, supply - chain security, vulnerability management, and incident response.
Senior Data Scientist developing anomaly detection, dashboards, and alerts for General Motors vehicle reliability. Supporting engineering, quality, warranty, and software teams with production analytics.
DevOps Specialist scaling AWS infrastructure and automating deployments for Portage CyberTech’s digital trust, identity, privacy, and security solutions. Maintaining reliable cloud environments and CI/CD pipelines.
SRE Manager leading platform engineering teams across reliability and developer experience. Improving Lightspeed’s global cloud commerce SaaS platform through scalable infrastructure and automation.
SRE Manager leading platform engineering, reliability, and Developer Experience for Lightspeed’s global cloud POS SaaS platform. Guiding teams, infrastructure automation, and strategic delivery.
Gestionnaire SRE dirigeant deux équipes de plateforme chez Lightspeed. Assurant la fiabilité et l’automatisation de sa plateforme SaaS mondiale de commerce infonuagique.
DevOps Engineer managing AWS, Kubernetes, CI/CD, and observability for StellarTech’s scalable global EdTech products. Improving reliability, security, and production operations across multiple products.