Site Reliability Engineering Team Lead – Principal SRE

Posted 6 days ago

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Principal SRE leading Cerence’s Site Reliability Engineering function for cloud-native automotive AI systems. Owning reliability strategy, incident escalation, observability, automation, and SLO governance.

Responsibilities

  • Lead Cerence's Site Reliability Engineering team and own the reliability, availability, and operational health of its cloud-native automotive AI platform
  • Help select, mentor, and technically develop the team across multiple locations
  • Set technical direction and priorities and contribute performance and growth input to team managers
  • Design and maintain a sustainable on-call rotation and monitor page load and team health
  • Own and drive the reliability roadmap across a 2–3 quarter horizon
  • Define and govern SLI/SLO/SLA frameworks for customer program availability targets up to 99.95%
  • Serve as Tier 2 technical escalation point for major incidents in partnership with the Global Operations Center
  • Champion blameless postmortem culture and ensure actionable outcomes
  • Lead and improve Production Readiness / NFR reviews with development teams
  • Contribute to root cause analysis and own systemic improvements
  • Approve high-risk and out-of-window production changes
  • Set strategic direction for metrics, dashboards, alerting, escalation, and automation
  • Drive CI/CD automation pipelines for service deployments, rollbacks, and operational tasks
  • Partner with DevOps and platform teams to evolve shared infrastructure
  • Embed reliability into the SDLC through collaboration with development managers and architects
  • Participate in service reliability consulting and architectural reviews
  • Communicate reliability posture and risk to technical and non-technical stakeholders

Requirements

  • 8+ years of hands-on experience in site reliability, DevOps, or cloud platform roles, including time leading a team or owning a function
  • A track record of setting technical direction and holding standards across a team — with or without formal authority
  • Hands-on experience with container orchestration frameworks (Kubernetes, Docker, Istio)
  • Experience with public cloud platforms (Azure primarily; AWS and Google Cloud)
  • Familiarity with observability tooling — metrics pipelines, dashboarding, and alerting (e.g., Zabbix, Prometheus, Grafana)
  • Experience with CI/CD pipelines and infrastructure-as-code practices (e.g., Terraform, Flux)
  • Proficiency in at least one scripting or programming language (Python, Go, Shell, etc.)
  • Strong UNIX/Linux background, including system configuration, performance debugging, and network fundamentals (Layer 4/5, DNS, HTTP/S, TLS)
  • Excellent written and verbal communication skills in English
  • Previous site reliability leadership experience
  • Experience leading distributed or multi-site technical teams
  • Background in high-availability service design (redundancy, failover, blast radius)
  • Experience with log aggregation and analytics platforms (Loki, Thanos)
  • Familiarity with ITSM and project tooling (Jira, Confluence)
  • Experience in automotive, embedded, or latency-sensitive production environments

Benefits

  • Annual bonus opportunity
  • Insurance coverage (medical, dental, vision, life, and disability)
  • Paid time off
  • Paid holidays
  • Company contribution to the RRSP (Registered Retirement Savings Plan)
  • Equity awards for certain positions and levels
  • Remote and/or hybrid work available depending on the position

Job type

Full Time

Experience level

Senior

Salary

Not specified

Degree requirement

No Education Requirement

Tech skills

AWSAzureCloudDNSDockerFluxGrafanaITSMKubernetesLinuxPrometheusPythonSDLCTerraformUnixGo

Location requirements

RemoteCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.