Senior Site Reliability Engineer at Rootly embedding with teams to enhance service performance and reliability. Own CI/CD pipelines and drive capacity planning efforts in a fast-paced environment.
Responsibilities
Embed with product teams to enhance observability, reliability, and performance of their services.
Own our CI/CD pipelines, observability tooling, monitoring systems, and incident response processes.
Build tools and automation to eliminate manual toil, improve engineering velocity and developer experience, and improve system reliability.
Collaborate deeply across engineering to understand systems at the code level and surface cross-cutting reliability, performance, and scaling concerns.
Architect and scale our infrastructure, ensuring best-in-class performance, availability, and operational excellence.
Drive capacity planning efforts to ensure our infrastructure is resilient and scalable as we grow.
Define and manage SLOs and error budgets in partnership with Engineering teams who own production services.
Be vocal - act as a strong voice and force of reliability, quality, performance, and scalability.
Requirements
5+ years of experience in an SRE, Platform, or Infrastructure Engineering role.
5+ years of experience writing software in a production environment.
Strong technical knowledge of cloud infrastructure, distributed systems, and reliability practices.
Strong understanding of observability, performance tuning, and scaling strategies.
Deep familiarity with incident response, monitoring, and CI/CD systems.
Hands-on experience supporting web or RPC services at meaningful scale.
You write code to solve infrastructure problems; not shell scripts alone, but production-grade software.
Benefits
Competitive compensation and early equity in a fast-growing, venture-backed company.
Comprehensive medical, dental, and vision coverage.
3 weeks of vacation, plus unlimited sick and mental health days, and a company-wide end-of-year shutdown to recharge.
$500 stipend for home office setup.
A fast-moving, high-impact environment where your leadership and ideas directly shape the future of the company.
DevOps Manager overseeing releases, enterprise tooling, and incident response for Delta Controls, a building - automation solutions manufacturer. Establishing standards across global product teams and offices.
DevOps Engineer building AWS infrastructure and automated systems for S&P Global’s financial data and technology solutions. Supporting resilient applications through Terraform, CI/CD, containerization, monitoring, and cloud operations.
Staff SRE leading GCP reliability, observability, and infrastructure automation for Calix’s broadband communications platform. Building resilient GKE, Kafka, database, and networking systems.
Senior Azure DevOps advisor governing platform evolution for Alithya, a digital transformation consulting firm. Defining standards, optimizing pipelines, dashboards, integrations, and AI capabilities.
Senior DevOps Engineer building Azure DevOps pipelines, Terraform infrastructure, and deployment automation. Supporting PLATO, a Canadian Indigenous - owned software testing and technology services company, across product and data teams.
Pilote Azure DevOps pour un organisme public de santé québécois en transformation numérique. Gouvernance, intégrations Power BI/Dynamics 365, administration de plateforme et accompagnement Agile.