Manager of Site Reliability Engineering at Docebo, overseeing platform health and leading engineering teams. Focus on operational efficiency and incident management for reliable SaaS delivery.
Responsibilities
Lead a team of Site Reliability Engineers to ensure operational health of the Docebo platform
Manage incident responses and day-to-day production operations
Conduct post-incident reviews and implement corrective actions
Collaborate with Product, Engineering, and Support teams for reliable practices
Requirements
6 to 10 years of experience in SRE, DevOps, or systems operations
Managing or coordinating critical incident responses and production operations in SaaS environments
Familiarity with AWS and advanced monitoring/observability tools
Collaboration with cross-functional teams to deliver results
Benefits
Employee Share Purchase Plan (ESPP) at a 15% discount
Senior DevOps Engineer improving Sherweb’s cloud - based IT solution delivery through CI/CD, IaC, GitOps, and AI automation. Supporting secure platforms, operational transitions, and developer self - service.
Senior DevOps / Cloud Infrastructure Engineer needed for hybrid role in North York, ON. Requires 10+ years experience with GCP, AWS, Kubernetes, Terraform, and CI/CD.
Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.
Mozilla release engineer optimizing Firefox build, test, and deployment pipelines at global scale. Improving developer experience, maintaining automation, and responding to critical service outages without an on - call rotation.