Senior Site Reliability Engineer managing reliability for a large scale SaaS platform. Leading incident response and improving infrastructure resilience with a global engineering team.
Responsibilities
Take point on major incidents, coordinating the response across engineering teams and making the real-time calls needed to restore service quickly.
Hold ownership over key pieces of the shared infrastructure, working continuously to improve its resilience and ability to scale.
Run structured reliability work such as capacity forecasting, performance tuning, and controlled failure testing.
Improve how the org detects problems before customers do, by refining metrics, alerting, and overall observability.
Run thorough retrospectives after incidents and make sure the resulting action items actually get built, not just logged.
Support less experienced SREs through code and design review, pairing, and ongoing feedback, and help spread strong reliability practices across Product and Engineering more broadly.
Requirements
4 to 8 years working in SRE, DevOps, or production engineering roles within SaaS environments.
Solid, practical grounding in Linux, containerized systems, and cloud native infrastructure, including hands-on experience with a major cloud provider and its networking layer.
Comfort writing code or scripts to manage infrastructure, with a genuine preference for infrastructure as code over manual configuration.
Real experience building and running CI/CD pipelines, observability tooling, and version control workflows at scale.
A track record of staying level headed during live incidents, and the ability to explain technical issues clearly to people who aren't engineers.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.