Site Reliability Engineer focused on ensuring reliability and scalability of CloudBlue’s SaaS platforms. Collaborating with global teams to monitor and improve multi-tenant service providers' systems.
Responsibilities
Define and implement SLIs, SLOs, and error budgets for critical CloudBlue services to ensure reliability and performance
Influence system architecture with a strong focus on reliability, scalability, and operability, designing systems for fault tolerance, graceful degradation, and self-healing
Reduce operational toil by identifying opportunities for automation and process improvement
Design and operate CloudBlue’s observability stack across metrics, logs, and traces using tools such as Datadog, Grafana, and Elastic Stack
Develop actionable alerting strategies and dashboards that provide clear insight into platform and business health
Design and maintain high-availability architectures, implementing redundancy, failover, and disaster recovery strategies across regions and availability zones
Conduct capacity planning, load testing, and performance optimization to ensure platform stability and scalability
Act as a senior responder during production incidents, leading incident coordination, communication, and service restoration
Own blameless postmortems and drive improvements that reduce incident frequency, MTTR, and customer impact
Improve reliability of Kubernetes-based platforms through health checks, autoscaling strategies, rollout safety, and resilience testing
Partner with engineering and DevOps teams to improve deployment safety, rollback strategies, and platform reliability
Maintain runbooks and operational documentation, and promote SRE best practices across engineering teams
Support other tasks or projects as assigned to meet team and business needs
Requirements
3+ years of experience as an SRE, DevOps Engineer, or Production Engineer, with strong ownership of production systems
DevOps Intern supporting CI/CD, cloud infrastructure, and automation for Ludia’s mobile game studio. Improving reliability and developer tools in production game environments.
Manager leading global SRE teams for Akamai's distributed Cloud IAM services. Improving reliability, scalability, security, and usability through cloud - native tooling and software.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.