Site Reliability Engineer at BMO focusing on code deployment, IT operations, and system reliability through automation and monitoring. Collaborating between development and operations teams to improve service health.
Responsibilities
Designs how code is deployed, configured, and monitored
Helps teams determine new features by using service-level agreements (SLAs) and service-level objectives (SLO)
Applies software engineering to automate IT operations tasks
Acts as a link between the development and operations teams
Conducts chaos tests and performance tests for critical business requirements
Debugs production issues across services and levels of the technology stack
Computes the cost of SLA breaches and assists management in calculating impact of system reliability
Improves service health visibility by recording metrics, logs, and traces across all services
Requirements
Typically between 4 - 6 years of relevant experience
Foundational level of proficiency: DevOps, Cybersecurity and privacy concepts
Emotional agility, IT infrastructure library, Robot Process Automation, Cloud Computing, Configuration Management, Container Orchestration, System Design and Implementation, Incident management, Learning Agility, Building and managing relationships
Intermediate level of proficiency: API Management, Automation and Automation Pipelines, Automated Testing, Quality Assurance and Control, Verbal & written communication skills, Collaboration & team skills, Analytical and problem solving skills, Data driven decision making
Post-secondary degree in related field of study or equivalent combination of education and experience
DevOps Intern supporting CI/CD, cloud infrastructure, and automation for Ludia’s mobile game studio. Improving reliability and developer tools in production game environments.
Manager leading global SRE teams for Akamai's distributed Cloud IAM services. Improving reliability, scalability, security, and usability through cloud - native tooling and software.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.