Site Reliability Engineer focusing on maintaining infrastructure and automating processes. Collaborating with a team while reporting to senior engineers in a hybrid setting.
Responsibilities
Maintain and enhance existing infrastructure components, scripts, and automation — with a focus on writing clean, maintainable code
Support CI/CD pipelines by contributing to configuration, troubleshooting failures, and implementing improvements under the direction of senior engineers
Write and maintain Infrastructure as Code (IaC) using Terraform, following established team standards and conventions
Contribute to observability solutions by adding or updating monitoring dashboards, alerts, and log queries (Prometheus, Coralogix, Datadog, or similar)
Participate in code reviews, learning from feedback and applying coding best practices to infrastructure-related services
Assist in incident response activities — helping with investigation, documenting findings, and implementing fixes with guidance
Apply programming skills in one or more languages (e.g., Python, Go, or JavaScript) to build or improve internal tooling
Learn and apply SRE principles as you contribute to the team’s ongoing reliability and performance work
Requirements
Bachelor’s degree in Computer Science, Software Engineering, or equivalent practical experience
1+ year of experience in a software engineering, DevOps, or SRE-adjacent role (internships and co-ops count)
Working knowledge of at least one programming language (Python, Go, Java, C#, or similar) with a willingness to learn more
Familiarity with IaC concepts; hands-on experience with Terraform is a plus
Basic understanding of CI/CD pipelines and version control (Git)
Exposure to cloud platforms (AWS or Azure) and containerization concepts (Docker, Kubernetes)
Curiosity about how distributed systems work and how to make them more reliable
Good communication skills — you ask questions when stuck, and document what you learn
Genuine interest in automation and operational excellence
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.