Senior Site Reliability Engineer maintaining and optimizing large-scale distributed infrastructure at Branch. Collaborating with cross-functional teams to support mission-critical services across the organization.
Responsibilities
Architect, design, and evolve complex distributed systems to improve reliability, operational efficiency, and performance at scale
Partner closely with product, security, and data engineering teams to translate business needs into resilient and scalable system designs
Drive reliability through automation and advanced observability
Lead and mentor in high stakes situations
Perform deep infrastructure cost audits
Own and maintain key distributed data platforms
Guide teams in defining SLIs/SLOs and operational best practices
Continuously identify and eliminate bottlenecks
Champion Infrastructure as Code (IaC) to automate provisioning, configuration, and lifecycle management
Lead our GitOps and deployment strategy using Argo CD
Requirements
6+ years in SRE, systems engineering, or software engineering roles
Proven track record as a senior reliability or production engineer
Expert level proficiency in Kubernetes, AWS, Linux internals, and distributed system fundamentals
Strong programming skills in Go, Python, Java, Kotlin, Bash, or similar languages
Hands-on experience with modern observability stacks (Prometheus, Grafana, AlertManager, Loki, PagerDuty)
Familiarity with large scale data and streaming ecosystems such as Kafka, Spark, Aerospike, FoundationDB, and the broader Hadoop ecosystem
Deep experience with Terraform, CloudFormation, or related IaC tooling
Proven incident management leadership in production SaaS systems
DevOps Intern supporting CI/CD, cloud infrastructure, and automation for Ludia’s mobile game studio. Improving reliability and developer tools in production game environments.
Manager leading global SRE teams for Akamai's distributed Cloud IAM services. Improving reliability, scalability, security, and usability through cloud - native tooling and software.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.