Senior Site Reliability Engineer maintaining and optimizing large-scale distributed infrastructure at Branch. Collaborating with cross-functional teams to support mission-critical services across the organization.
Responsibilities
Architect, design, and evolve complex distributed systems to improve reliability, operational efficiency, and performance at scale
Partner closely with product, security, and data engineering teams to translate business needs into resilient and scalable system designs
Drive reliability through automation and advanced observability
Lead and mentor in high stakes situations
Perform deep infrastructure cost audits
Own and maintain key distributed data platforms
Guide teams in defining SLIs/SLOs and operational best practices
Continuously identify and eliminate bottlenecks
Champion Infrastructure as Code (IaC) to automate provisioning, configuration, and lifecycle management
Lead our GitOps and deployment strategy using Argo CD
Requirements
6+ years in SRE, systems engineering, or software engineering roles
Proven track record as a senior reliability or production engineer
Expert level proficiency in Kubernetes, AWS, Linux internals, and distributed system fundamentals
Strong programming skills in Go, Python, Java, Kotlin, Bash, or similar languages
Hands-on experience with modern observability stacks (Prometheus, Grafana, AlertManager, Loki, PagerDuty)
Familiarity with large scale data and streaming ecosystems such as Kafka, Spark, Aerospike, FoundationDB, and the broader Hadoop ecosystem
Deep experience with Terraform, CloudFormation, or related IaC tooling
Proven incident management leadership in production SaaS systems
DevOps Manager overseeing releases, enterprise tooling, and incident response for Delta Controls, a building - automation solutions manufacturer. Establishing standards across global product teams and offices.
DevOps Engineer building AWS infrastructure and automated systems for S&P Global’s financial data and technology solutions. Supporting resilient applications through Terraform, CI/CD, containerization, monitoring, and cloud operations.
Staff SRE leading GCP reliability, observability, and infrastructure automation for Calix’s broadband communications platform. Building resilient GKE, Kafka, database, and networking systems.
Senior Azure DevOps advisor governing platform evolution for Alithya, a digital transformation consulting firm. Defining standards, optimizing pipelines, dashboards, integrations, and AI capabilities.
Senior DevOps Engineer building Azure DevOps pipelines, Terraform infrastructure, and deployment automation. Supporting PLATO, a Canadian Indigenous - owned software testing and technology services company, across product and data teams.