Staff SRE leading GCP reliability, observability, and infrastructure automation for Calix’s broadband communications platform. Building resilient GKE, Kafka, database, and networking systems.
Responsibilities
Lead global platform reliability and drive the next-generation observability strategy on Google Cloud Platform
Architect, optimize, and troubleshoot networking infrastructure across OSI Layers 1–7
Design, scale, and optimize the Grafana Labs observability stack, including Grafana, Mimir, Loki, Tempo, and Beyla
Deploy machine-learning models and automated anomaly detection to reduce alert fatigue and predict bottlenecks
Architect, scale, secure, and manage production Google Kubernetes Engine clusters
Tune and maintain high-throughput Apache Kafka clusters for low-latency event delivery and high availability
Ensure performance, scalability, and disaster-recovery readiness across PostgreSQL, AlloyDB, and BigQuery
Automate incident triage, root-cause analysis, and remediation through Grafana workflows and AIOps insights
Champion the technical roadmap for distributed infrastructure engineering and GCP cloud-native observability standards
Coach engineers on advanced debugging, distributed-systems thinking, and intelligent operations
Requirements
Proven track record of high autonomy and successful delivery in a 100% remote engineering environment
8+ years of experience in SRE, Production Engineering, or Distributed Systems infrastructure roles
Deep technical knowledge across OSI Layers 1–7
Physical/fiber infrastructure awareness, switching, BGP, and OSPF
TCP congestion control, UDP, and QUIC
Session management, TLS termination, DNS architecture, HTTP/3, and gRPC
Expert-level mastery of GKE internals, custom controllers, multi-cluster networking, and GitOps workflows
Experience managing high-throughput Apache Kafka pipelines and large-scale PostgreSQL, AlloyDB, and BigQuery environments
Hands-on experience with Grafana Enterprise/Cloud, Prometheus/Mimir, Loki, and Tempo at scale
Experience applying AI/ML to time-series anomaly detection, log clustering, and correlation
Advanced production-scale expertise with HashiCorp Terraform for multi-region GCP architectures
High proficiency in Go and Python
Exceptional written and verbal communication skills for asynchronous alignment
Deep knowledge of Google Cloud architecture, Cloud SDN, Cloud Armor, Interconnect, IAM, and cost optimization
Understanding of Linux internals, eBPF-based monitoring, kernel-level networking, Wireshark, and tcpdump
Benefits
Eligibility for a bonus as part of the total compensation package
Benefits package available; details provided via the employer's benefits information
DevOps Intern supporting CI/CD, cloud infrastructure, and automation for Ludia’s mobile game studio. Improving reliability and developer tools in production game environments.
Manager leading global SRE teams for Akamai's distributed Cloud IAM services. Improving reliability, scalability, security, and usability through cloud - native tooling and software.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.