Staff SRE leading GCP reliability, observability, and infrastructure automation for Calix’s broadband communications platform. Building resilient GKE, Kafka, database, and networking systems.
Responsibilities
Lead global platform reliability and drive the next-generation observability strategy on Google Cloud Platform
Architect, optimize, and troubleshoot networking infrastructure across OSI Layers 1–7
Design, scale, and optimize the Grafana Labs observability stack, including Grafana, Mimir, Loki, Tempo, and Beyla
Deploy machine-learning models and automated anomaly detection to reduce alert fatigue and predict bottlenecks
Architect, scale, secure, and manage production Google Kubernetes Engine clusters
Tune and maintain high-throughput Apache Kafka clusters for low-latency event delivery and high availability
Ensure performance, scalability, and disaster-recovery readiness across PostgreSQL, AlloyDB, and BigQuery
Automate incident triage, root-cause analysis, and remediation through Grafana workflows and AIOps insights
Champion the technical roadmap for distributed infrastructure engineering and GCP cloud-native observability standards
Coach engineers on advanced debugging, distributed-systems thinking, and intelligent operations
Requirements
Proven track record of high autonomy and successful delivery in a 100% remote engineering environment
8+ years of experience in SRE, Production Engineering, or Distributed Systems infrastructure roles
Deep technical knowledge across OSI Layers 1–7
Physical/fiber infrastructure awareness, switching, BGP, and OSPF
TCP congestion control, UDP, and QUIC
Session management, TLS termination, DNS architecture, HTTP/3, and gRPC
Expert-level mastery of GKE internals, custom controllers, multi-cluster networking, and GitOps workflows
Experience managing high-throughput Apache Kafka pipelines and large-scale PostgreSQL, AlloyDB, and BigQuery environments
Hands-on experience with Grafana Enterprise/Cloud, Prometheus/Mimir, Loki, and Tempo at scale
Experience applying AI/ML to time-series anomaly detection, log clustering, and correlation
Advanced production-scale expertise with HashiCorp Terraform for multi-region GCP architectures
High proficiency in Go and Python
Exceptional written and verbal communication skills for asynchronous alignment
Deep knowledge of Google Cloud architecture, Cloud SDN, Cloud Armor, Interconnect, IAM, and cost optimization
Understanding of Linux internals, eBPF-based monitoring, kernel-level networking, Wireshark, and tcpdump
Benefits
Eligibility for a bonus as part of the total compensation package
Benefits package available; details provided via the employer's benefits information
DevOps Engineer building AWS infrastructure and automated systems for S&P Global’s financial data and technology solutions. Supporting resilient applications through Terraform, CI/CD, containerization, monitoring, and cloud operations.
Senior Azure DevOps advisor governing platform evolution for Alithya, a digital transformation consulting firm. Defining standards, optimizing pipelines, dashboards, integrations, and AI capabilities.
Senior DevOps Engineer building Azure DevOps pipelines, Terraform infrastructure, and deployment automation. Supporting PLATO, a Canadian Indigenous - owned software testing and technology services company, across product and data teams.
Pilote Azure DevOps pour un organisme public de santé québécois en transformation numérique. Gouvernance, intégrations Power BI/Dynamics 365, administration de plateforme et accompagnement Agile.
AWS Cloud Engineer optimizing Amazon Connect and enterprise cloud infrastructure for Miratech, a global IT services and consulting company. Automating operations, reliability, observability, and deployments across AWS platforms.
Senior SRE strengthening PostgreSQL, Kubernetes, and observability for Alpaca’s global brokerage infrastructure. Operating production systems, improving reliability, and mentoring engineers across cloud and database operations.