Staff Site Reliability Operations Engineer

Posted 3 days ago

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Staff SRE leading GCP reliability, observability, and infrastructure automation for Calix’s broadband communications platform. Building resilient GKE, Kafka, database, and networking systems.

Responsibilities

  • Lead global platform reliability and drive the next-generation observability strategy on Google Cloud Platform
  • Architect, optimize, and troubleshoot networking infrastructure across OSI Layers 1–7
  • Design, scale, and optimize the Grafana Labs observability stack, including Grafana, Mimir, Loki, Tempo, and Beyla
  • Deploy machine-learning models and automated anomaly detection to reduce alert fatigue and predict bottlenecks
  • Architect, scale, secure, and manage production Google Kubernetes Engine clusters
  • Tune and maintain high-throughput Apache Kafka clusters for low-latency event delivery and high availability
  • Ensure performance, scalability, and disaster-recovery readiness across PostgreSQL, AlloyDB, and BigQuery
  • Automate incident triage, root-cause analysis, and remediation through Grafana workflows and AIOps insights
  • Champion the technical roadmap for distributed infrastructure engineering and GCP cloud-native observability standards
  • Coach engineers on advanced debugging, distributed-systems thinking, and intelligent operations

Requirements

  • Proven track record of high autonomy and successful delivery in a 100% remote engineering environment
  • 8+ years of experience in SRE, Production Engineering, or Distributed Systems infrastructure roles
  • Deep technical knowledge across OSI Layers 1–7
  • Physical/fiber infrastructure awareness, switching, BGP, and OSPF
  • TCP congestion control, UDP, and QUIC
  • Session management, TLS termination, DNS architecture, HTTP/3, and gRPC
  • Expert-level mastery of GKE internals, custom controllers, multi-cluster networking, and GitOps workflows
  • Experience managing high-throughput Apache Kafka pipelines and large-scale PostgreSQL, AlloyDB, and BigQuery environments
  • Hands-on experience with Grafana Enterprise/Cloud, Prometheus/Mimir, Loki, and Tempo at scale
  • Experience applying AI/ML to time-series anomaly detection, log clustering, and correlation
  • Advanced production-scale expertise with HashiCorp Terraform for multi-region GCP architectures
  • High proficiency in Go and Python
  • Exceptional written and verbal communication skills for asynchronous alignment
  • Deep knowledge of Google Cloud architecture, Cloud SDN, Cloud Armor, Interconnect, IAM, and cost optimization
  • Understanding of Linux internals, eBPF-based monitoring, kernel-level networking, Wireshark, and tcpdump

Benefits

  • Eligibility for a bonus as part of the total compensation package
  • Benefits package available; details provided via the employer's benefits information

Job type

Full Time

Experience level

Lead

Salary

$136,000 - $231,000 per year

Degree requirement

No Education Requirement

Tech skills

ApacheBigQueryCloudDistributed SystemsDNSGoogle Cloud PlatformGrafanaGRPCKafkaKubernetesLinuxPostgresPrometheusPythonSwitchingTerraformGo

Location requirements

RemoteUnited States

Report this job

Found something wrong with the page? Please let us know by submitting a report below.