Site Reliability Engineer improving AXON Networks’ AI-driven ISP orchestration platform and high-speed router services. Enhancing cloud-to-device reliability, observability, automation and incident response.
Responsibilities
Improve the availability, performance, scalability and recoverability of AXON Networks cloud solutions
Combine software engineering with hands-on NOC operations to make the cloud-to-device service path observable, supportable and resilient at fleet scale
Establish practical SRE capabilities inside the NOC while partnering with Support, Operations, cloud and DevOps Engineering
Participate in a sustainable on-call rotation and improve diagnosis of customer-impacting issues
Own reliability outcomes for assigned cloud services
Improve observability, capacity, resilience and recovery
Define and operationalize SLIs, SLOs and actionable alerting
Automate repetitive NOC work and create safe, testable mechanisms for diagnosis, recovery, device operations and routine production changes
Lead technically during incidents, drive evidence-based learning and complete corrective actions
Establish reliability baselines, SLIs, SLOs and error budgets for cloud services and device-management workflows
Trace failures across cloud APIs, microservices, Kubernetes, infrastructure, databases, messaging, networks, device-management protocols and devices
Identify fleet-wide and customer-specific failure patterns
Contribute operability requirements and production evidence during design and readiness reviews
Maintain NOC dashboards for service health, device reachability, provisioning, commands, telemetry, firmware adoption and customer impact
Participate in NOC production on-call rotation as technical incident lead or senior troubleshooter
Diagnose complex failures across applications, cloud infrastructure, Kubernetes, APIs, networking, DNS/TLS, databases, messaging, device-management sessions and CPE behavior
Coordinate evidence gathering and technical escalation with customers, Engineering, firmware, DevOps and vendors
Lead or contribute to post-incident reviews and convert recurring failures into measurable corrective actions
Develop production-grade software, scripts and workflows for diagnosis, remediation, deployment safety, fleet analysis, scaling, maintenance and recovery
Improve CI/CD and GitOps practices, including automated testing, release validation, progressive delivery and rollback readiness
Manage or contribute to infrastructure as code, configuration as code and reusable self-service patterns
Measure NOC toil and prioritize durable platform capabilities with Automation & Tools Engineers
Develop capacity models for service-provider growth, managed-device populations, telemetry, messaging, APIs and rollout events
Create and maintain runbooks, troubleshooting decision trees, service maps, dependency records, known-error guidance and operational knowledge
Coach NOC and Support personnel on diagnosis, mitigation, evidence capture and escalation
Build self-service diagnostic views and tools for determining scope, affected customers, device cohorts, fault domain and next action
Share reliability insights with Engineering and Product and contribute to reliability and operational-readiness reviews
Requirements
5+ years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering or a closely related role
Strong software or automation skills in Python, Go, Java, Bash or a comparable language, with experience producing maintainable operational code
Hands-on experience operating distributed production systems in a public cloud environment and troubleshooting application, infrastructure, network and device-integration layers
Experience with Google Cloud Platform, Oracle Cloud Infrastructure and production Kubernetes environments
Experience with Terraform, Helm, Git-based CI/CD and policy-as-code
Strong Linux, containers and Kubernetes fundamentals, including deployment behavior, resource management, networking and failure diagnosis
Strong troubleshooting and debugging skills in Kubernetes platforms
Experience with observability practices and tools across metrics, logs, traces, alerting, dashboards and synthetic monitoring
Familiarity with Prometheus, Grafana, OpenTelemetry or equivalent observability ecosystems
Familiarity with Apache Pulsar or similar distributed messaging and streaming platforms handling requests from millions of devices
Experience participating in an on-call rotation and responding to high-severity, customer-impacting production incidents
Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety and blameless incident learning
Strong networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing and packet- or session-level troubleshooting
Clear communication, disciplined documentation and collaboration across NOC, cloud, DevOps, firmware and service-provider teams
Bachelor’s degree in computer science, engineering or equivalent practical experience
Preferred: experience supporting cloud-managed CPEs in a service-provider environment
Preferred: familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry and remote lifecycle management
Preferred: experience with Apache Pulsar or Kafka, APIs and highly available databases used in device-management control planes
Preferred: understanding of GPON/XGS-PON, DOCSIS, Ethernet or fixed wireless access technologies
Preferred: experience with firmware rollout automation, canary or cohort deployments, fleet health analysis and safe rollback practices
Preferred: experience building auto-remediation, safe self-service operations or internal reliability platforms
Preferred: experience supporting multiple service-provider customers in a 24×7 telecommunications, broadband or managed-network environment
We’re looking for an Azure & Databricks DevOps Engineer (12 - month renewable contract) to support a large - scale Azure data platform initiative. 📍 Hybrid role in
Senior DevOps Engineer building secure, scalable Azure platforms for CARET’s legal and accounting practice - management software. Leading infrastructure, Kubernetes, CI/CD, security, observability, and reliability initiatives.
Senior SRE operating Kubernetes and cloud infrastructure for Penn Entertainment’s sports betting and media platforms. Leading migrations, automation, observability, and incident response across regulated production services.
DevOps Engineer operating multi - cloud Kubernetes infrastructure for InfluxData’s time - series platform. Automating operations and supporting highly available distributed services.
Senior Reliability Engineer improving embedded protection, control, and software products for utility grids. Leading reliability testing, failure analysis, and modernization initiatives for resilient energy systems.
Senior Reliability Engineer improving embedded grid automation reliability for utility - scale energy systems. Leading testing, failure analysis, KPIs, and modernization initiatives with utilities.