Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Site Reliability Engineer improving AXON Networks’ AI-driven ISP orchestration platform and high-speed router services. Enhancing cloud-to-device reliability, observability, automation and incident response.

Responsibilities

  • Improve the availability, performance, scalability and recoverability of AXON Networks cloud solutions
  • Combine software engineering with hands-on NOC operations to make the cloud-to-device service path observable, supportable and resilient at fleet scale
  • Establish practical SRE capabilities inside the NOC while partnering with Support, Operations, cloud and DevOps Engineering
  • Participate in a sustainable on-call rotation and improve diagnosis of customer-impacting issues
  • Own reliability outcomes for assigned cloud services
  • Improve observability, capacity, resilience and recovery
  • Define and operationalize SLIs, SLOs and actionable alerting
  • Automate repetitive NOC work and create safe, testable mechanisms for diagnosis, recovery, device operations and routine production changes
  • Lead technically during incidents, drive evidence-based learning and complete corrective actions
  • Establish reliability baselines, SLIs, SLOs and error budgets for cloud services and device-management workflows
  • Trace failures across cloud APIs, microservices, Kubernetes, infrastructure, databases, messaging, networks, device-management protocols and devices
  • Identify fleet-wide and customer-specific failure patterns
  • Contribute operability requirements and production evidence during design and readiness reviews
  • Maintain NOC dashboards for service health, device reachability, provisioning, commands, telemetry, firmware adoption and customer impact
  • Participate in NOC production on-call rotation as technical incident lead or senior troubleshooter
  • Diagnose complex failures across applications, cloud infrastructure, Kubernetes, APIs, networking, DNS/TLS, databases, messaging, device-management sessions and CPE behavior
  • Coordinate evidence gathering and technical escalation with customers, Engineering, firmware, DevOps and vendors
  • Lead or contribute to post-incident reviews and convert recurring failures into measurable corrective actions
  • Develop production-grade software, scripts and workflows for diagnosis, remediation, deployment safety, fleet analysis, scaling, maintenance and recovery
  • Improve CI/CD and GitOps practices, including automated testing, release validation, progressive delivery and rollback readiness
  • Manage or contribute to infrastructure as code, configuration as code and reusable self-service patterns
  • Measure NOC toil and prioritize durable platform capabilities with Automation & Tools Engineers
  • Develop capacity models for service-provider growth, managed-device populations, telemetry, messaging, APIs and rollout events
  • Create and maintain runbooks, troubleshooting decision trees, service maps, dependency records, known-error guidance and operational knowledge
  • Coach NOC and Support personnel on diagnosis, mitigation, evidence capture and escalation
  • Build self-service diagnostic views and tools for determining scope, affected customers, device cohorts, fault domain and next action
  • Share reliability insights with Engineering and Product and contribute to reliability and operational-readiness reviews

Requirements

  • 5+ years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering or a closely related role
  • Strong software or automation skills in Python, Go, Java, Bash or a comparable language, with experience producing maintainable operational code
  • Hands-on experience operating distributed production systems in a public cloud environment and troubleshooting application, infrastructure, network and device-integration layers
  • Experience with Google Cloud Platform, Oracle Cloud Infrastructure and production Kubernetes environments
  • Experience with Terraform, Helm, Git-based CI/CD and policy-as-code
  • Strong Linux, containers and Kubernetes fundamentals, including deployment behavior, resource management, networking and failure diagnosis
  • Strong troubleshooting and debugging skills in Kubernetes platforms
  • Experience with observability practices and tools across metrics, logs, traces, alerting, dashboards and synthetic monitoring
  • Familiarity with Prometheus, Grafana, OpenTelemetry or equivalent observability ecosystems
  • Familiarity with Apache Pulsar or similar distributed messaging and streaming platforms handling requests from millions of devices
  • Experience participating in an on-call rotation and responding to high-severity, customer-impacting production incidents
  • Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety and blameless incident learning
  • Strong networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing and packet- or session-level troubleshooting
  • Clear communication, disciplined documentation and collaboration across NOC, cloud, DevOps, firmware and service-provider teams
  • Bachelor’s degree in computer science, engineering or equivalent practical experience
  • Preferred: experience supporting cloud-managed CPEs in a service-provider environment
  • Preferred: familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry and remote lifecycle management
  • Preferred: experience with Apache Pulsar or Kafka, APIs and highly available databases used in device-management control planes
  • Preferred: understanding of GPON/XGS-PON, DOCSIS, Ethernet or fixed wireless access technologies
  • Preferred: experience with firmware rollout automation, canary or cohort deployments, fleet health analysis and safe rollback practices
  • Preferred: experience building auto-remediation, safe self-service operations or internal reliability platforms
  • Preferred: experience supporting multiple service-provider customers in a 24×7 telecommunications, broadband or managed-network environment

Benefits

  • Equal opportunity recruitment process
  • Inclusive and diverse working environment

Job type

Contract

Experience level

Mid levelSenior

Salary

CA$80 - CA$110 per hour

Degree requirement

Bachelor's Degree

Tech skills

ApacheCloudDNSGoogle Cloud PlatformGrafanaJavaKafkaKubernetesLinuxMicroservicesOraclePrometheusPulsarPythonTCP/IPTerraformGo

Location requirements

RemoteCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.