Senior SRE contractor needed for 6-12 month remote role in Canada. Requires 8+ years experience with Dynatrace, ELK, Splunk, PagerDuty, AKS, Terraform, and incident management.
Responsibilities
We are hiring a Senior Site Reliability Engineering (SRE) contractor for a 6-12 month remote engagement in Canada. The role requires deep expertise in observability, SRE, and DevOps practices, with hands-on experience in Dynatrace, ELK, Splunk, PagerDuty, Azure Kubernetes Service (AKS), Terraform, and Azure managed services. You will support Node.js and .NET applications in microservices and event-driven architectures, perform distributed tracing, metrics collection, and log aggregation, and lead incident management and root cause analysis across distributed systems.
Requirements
8+ years of experience. Strong experience in Observability, SRE, and DevOps practices. Deep expertise with Dynatrace, ELK, Splunk, and PagerDuty. Strong understanding of observability principles, instrumentation, correlation IDs, and SLI/SLO frameworks. Hands-on experience with Azure Kubernetes Service (AKS). Advanced proficiency with Terraform and Infrastructure as Code (IaC). Experience with Azure managed services including SQL MI, Redis, Functions, and Event Grid. Strong experience with distributed tracing, metrics collection, and log aggregation. Experience supporting Node.js and .NET applications in microservices and event-driven architectures. Strong troubleshooting and root cause analysis skills across distributed systems, APIs, databases, and caches. Experience with incident management tools such as PagerDuty and ServiceNow. Knowledge of incident, problem, and change management processes. Familiarity with CI/CD pipelines, automation, and operational resilience practices. Experience with chaos engineering and blameless postmortems. Strong communication, leadership, and cross-functional collaboration skills. Only candidates holding Canadian PR / OWP / Citizenship will be considered.
DevOps Intern supporting CI/CD, cloud infrastructure, and automation for Ludia’s mobile game studio. Improving reliability and developer tools in production game environments.
Manager leading global SRE teams for Akamai's distributed Cloud IAM services. Improving reliability, scalability, security, and usability through cloud - native tooling and software.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.