Staff SRE leading BeyondTrust’s Password Safe identity-security platform across cloud and on-premises infrastructure. Driving GitOps, CI/CD, observability, resilience, and reliability strategy.
Responsibilities
Lead the evolution of the Password Safe platform, infrastructure, and deployment ecosystem.
Architect highly available and resilient systems across cloud and on-premises environments.
Own the reliability, scalability, and efficiency of shared services, CI/CD pipelines, and platform engineering initiatives.
Design, scale, and maintain secure, resilient systems across AWS/Azure and on-premises environments.
Champion platform engineering initiatives that reduce cognitive load and accelerate software engineering velocity.
Own and optimize API gateways, service meshes, caches, configuration management, and secrets management.
Drive an Everything as Code culture using declarative, version-controlled infrastructure, pipelines, and configurations deployed through GitOps workflows.
Standardize, secure, and optimize CI/CD pipelines for safe, repeatable, rapid deployments.
Design and implement chaos engineering frameworks and disaster recovery simulations.
Architect and mature telemetry using metrics, logs, traces, Grafana Cloud, Datadog, and OpenTelemetry.
Define and enforce SLOs and SLIs across critical applications.
Partner with engineering leadership on long-term SRE strategy.
Mentor and coach engineers, lead design and incident reviews, and build reliability documentation and tooling.
Requirements
7+ years of experience in SRE, DevOps, or Platform Engineering, including at least 2 years at Senior or Staff level
Proven experience managing automation and infrastructure across cloud and on-premises environments
Experience with Docker and Kubernetes, including cluster administration, networking, and security primitives
Prior experience with canary or blue-green release orchestration strategies
Proficiency in at least one systems language such as Go, Java, or C#
Deep understanding of OpenTelemetry and monitoring and observability best practices
Familiarity with GitOps workflows and Infrastructure as Code and Configuration as Code best practices
Nice to have: UI automation testing knowledge
Nice to have: experience designing and testing microservice-based applications
Nice to have: experience with virtual machines and test environments
Nice to have: experience in continuous integration environments
Nice to have: deep understanding of Linux and Windows internals
Nice to have: experience migrating workloads across on-premises and cloud environments
Nice to have: understanding of modern DevSecOps practices
Senior DevOps / Cloud Infrastructure Engineer needed for hybrid role in North York, ON. Requires 10+ years experience with GCP, AWS, Kubernetes, Terraform, and CI/CD.
Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.
Mozilla release engineer optimizing Firefox build, test, and deployment pipelines at global scale. Improving developer experience, maintaining automation, and responding to critical service outages without an on - call rotation.