Staff SRE securing and scaling IAM systems at RBC, a Canadian bank. Designing resilient infrastructure, automating operations, and leading incident response.
Responsibilities
Serve as the senior-most technical voice for IAM reliability, setting architecture direction and reliability standards
Own end-to-end service reliability for IAM systems by defining and maintaining SLOs, SLIs, and error budgets
Design and implement resilient, highly available IAM infrastructure and services across multi-region and hybrid-cloud architectures
Write and review production-grade services, APIs, automation frameworks, and internal tools
Build self-service platforms, reusable modules, and golden-path automation for IAM services
Champion Infrastructure as Code and GitOps practices using Terraform, Ansible, Puppet, Kubernetes, and Helm
Own and improve CI/CD pipelines and release engineering for IAM services
Lead high-severity incident response, root cause analysis, blameless postmortems, and structural remediation
Build and evolve observability, monitoring, and alerting pipelines using metrics, logs, and traces
Develop and test failover strategies, backup validation, chaos engineering exercises, and disaster recovery simulations
Orchestrate workload automation, scheduling, and release pipelines across enterprise systems
Partner with security, infrastructure, application, and compliance teams on business continuity and resilience strategy
Mentor and coach engineers on reliability, automation-first thinking, and software engineering practices
Requirements
5+ years of experience in Site Reliability Engineering, DevOps or Platform Engineering with demonstrated staff/senior-level technical leadership across SLOs/SLIs, error budgets, and driving continuous reliability improvement at scale
Proficient in at least one modern language (Python, Go, Java, or similar)
Experience designing, implementing, and operating highly available, fault-tolerant, and scalable systems in production, including hybrid and multi-cloud environments
Experience owning CI/CD pipelines and release engineering practices
Track record of turning recurring operational work into self-service tooling, reusable modules, or golden-path automation
Experience with Docker and Kubernetes in production environments
Deep knowledge of monitoring, alerting, and observability platforms
Proven incident management skills, including high-severity incident response, root cause analysis, and postmortems
Proficient in disaster recovery, failover strategies, and resilience testing
Understanding of AWS, Azure, and hybrid environments
Excellent collaboration and communication skills
Nice to have: Experience with IAM platforms and security-focused systems
Nice to have: Infrastructure as Code and configuration management with Terraform, Ansible, and Puppet; scripting with Python, PowerShell, and Bash
Nice to have: Familiarity with OAuth2, OIDC, SAML, LDAP, and SCIM
Nice to have: Knowledge of enterprise security architecture and compliance frameworks
Nice to have: Exposure to AIOps or ML-based anomaly detection
Benefits
A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable
Leaders who support your development through coaching and managing opportunities
Ability to make a difference and lasting impact
Work in a dynamic, collaborative, progressive, and high-performing team
A world-class training program in financial services
DevOps Engineer designing production - style CI/CD, cloud, and infrastructure tasks for Your Software Supplier. Reviewing AI - generated solutions and ensuring correctness and reproducibility.
Senior DevOps Engineer modernizing AWS infrastructure and agentic automation for Big Viking Games’ live - service games. Improving reliability, security, observability, and production operations.
Support and Deployment Engineer deploying and optimizing ATEME video delivery solutions for broadcast and streaming customers. Troubleshooting Linux, networking, compression, and streaming systems across North America.
Senior Site Reliability Engineer deploying Kubernetes - based AI infrastructure on NVIDIA - certified hardware for Mirantis. Ensuring reliable, secure, scalable cloud operations and customer delivery.
DevOps Engineer building and maintaining cloud infrastructure, automation, and CI/CD pipelines for Calliere's software platform. Operating containers, observability tooling, and production workloads across public clouds.
Senior DevOps Engineer improving Sherweb’s cloud - based IT solution delivery through CI/CD, IaC, GitOps, and AI automation. Supporting secure platforms, operational transitions, and developer self - service.
Senior DevOps / Cloud Infrastructure Engineer needed for hybrid role in North York, ON. Requires 10+ years experience with GCP, AWS, Kubernetes, Terraform, and CI/CD.
Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.