Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
Responsibilities
Drive reliability engineering initiatives and operational excellence for mission-critical services running on AWS and Kubernetes
Design, implement, and improve deployment, release, and rollback strategies across distributed systems
Establish secure-by-default CI/CD pipelines with automation, governance, and policy-driven controls
Enhance platform observability through metrics, logs, tracing, and actionable alerting
Define and mature SLIs, SLOs, and reliability standards
Lead high-severity incident response and post-incident reviews
Partner with engineering teams to improve platform standards, service resilience, runtime performance, and reliability practices
Mentor engineers on cloud-native technologies, SRE principles, and operational excellence
Design, build, and evolve foundational systems, tooling, and operational practices
Architect scalable distributed systems and optimize Kubernetes and AWS infrastructure
Reduce operational toil, improve system performance, increase platform reliability, and support business growth
Requirements
8+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or related cloud-native engineering roles
Deep expertise in AWS, including EKS, IAM, VPC, Lambda, CloudFront, S3, and cloud networking/security
Advanced experience operating and scaling production Kubernetes environments
Strong hands-on experience with Istio service mesh
Proven expertise with Infrastructure as Code, preferably AWS CDK
Experience building and managing CI/CD pipelines using GitHub Actions or similar platforms
Strong troubleshooting, performance optimization, and incident management experience in distributed systems
Excellent communication, collaboration, and technical leadership skills
Experience designing and operating monitoring, logging, tracing, and alerting solutions
Knowledge of AWS CloudWatch, OpenTelemetry, AWS X-Ray, and Kubernetes observability tooling
Experience defining SLIs, SLOs, alerting strategies, runbooks, and reliability metrics
Proficiency in TypeScript and Node.js
Experience building scalable backend services, APIs, and event-driven systems
Understanding of Kubernetes architecture, controllers, Gateway API, ingress management, and service networking
Experience implementing zero-trust architectures, mTLS, and service-to-service security controls
Automated testing, code reviews, and observability-driven development practices
Understanding of resilience engineering, autoscaling, disruption management, failure testing, and safe deployment strategies
Experience with progressive delivery, regulated or security-sensitive SaaS environments, FinOps, internal developer platforms, or cloud-native certifications is nice to have
Benefits
Flexible work options
Remote opportunities
Generous time-off policies
Competitive salary
Health insurance
Retirement plans
Recognition programs
Performance bonuses
Opportunities for career growth
International projects
Collaboration with a diverse, global team
Inclusive and equitable workplace
Accommodations during the application or interview process
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.