Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
Responsibilities
Drive reliability engineering initiatives and operational excellence for mission-critical services running on AWS and Kubernetes
Design, implement, and improve deployment, release, and rollback strategies across distributed systems
Establish secure-by-default CI/CD pipelines with automation, governance, and policy-driven controls
Enhance platform observability through metrics, logs, tracing, and actionable alerting
Define and mature SLIs, SLOs, and reliability standards
Lead high-severity incident response and post-incident reviews
Partner with engineering teams to improve platform standards, service resilience, runtime performance, and reliability practices
Mentor engineers on cloud-native technologies, SRE principles, and operational excellence
Design, build, and evolve foundational systems, tooling, and operational practices
Architect scalable distributed systems and optimize Kubernetes and AWS infrastructure
Reduce operational toil, improve system performance, increase platform reliability, and support business growth
Requirements
8+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or related cloud-native engineering roles
Deep expertise in AWS, including EKS, IAM, VPC, Lambda, CloudFront, S3, and cloud networking/security
Advanced experience operating and scaling production Kubernetes environments
Strong hands-on experience with Istio service mesh
Proven expertise with Infrastructure as Code, preferably AWS CDK
Experience building and managing CI/CD pipelines using GitHub Actions or similar platforms
Strong troubleshooting, performance optimization, and incident management experience in distributed systems
Excellent communication, collaboration, and technical leadership skills
Experience designing and operating monitoring, logging, tracing, and alerting solutions
Knowledge of AWS CloudWatch, OpenTelemetry, AWS X-Ray, and Kubernetes observability tooling
Experience defining SLIs, SLOs, alerting strategies, runbooks, and reliability metrics
Proficiency in TypeScript and Node.js
Experience building scalable backend services, APIs, and event-driven systems
Understanding of Kubernetes architecture, controllers, Gateway API, ingress management, and service networking
Experience implementing zero-trust architectures, mTLS, and service-to-service security controls
Automated testing, code reviews, and observability-driven development practices
Understanding of resilience engineering, autoscaling, disruption management, failure testing, and safe deployment strategies
Experience with progressive delivery, regulated or security-sensitive SaaS environments, FinOps, internal developer platforms, or cloud-native certifications is nice to have
Benefits
Flexible work options
Remote opportunities
Generous time-off policies
Competitive salary
Health insurance
Retirement plans
Recognition programs
Performance bonuses
Opportunities for career growth
International projects
Collaboration with a diverse, global team
Inclusive and equitable workplace
Accommodations during the application or interview process
Senior DevOps / Cloud Infrastructure Engineer needed for hybrid role in North York, ON. Requires 10+ years experience with GCP, AWS, Kubernetes, Terraform, and CI/CD.
Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.
Mozilla release engineer optimizing Firefox build, test, and deployment pipelines at global scale. Improving developer experience, maintaining automation, and responding to critical service outages without an on - call rotation.
DevOps Manager overseeing releases, enterprise tooling, and incident response for Delta Controls, a building - automation solutions manufacturer. Establishing standards across global product teams and offices.