Staff Site Reliability Engineer

Posted 22 hours ago

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.

Responsibilities

  • Drive reliability engineering initiatives and operational excellence for mission-critical services running on AWS and Kubernetes
  • Design, implement, and improve deployment, release, and rollback strategies across distributed systems
  • Establish secure-by-default CI/CD pipelines with automation, governance, and policy-driven controls
  • Enhance platform observability through metrics, logs, tracing, and actionable alerting
  • Define and mature SLIs, SLOs, and reliability standards
  • Lead high-severity incident response and post-incident reviews
  • Partner with engineering teams to improve platform standards, service resilience, runtime performance, and reliability practices
  • Mentor engineers on cloud-native technologies, SRE principles, and operational excellence
  • Design, build, and evolve foundational systems, tooling, and operational practices
  • Architect scalable distributed systems and optimize Kubernetes and AWS infrastructure
  • Reduce operational toil, improve system performance, increase platform reliability, and support business growth

Requirements

  • 8+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or related cloud-native engineering roles
  • Deep expertise in AWS, including EKS, IAM, VPC, Lambda, CloudFront, S3, and cloud networking/security
  • Advanced experience operating and scaling production Kubernetes environments
  • Strong hands-on experience with Istio service mesh
  • Proven expertise with Infrastructure as Code, preferably AWS CDK
  • Experience building and managing CI/CD pipelines using GitHub Actions or similar platforms
  • Strong troubleshooting, performance optimization, and incident management experience in distributed systems
  • Excellent communication, collaboration, and technical leadership skills
  • Experience designing and operating monitoring, logging, tracing, and alerting solutions
  • Knowledge of AWS CloudWatch, OpenTelemetry, AWS X-Ray, and Kubernetes observability tooling
  • Experience defining SLIs, SLOs, alerting strategies, runbooks, and reliability metrics
  • Proficiency in TypeScript and Node.js
  • Experience building scalable backend services, APIs, and event-driven systems
  • Understanding of Kubernetes architecture, controllers, Gateway API, ingress management, and service networking
  • Experience implementing zero-trust architectures, mTLS, and service-to-service security controls
  • Automated testing, code reviews, and observability-driven development practices
  • Understanding of resilience engineering, autoscaling, disruption management, failure testing, and safe deployment strategies
  • Experience with progressive delivery, regulated or security-sensitive SaaS environments, FinOps, internal developer platforms, or cloud-native certifications is nice to have

Benefits

  • Flexible work options
  • Remote opportunities
  • Generous time-off policies
  • Competitive salary
  • Health insurance
  • Retirement plans
  • Recognition programs
  • Performance bonuses
  • Opportunities for career growth
  • International projects
  • Collaboration with a diverse, global team
  • Inclusive and equitable workplace
  • Accommodations during the application or interview process

Job type

Full Time

Experience level

Lead

Salary

Not specified

Degree requirement

No Education Requirement

Tech skills

AWSCloudDistributed SystemsJavaScriptKubernetesNode.jsRayTypeScript

Location requirements

RemoteCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.