Principal Site Reliability Engineer responsible for building shared infrastructure services for a mission-critical SaaS platform used by global enterprises. Influencing architecture and engineering culture while solving complex reliability challenges.
Responsibilities
Work on a mission-critical SaaS platform used by global enterprises
Solve complex reliability challenges at scale
Influence architecture and engineering culture at a company level
Design, build, and maintain shared infrastructure services and platforms
Create reusable, reliable, and scalable solutions
Architect, implement, and manage highly available and scalable Kubernetes platforms
Develop robust, internal-facing tools and automation using Go (Golang)
Design and implement shared Event-Driven Architecture components
Develop and maintain robust CI/CD pipelines as a service
Design and build resilient Distributed Systems components
Manage and optimize shared infrastructure across Multi-Region Cloud Environments
Establish and enhance centralized Observability and Monitoring platforms
Define and implement clear, well-documented RESTful API designs
Implement and manage Service Mesh capabilities
Design, implement, and optimize highly available Relational Database services
Collaborate closely with product development teams
Participate in on-call rotations to support critical shared infrastructure
Requirements
9+ years of experience in an Infrastructure Development, Platform Engineering, or Site Reliability Engineering role
Deep expertise with Kubernetes in production environments
Strong programming skills in Go (Golang) and Python
Extensive hands-on experience with at least one major Cloud Provider (AWS, GCP, or Azure)
Proven experience designing and implementing Event-Driven Architecture and message queuing systems
Solid understanding and practical experience with CI/CD pipeline tools
Demonstrable experience designing and operating Distributed Systems
Familiarity with Multi-Region Cloud Environments
Proficiency in establishing and utilizing comprehensive Observability and Monitoring platforms
Strong experience with RESTful API design principles
Knowledge of Service Mesh concepts
Hands-on experience with Relational Databases
Excellent communication skills
A strong customer-centric mindset
Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience or equivalent military experience required
Benefits
Competitive compensation, benefits, and growth opportunities
Security & Compliance: Compliance with Saviynt’s information security and privacy policies, including annual security training
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.