Senior Staff DevOps Engineer scaling AWS Kubernetes infrastructure for SailPoint’s identity security platform.
Leading enterprise service mesh adoption, PCI-compliant operations, and cloud-native reliability across global teams.
Responsibilities
Design, build, and operate SailPoint’s global Identity Security Cloud infrastructure on AWS
Lead the full lifecycle of service mesh implementation across hundreds of microservices and production EKS clusters
Establish organization-wide service mesh standards, governance, onboarding runbooks, sidecar injection policies, traffic policy templates, and failure-mode playbooks
Mentor and upskill engineers on service mesh architecture, troubleshooting, performance tuning, and capacity planning
Design, operate, and optimize production Kubernetes clusters on AWS EKS, including architecture, upgrades, node management, networking, storage, and multi-tenant isolation
Define and drive Kubernetes standards and best practices across teams
Design and scale infrastructure for customer demand, data sovereignty requirements, and regional expansion
Automate deployment, monitoring, incident response, and capacity management using GitOps and CI/CD
Develop and improve operational practices, runbooks, and platform engineering standards
Collaborate with development teams to bring new features and services into production safely
Support SailPoint’s PCI compliance initiative through secure platform design, controls implementation, and audit readiness
Participate in and improve the global follow-the-sun on-call rotation
Drive post-incident reviews and systemic fixes
Influence architecture direction, mentor engineers, and drive operational excellence without people management
Requirements
Bachelor's and/or Master's degree in Computer Science or equivalent technical experience
3+ years of hands-on experience designing, implementing, and operating a service mesh at scale in production Kubernetes environments
10+ years of experience in 24x7 production operations supporting highly available SaaS or cloud service environments
10+ years of experience with containerization, virtualization, and configuration management technologies
5+ years of hands-on experience with Kubernetes in production at scale
5+ years of experience with Terraform, managing infrastructure across multiple AWS accounts and regions
5+ years of experience designing and implementing CI/CD pipelines, especially for Terraform, Kubernetes, and microservices
5+ years of experience with scripting/programming languages such as Python or Go
Strong shell scripting proficiency
Strong understanding of Linux, networking, distributed systems, and production troubleshooting
Demonstrated experience scaling a service mesh across a large microservices fleet, including phased adoption strategy, sidecar resource management, control plane scaling, and performance tuning
Experience with monitoring and logging stacks such as Prometheus, Grafana, and OpenSearch or equivalent
Prior experience as a technical lead or Staff+ individual contributor in a global engineering organization
Strong interpersonal and teaming skills, with ability to set and enforce process and influence engineers across teams and geographies
Ability to operate effectively in an agile, entrepreneurial environment with global stakeholders
Candidates must be eligible to work in the United States or Canada
Availability for collaboration with global teams across US, EMEA, and APAC time zones
Participation in an on-call rotation is required
Preferred: multi-cluster or multi-region service mesh federation experience in production
Preferred: familiarity with PCI DSS and FedRAMP-adjacent practices
Benefits
Medical, dental, and vision insurance
Short-term and long-term disability coverage
Life insurance and Accidental Death & Dismemberment (AD&D)
Supplemental life insurance for employees, spouses, and children
Flexible spending accounts for health care and dependent care
Limited purpose flexible spending account
401(k) Savings and Investment Plan with company matching
Flexible vacation policy
8 paid holidays annually
Sick leave
Paid parental leave
Employee Assistance Program (EAP) and Care Counselors
Voluntary Legal Assistance, Critical Illness, Accident, Hospital Indemnity and Pet Insurance options
Health Savings Account (HSA) with employer contribution
May be eligible for the SailPoint Corporate Bonus Plan or role-specific commission
Principal SRE leading Cerence’s Site Reliability Engineering function for cloud - native automotive AI systems. Owning reliability strategy, incident escalation, observability, automation, and SLO governance.
Senior Data Scientist developing anomaly detection, dashboards, and alerts for General Motors vehicle reliability. Supporting engineering, quality, warranty, and software teams with production analytics.
DevOps Specialist scaling AWS infrastructure and automating deployments for Portage CyberTech’s digital trust, identity, privacy, and security solutions. Maintaining reliable cloud environments and CI/CD pipelines.
SRE Manager leading platform engineering teams across reliability and developer experience. Improving Lightspeed’s global cloud commerce SaaS platform through scalable infrastructure and automation.
SRE Manager leading platform engineering, reliability, and Developer Experience for Lightspeed’s global cloud POS SaaS platform. Guiding teams, infrastructure automation, and strategic delivery.
Gestionnaire SRE dirigeant deux équipes de plateforme chez Lightspeed. Assurant la fiabilité et l’automatisation de sa plateforme SaaS mondiale de commerce infonuagique.
DevOps Engineer managing AWS, Kubernetes, CI/CD, and observability for StellarTech’s scalable global EdTech products. Improving reliability, security, and production operations across multiple products.
Principal DevOps Engineer leading AWS, Kubernetes, and Terraform infrastructure for Campspot’s campground reservation software and camping marketplace. Driving automation, reliability, security, and cloud - cost optimization.