Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
Responsibilities
Lead the Site Reliability Engineering team as a 60% individual contributor and 40% leadership role
Own direct reports' onboarding, feedback, performance assessment, progression and hiring
Steer the team's focus against company goals and represent the team across engineering and senior leadership
Set technical direction and review the team's work
Own SRE goals, prioritization, support rotation and on-call model
Oversee Kubernetes, AWS, PostgreSQL, DNS and TLS, and CI infrastructure
Develop the reliability practice, including SLOs, error budgets, incident response and observability
Partner with Security on threats, patching, infrastructure controls, audit and compliance obligations
Manage platform vendor relationships, renewals and commercial conversations
Improve operational load versus project delivery balance and extend the SLO framework across teams
Requirements
Experience leading an SRE, infrastructure or platform engineering team
Ownership of reports' growth, performance and career progression
Experience coaching technical craft and soft skills
Experience handling underperformance directly and early
Experience hiring engineers and assessing engineering quality
Hands-on background in site reliability, DevOps or cloud infrastructure engineering
Kubernetes in production
AWS at meaningful scale
Hands-on AI building, enablement, and scaling AI infrastructure
Solid observability practices and principles
Infrastructure as code with Terraform
CI/CD systems such as GitLab CI, GitHub Actions or Jenkins
Docker and shell scripting
Experience running a reliability practice: incident response, on-call, SLOs and error budgets
Understanding and history of working in regulated environments
Ability to prioritize operational load and project work
Clear written communication for asynchronous work
Ability to build relationships across teams
English application materials required
Benefits
work from anywhere
flexible paid time off
flexible working hours (we are async)
16 weeks paid parental leave
budget towards co-working spaces, learning and wellness (including gym memberships)
mental health support services
stock options
home office budget & IT equipment
life-work balance and schedule flexibility
employee resource groups (Women, Disability, Queer, Minorities in Tech)
accommodation support during interviews and beyond
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.
Senior Site Reliability Engineer securing AuthZed’s cloud infrastructure and authorization platform, including SpiceDB. Building Kubernetes guardrails, supply - chain security, vulnerability management, and incident response.
Principal SRE leading Cerence’s Site Reliability Engineering function for cloud - native automotive AI systems. Owning reliability strategy, incident escalation, observability, automation, and SLO governance.
Senior Data Scientist developing anomaly detection, dashboards, and alerts for General Motors vehicle reliability. Supporting engineering, quality, warranty, and software teams with production analytics.
DevOps Specialist scaling AWS infrastructure and automating deployments for Portage CyberTech’s digital trust, identity, privacy, and security solutions. Maintaining reliable cloud environments and CI/CD pipelines.
SRE Manager leading platform engineering teams across reliability and developer experience. Improving Lightspeed’s global cloud commerce SaaS platform through scalable infrastructure and automation.
SRE Manager leading platform engineering, reliability, and Developer Experience for Lightspeed’s global cloud POS SaaS platform. Guiding teams, infrastructure automation, and strategic delivery.
Gestionnaire SRE dirigeant deux équipes de plateforme chez Lightspeed. Assurant la fiabilité et l’automatisation de sa plateforme SaaS mondiale de commerce infonuagique.
DevOps Engineer managing AWS, Kubernetes, CI/CD, and observability for StellarTech’s scalable global EdTech products. Improving reliability, security, and production operations across multiple products.