SRE Manager leading platform engineering teams across reliability and developer experience. Improving Lightspeed’s global cloud commerce SaaS platform through scalable infrastructure and automation.
Responsibilities
Lead two teams of 4–6 platform engineers across Site Reliability Engineering and Developer Experience
Ensure reliability, scalability, and optimal performance of Lightspeed’s global cloud-based POS SaaS platform
Coach, mentor, guide, and motivate senior cross-functional cloud platform engineers
Drive performance management, goal setting, and individual development plans
Foster teamwork, creativity, innovation, and high performance
Develop and execute the SRE team’s strategic vision aligned with company objectives
Understand business-unit roadmaps and cross-business interactions
Define the team vision and prioritize projects improving developer experience, production stability, and scalability
Build and project-manage the team roadmap and establish execution processes
Lead scalable, reliable, and efficient infrastructure design and implementation using modern cloud technologies
Drive automation for deployment, monitoring, and operational processes
Collaborate with engineering, product, and SRE teams across departments
Drive cross-organization projects
Communicate with stakeholders, gather requirements, provide updates, and address concerns
Participate in the on-call rotation roster when required
Requirements
Proven experience working autonomously on projects across multiple streams of work
Track record of successful project delivery through high-performing teams
Extensive hands-on experience with cloud infrastructure providers such as AWS and GCP
Proven expertise in orchestrating and managing infrastructure through code
Experience across engineering and leadership roles
Ability to lead large teams across complex projects and shape the strategic and technical direction of the platform engineering team
Ability to provide strategic direction and set long-term team objectives
Proven experience implementing improvements to processes spanning team boundaries
Experience leading medium to large-sized teams of platform engineers
Applicants must disclose criminal convictions; criminal record checks are part of the hiring process
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.
Senior Site Reliability Engineer securing AuthZed’s cloud infrastructure and authorization platform, including SpiceDB. Building Kubernetes guardrails, supply - chain security, vulnerability management, and incident response.
Principal SRE leading Cerence’s Site Reliability Engineering function for cloud - native automotive AI systems. Owning reliability strategy, incident escalation, observability, automation, and SLO governance.
Senior Data Scientist developing anomaly detection, dashboards, and alerts for General Motors vehicle reliability. Supporting engineering, quality, warranty, and software teams with production analytics.
DevOps Specialist scaling AWS infrastructure and automating deployments for Portage CyberTech’s digital trust, identity, privacy, and security solutions. Maintaining reliable cloud environments and CI/CD pipelines.
SRE Manager leading platform engineering, reliability, and Developer Experience for Lightspeed’s global cloud POS SaaS platform. Guiding teams, infrastructure automation, and strategic delivery.
Gestionnaire SRE dirigeant deux équipes de plateforme chez Lightspeed. Assurant la fiabilité et l’automatisation de sa plateforme SaaS mondiale de commerce infonuagique.
DevOps Engineer managing AWS, Kubernetes, CI/CD, and observability for StellarTech’s scalable global EdTech products. Improving reliability, security, and production operations across multiple products.