Engineering Manager leading Site Reliability Engineers in developing reliable cloud infrastructure at Tempo. Ensure stability, cost efficiency, and effective team management in a SaaS environment.
Responsibilities
Lead, mentor, and grow a team of Site Reliability Engineers, focusing on career development, performance management, and hiring.
Define the team's roadmap and strategy for platform reliability, scaling, and operational efficiency.
Provide technical oversight and direction for the design and implementation of key infrastructure projects, including CI/CD pipelines and automation for build, release, and deployment processes.
Partner closely with engineering teams and product managers to ensure the reliability and performance requirements of new products and features are met.
Oversee the maintenance and continuous improvement of the AWS-based platform to ensure it scales effectively.
Drive the adoption of AI tooling to enhance SRE productivity and introduce intelligent automation of SRE processes.
Champion SRE best practices, including error budget management, effective on-call rotations, incident response, and post-mortem processes.
Requirements
6+ years of progressive experience in a SaaS environment, with 2+ years of experience managing or leading high-performing SRE or Infrastructure teams.
Proven experience in defining strategy and overseeing the deployment of complex software solutions in a fast-paced, cloud environment.
Working knowledge of AWS or other cloud service providers.
Solid understanding of SRE and DevOps principles, software design patterns, and infrastructure operations.
Passionate about containerization and orchestration technologies like Kubernetes.
Familiarity with monitoring, alerting, and observability tools, including RUM (Real User Metrics), tracing, and other vital metrics.
Demonstrated ability to lead cross-functional projects, manage ambiguity, and drive technical decision-making.
Exceptional communication, collaboration, and analytical skills, with a passion for solving tough technical and organizational problems.
Benefits
Remote First work environment
Unlimited vacation in most of our locations!!
Great benefits including health, dental, vision and savings plan.
Perks such as training reimbursement, WFH reimbursement, and more.
Diverse and dynamic teams with challenging and exciting work.
An opportunity to have a real impact on our business.
A great range of social activities (both in person and virtual).
Optional in person meet-ups and the ability to travel to our international offices
DevOps Intern supporting CI/CD, cloud infrastructure, and automation for Ludia’s mobile game studio. Improving reliability and developer tools in production game environments.
Manager leading global SRE teams for Akamai's distributed Cloud IAM services. Improving reliability, scalability, security, and usability through cloud - native tooling and software.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.