Senior Site Reliability Engineer managing cloud-based AI solutions at Mirantis. Focusing on Kubernetes infrastructure deployment and optimization while leading technical tasks.
Responsibilities
Work with geographically distributed international teams on technical challenges and process improvements.
Develop, implement, maintain, and troubleshoot cloud and AI infrastructure solutions based on open source software.
Collaborate with stakeholders to gather and refine technical requirements.
Optimize system performance, reliability, and scalability.
Troubleshoot, debug, and resolve complex technical issues.
Participate in code reviews to maintain high quality standards.
Stay up to date with industry trends and best practices in cloud operations and development.
Design and implement AI-driven automation across the DevOps lifecycle, including code development and maintenance.
Facilitate knowledge transfer to customers during the delivery phases.
Requirements
5+ years of professional experience in DevOps, with a strong focus on Cloud, infrastructure technologies and Kubernetes
Experience with high-performance data center processing, networking, and storage
Exposure to Golang and working knowledge of other programming languages (Python, JavaScript).
Strong knowledge of distributed systems, microservices architecture, and CI/CD pipelines.
Exceptional problem-solving and debugging skills across networking and storage (hardware and software), Linux, and Kubernetes, with attention to performance optimization and security.
Demonstrated ability to lead technical tasks and collaborate effectively with diverse teams.
Comfortable making independent judgment calls when working directly with customers, often with limited day-to-day oversight.
Excellent written and spoken English.
Excellent customer-facing communication skills.
A commitment to innovation, continuous learning, and delivering high-quality results.
Ability to travel up to 25% if needed, including internationally.
Benefits
Work with an established Silicon Valley leader in the cloud infrastructure industry;
Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
Be a part of cutting-edge, open-source innovation;
Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
Professional development and training;
Attend conferences and working groups;
Company outings, happy hours, hackathons, and tech talks;
Receive a competitive compensation package with a strong benefits plan.
Senior DevOps / Cloud Infrastructure Engineer needed for hybrid role in North York, ON. Requires 10+ years experience with GCP, AWS, Kubernetes, Terraform, and CI/CD.
Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.
Mozilla release engineer optimizing Firefox build, test, and deployment pipelines at global scale. Improving developer experience, maintaining automation, and responding to critical service outages without an on - call rotation.