Senior Site Reliability Engineer deploying Kubernetes-based AI infrastructure on NVIDIA-certified hardware for Mirantis. Ensuring reliable, secure, scalable cloud operations and customer delivery.
Responsibilities
Work with geographically distributed international teams on technical challenges and process improvements
Develop, implement, maintain, and troubleshoot cloud and AI infrastructure solutions based on open source software
Deploy AI infrastructure built on NVIDIA-certified hardware
Collaborate with stakeholders to gather and refine technical requirements
Optimize system performance, reliability, and scalability
Troubleshoot, debug, and resolve complex technical issues
Participate in code reviews to maintain high quality standards
Stay up to date with industry trends and best practices in cloud operations and development
Design and implement AI-driven automation across the DevOps lifecycle, including code development and maintenance
Facilitate knowledge transfer to customers during delivery phases
Mentor team members and Mirantis customers
Define technical strategies and ensure seamless integration of cloud and software services
Ensure reliability, security, and performance of container infrastructure
Requirements
5+ years of professional experience in DevOps, with a strong focus on Cloud, infrastructure technologies and Kubernetes
Experience with high-performance data center processing, networking, and storage
Exposure to Golang and working knowledge of other programming languages (Python, JavaScript)
Strong knowledge of distributed systems, microservices architecture, and CI/CD pipelines
Exceptional problem-solving and debugging skills across networking and storage (hardware and software), Linux, and Kubernetes
Knowledge of performance optimization and security
Ability to lead technical tasks and collaborate effectively with diverse teams
Comfortable making independent judgment calls when working directly with customers, often with limited day-to-day oversight
Excellent written and spoken English
Excellent customer-facing communication skills
Commitment to innovation, continuous learning, and delivering high-quality results
Ability to travel up to 25% if needed, including internationally
Bachelor's degree in Computer Science or a related field, or equivalent experience
At least 5 years of DevOps or Software Development experience or in a similar role
Benefits
Professional development and training
Attend conferences and working groups
Company outings, happy hours, hackathons, and tech talks
Competitive compensation package with a strong benefits plan
DevOps Engineer building and maintaining cloud infrastructure, automation, and CI/CD pipelines for Calliere's software platform. Operating containers, observability tooling, and production workloads across public clouds.
Senior DevOps Engineer improving Sherweb’s cloud - based IT solution delivery through CI/CD, IaC, GitOps, and AI automation. Supporting secure platforms, operational transitions, and developer self - service.
Senior DevOps / Cloud Infrastructure Engineer needed for hybrid role in North York, ON. Requires 10+ years experience with GCP, AWS, Kubernetes, Terraform, and CI/CD.
Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.
Mozilla release engineer optimizing Firefox build, test, and deployment pipelines at global scale. Improving developer experience, maintaining automation, and responding to critical service outages without an on - call rotation.