Kubernetes/DevOps Engineer operating Kubernetes infrastructure for a petabyte-scale social media platform. Building distributed systems and machine-learning workloads for Capgemini Engineering’s client.
Responsibilities
Work as a Kubernetes/DevOps Engineer on a large-scale social media platform processing petabytes of data daily
Contribute as part of an R&D self-organized team in a challenging, innovative environment
Obtain tasks from the project lead or Team Lead
Prepare functional and design specifications and obtain stakeholder approval
Provide estimations, agree task duration with the manager, and contribute to the project plan
Update and optimize local and metadata models
Resolve crisis situations within the assigned area of responsibility
Understand business drivers and analytical use cases and translate them into data products
Initiate and conduct code reviews
Create code standards, conventions, and guidelines
Suggest technical and functional improvements to add value to the product
Requirements
Over 5+ years of experience as a Kubernetes/DevOps Engineer
University degree in Computer Science or a related field
Strong experience with Kubernetes and batch schedulers such as Volcano and Kueue
Understanding of distributed systems, containerization, and cloud-native architectures
Hands-on experience managing workloads on Kubernetes clusters such as AWS EKS
Proven track record working on infrastructure migrations
Proficiency in Python and Golang
Experience with infrastructure-as-code tools such as Terraform and Helm
Background in observability and monitoring for GPU clusters, such as Prometheus and Grafana
Strong problem-solving skills
Ability to operate effectively in a fast-paced, collaborative environment
Benefits
Paid time off based on employee grade: Vacation 12-25 days, depending on grade
Company paid holidays
Personal Days
Sick Leave
Medical, dental, and vision coverage or provincial healthcare coordination in Canada
Retirement savings plans, including RRSP in Canada
Life and disability insurance
Employee assistance programs
Other benefits as provided by local policy and eligibility
Potential eligibility for variable incentives, bonuses, or commissions
Senior AI Platform Developer building scalable AI services and LLM workflows for MaintainX’s industrial work execution platform. Improving reliability, observability, performance, and cost efficiency.
Infrastructure team lead building and operating Spare’s GCP and Kubernetes platform for on - demand transit. Leading developers while improving reliability, security, AI SRE, and cloud cost efficiency.
Senior AI Platform Developer building reusable, secure AI - agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Driving Azure - based platform capabilities across orchestration, evaluation, observability, and governance.
AI Platform Developer building Azure - based agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Creating secure, observable, governed AI services for product teams.
Lead AI Platform Engineer securing and operating EQ Bank’s enterprise AI platforms. Driving Azure engineering, observability, automation, governance, and production readiness.
Senior Platform Engineer building and operating GitLab Orbit, a Rust - based knowledge graph service for GitLab’s DevSecOps platform. Improving distributed - system reliability, observability, cloud infrastructure, and data workflows.
Platform Software Engineer supporting Invoice Simple, EverCommerce’s invoicing SaaS for small businesses. Building cloud infrastructure, CI/CD pipelines, and reliable production systems.
Software Engineer building cloud infrastructure and deployment systems for Invoice Simple, EverCommerce’s invoicing platform for small businesses. Improving reliability, observability, and production operations.
Senior cybersecurity engineer securing L3Harris’s mission - critical platform - management systems for military and cruise ships. Hardening Linux, networks, software, and infrastructure across the product lifecycle.