Senior DevOps Engineer designing and operating cloud-native infrastructure for distributed systems at ELITS. Collaborating with teams to ensure reliable streaming and high availability in production.
Responsibilities
Design, deploy and operate containerized microservices and distributed systems in production Kubernetes environments.
Build and maintain CI/CD pipelines to enable frequent, reliable releases and automated testing.
Implement and manage real‑time streaming data platforms (for example, Kafka or similar technologies) for low‑latency, high‑throughput workloads.
Design and operate infrastructure with a strong focus on reliability, performance and cost‑efficiency across cloud and on‑prem/hybrid environments.
Own infrastructure as code (IaC) using tools such as Terraform and Helm for repeatable, auditable environments.
Monitor, troubleshoot and optimize Linux‑based systems, containers and services, including performance tuning and incident response.
Collaborate with development teams to improve operability, observability and resilience of services (SRE mindset).
Document architectures, runbooks and operational procedures, and contribute to continuous improvement of processes and tooling.
Requirements
10+ years of experience in DevOps, SRE, Platform Engineering or similar roles.
Strong hands‑on experience with streaming technologies and real‑time data processing (for example, Apache Kafka, Kinesis, Pulsar or equivalent).
Solid background in distributed systems: microservices, event‑driven architectures, scalability and fault tolerance.
Strong understanding of hardware and infrastructure concepts (servers, networking, storage) and experience with on‑prem or hybrid environments.
Deep knowledge of Linux/Unix operating systems, system internals, performance and troubleshooting.
Extensive experience with cloud‑native technologies: • Containers and orchestration: Docker, Kubernetes (AKS/EKS/GKE or similar) • Infrastructure as Code: Terraform, Helm (and/or similar tools) • CI/CD pipelines: GitHub Actions, Jenkins, Argo CD or equivalent • Observability: monitoring, logging and alerting (for example, ELK/EFK, Prometheus, Grafana).
Experience with at least one major cloud provider (Azure, AWS or GCP); Azure experience is a strong asset.
Good understanding of networking (VPN, IPsec, load balancing, DNS, certificates).
Experience with agile ways of working and tools such as JIRA and Git.
Strong debugging and troubleshooting abilities across multiple layers (application, infrastructure, network).
Ability to understand users’ technical issues and provide clear, pragmatic recommendations.
DevOps Intern supporting CI/CD, cloud infrastructure, and automation for Ludia’s mobile game studio. Improving reliability and developer tools in production game environments.
Manager leading global SRE teams for Akamai's distributed Cloud IAM services. Improving reliability, scalability, security, and usability through cloud - native tooling and software.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.