Lead DevOps Engineer managing a distributed team across North America and India. Transforming engineering discipline with a focus on AI-enabled practices and efficient CI/CD workflows.
Responsibilities
Lead and mentor a distributed DevOps team spanning North America and India, including an infrastructure security-focused sub-team.
Serve as the primary technical decision-maker for the DevOps function — architecture, tooling choices, prioritization, and delivery standards.
Partner with Engineering, QA, and Product leadership to reduce delivery friction and improve DORA metrics (lead time, deployment frequency, MTTR, change fail rate).
Represent the DevOps function at the leadership level, including communicating roadmap, risks, and platform health to the Head of Technology and broader Technology leadership.
Own the self-hosted GitLab platform — upgrades, runner fleet management (VMware-hosted and cloud), and platform health.
Drive maturity of the CI/CD Catalog and shared template library (ci-common/gitlab-templates), ensuring teams can self-serve without bespoke pipeline configuration.
Establish and enforce merge request standards, branch protection policies, and CODEOWNERS governance across the GitLab organization.
Own EKS day-2 operations: cluster upgrades, node group management, networking (private API endpoints, Cloudflare tunnel integration), and reliability posture.
Manage multi-cloud infrastructure across AWS (primary) and GCP, including resource lifecycle, cloud cost optimization, and account governance.
Lead the rationalization of legacy infrastructure (on-prem Nexus, Concourse CI) and drive the migration to cloud-native equivalents where appropriate.
Maintain and improve the Terraform IaC estate, including drift detection, module governance, and GitLab CI-driven plan/apply workflows.
Drive the rollout and stabilization of SSO federation across vCenter/VMware, AWS IAM Identity Center, and Azure AD groups.
Own the security tooling stack: Qualys vulnerability scanning, Defender alert triage, container scanning pipelines, and SBOM/CVE reporting for product releases.
Establish and enforce secrets management standards using 1Password across pipelines and infrastructure automation.
Ensure data security in transit and at rest as automation and self-service capabilities expand.
Build and own the internal developer platform vision — reducing cognitive load on engineers, QA, and program managers through self-service tooling and automation.
Lead the observability stack: Grafana (Helm-deployed on EKS), alerting pipelines, and infrastructure/application performance monitoring.
Drive a metrics-first culture for the DevOps function, using DORA metrics and custom platform health indicators to guide roadmap decisions.
Evaluate and recommend tooling investments that improve developer experience, pipeline performance, and release confidence.
Requirements
5+ years of progressive DevOps/platform engineering experience, with at least 2 years in a technical lead or staff-level role.
Production Kubernetes experience (preferably EKS):
- Cluster upgrades, node management, networking, and RBAC
- Day-2 operations and reliability engineering
- GitLab-driven deployment workflows
Multi-cloud infrastructure proficiency across AWS and at least one of GCP/Azure:
- AWS IAM, Organizations, SSO/IAM Identity Center
- VPC networking, EKS, ECR, and cloud cost optimization
Infrastructure as Code with Terraform:
- Module design, remote state, drift detection
- CI/CD-driven plan/apply pipelines
Identity and access management:
- Azure AD / Microsoft Entra ID — SSO federation and group-based access
- Experience federating VMware vCenter, AWS, or similar platforms with AD/LDAP
Security tooling experience: vulnerability scanning (Qualys or equivalent), secrets management (1Password, Vault, or equivalent), SBOM/CVE pipeline integration.
Fluency in at least one scripting language (Bash, Python, or similar) for automation and tooling.
Strong written and verbal communication — able to write clear design documents, drive technical alignment, and represent the team in cross-functional and leadership conversations.
Demonstrated experience using AI tooling in an engineering context — whether in pipelines, developer tooling, observability, or infrastructure automation — and a clear point of view on where it creates genuine leverage vs. hype.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.