Staff DevOps Engineer responsible for infrastructure that powers AI development pipeline for real estate platform. Collaborate with teams to enhance environments and tooling across AWS and EKS.
Responsibilities
Design and operate ephemeral, pre-warmed development environments that agents and engineers can spin up on demand.
Own the infrastructure-level gates that prove a deploy is safe before it reaches production.
Build containerized mock services (generated from OpenAPI specs) so contributors can validate integration code against realistic third-party dependencies locally.
Build the review tooling, metrics dashboards, and operational controls that make our pipelines observable and improvable.
Requirements
6–10+ years in DevOps, Platform Engineering, or SRE roles building and operating production systems at scale.
Active user of AI development tools (Claude Code, Codex, etc) in your infrastructure workflow.
Expertise with Kubernetes (EKS) and AWS (IAM, VPC, ECR, SSM/Secrets Manager, S3, SQS, Lambda, RDS/Aurora).
Strong IaC experience (Terraform preferred) and GitOps workflows (Argo CD or similar).
Proven track record building ephemeral environments, developer tooling, or internal platforms (CLIs, scaffolding tools, developer portals).
Experience with load testing frameworks (k6, Locust, Gatling, or similar) and automating performance gates in CI/CD pipelines.
Examples of building mock or stub infrastructure for integration testing at scale — containerized services, API mocking, dependency isolation.
CI/CD depth (CircleCI, GitHub Actions, or similar) including caching/parallelism, artifact management, test reliability, and pipeline observability.
Experience with release strategies (canary/blue-green, automated rollbacks) and progressive delivery.
Observability fundamentals (Datadog, OpenTelemetry) with the ability to define SLIs/SLOs and wire them to delivery decisions.
Excellent cross-team communicator who can translate platform constraints into developer-friendly solutions and documentation.
DevOps Intern supporting CI/CD, cloud infrastructure, and automation for Ludia’s mobile game studio. Improving reliability and developer tools in production game environments.
Manager leading global SRE teams for Akamai's distributed Cloud IAM services. Improving reliability, scalability, security, and usability through cloud - native tooling and software.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.