Senior Site Reliability Engineer optimizing Kubernetes infrastructure and AI tooling at SecurityScorecard. Driving best practices in automation, observability, and resilience through team collaboration.
Responsibilities
Design, build, and scale Kubernetes infrastructure for secure, multi-tenant, high-availability applications.
Build and operate AI tooling infrastructure — stand up MCP servers and establish secure, governed AI access and guardrails for production systems.
Optimize and maintain CI/CD pipelines, improving reliability, speed, and rollback safety.
Implement progressive delivery strategies such as blue/green and canary deployments.
Advance Infrastructure as Code with Terraform, Helm, and Argo CD, defining reusable patterns for the org.
Operate and optimize streaming and analytics infrastructure: Kafka, Flink, and ClickHouse.
Build automated testing into the CI/CD lifecycle.
Improve system observability — define SLOs, alerts, and dashboards.
Lead incident response and postmortems, focusing on root cause and durable fixes.
Mentor engineers across teams on Kubernetes, CI/CD, and cloud infrastructure.
Requirements
6+ years in SRE, DevOps, or Infrastructure roles, with significant production Kubernetes experience.
Hands-on experience integrating AI/LLM tooling into engineering or operational workflows (e.g., MCP servers, AI agents acting on infrastructure), and a clear grasp of the security and governance considerations of giving AI access to production.
Proven success building CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, or similar).
Strong with Kubernetes internals and managed services like EKS, GKE, or AKS.
Expertise with Infrastructure as Code (Terraform, Helm, Pulumi) and GitOps.
Proficient in Python, Bash, or Go.
Knowledge of observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry).
Production experience with Kafka, Flink, and ClickHouse.
Strong communication and cross-team collaboration skills.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.