DevOps Engineer managing AWS, Kubernetes, CI/CD, and observability for StellarTech’s scalable global EdTech products. Improving reliability, security, and production operations across multiple products.
Responsibilities
Design, implement, and maintain scalable cloud infrastructure
Establish and manage CI/CD pipelines for reliable software deployment
Manage Kubernetes services and troubleshoot cluster failures
Optimise performance across distributed systems
Handle production incidents and conduct post-mortems
Identify root causes and improve reliability processes
Implement and improve logging, monitoring, tracing, and alerting
Monitor infrastructure resource usage and optimise cloud efficiency
Manage global CDN configurations
Maintain SSO systems, databases, backups, and recovery processes
Support development, staging, and production environments
Automate infrastructure and operational processes using scripts and Infrastructure as Code
Collaborate with Backend, Product, and Engineering teams across multiple products
Requirements
Strong hands-on experience with AWS and cloud architecture and operations
Strong knowledge of Kubernetes and Docker, including production-grade cluster management
Experience with managed Kubernetes platforms such as EKS, GKE, or AKS
Advanced experience with Terraform, Helm, and Git
Hands-on experience implementing and maintaining GitHub Actions for CI/CD pipelines
Experience with Prometheus, Grafana, and OpenTelemetry
Ability to automate system and deployment tasks using Python and Bash
Strong troubleshooting skills and experience with distributed production systems
Experience using AI-assisted development tools such as Claude, Cursor, GitHub Copilot, ChatGPT, Codex, or similar
Hands-on experience managing API gateways for high-traffic applications
Experience with machine learning projects, including deploying, monitoring, and scaling ML workloads
Familiarity with machine learning platforms, frameworks, or libraries
Experience in a DBA role, including managing, maintaining, and optimising large-scale databases
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.
Senior Site Reliability Engineer securing AuthZed’s cloud infrastructure and authorization platform, including SpiceDB. Building Kubernetes guardrails, supply - chain security, vulnerability management, and incident response.
Principal SRE leading Cerence’s Site Reliability Engineering function for cloud - native automotive AI systems. Owning reliability strategy, incident escalation, observability, automation, and SLO governance.
Senior Data Scientist developing anomaly detection, dashboards, and alerts for General Motors vehicle reliability. Supporting engineering, quality, warranty, and software teams with production analytics.
DevOps Specialist scaling AWS infrastructure and automating deployments for Portage CyberTech’s digital trust, identity, privacy, and security solutions. Maintaining reliable cloud environments and CI/CD pipelines.
SRE Manager leading platform engineering teams across reliability and developer experience. Improving Lightspeed’s global cloud commerce SaaS platform through scalable infrastructure and automation.
SRE Manager leading platform engineering, reliability, and Developer Experience for Lightspeed’s global cloud POS SaaS platform. Guiding teams, infrastructure automation, and strategic delivery.
Gestionnaire SRE dirigeant deux équipes de plateforme chez Lightspeed. Assurant la fiabilité et l’automatisation de sa plateforme SaaS mondiale de commerce infonuagique.