DevOps Engineer responsible for multi-cloud infrastructure across Azure, AWS, and GCP. Collaborate with teams to build CI/CD pipelines and implement automation for AI applications.
Responsibilities
Design, build, and maintain CI/CD pipelines that support continuous delivery across multiple cloud platforms (Azure DevOps, GitHub Actions, GitLab CI)
Architect and manage cloud infrastructure across Azure, AWS, and GCP using Infrastructure as Code (Terraform, Bicep, CloudFormation)
Manage containerized application workloads using Kubernetes (AKS, EKS, or GKE) and Docker
Implement and maintain cloud security best practices: IAM policies, network segmentation, secrets management, vulnerability scanning
Design and maintain observability stacks — logging, metrics, alerting — using tools such as Azure Monitor, CloudWatch, Datadog, or Grafana/Prometheus
Collaborate with software and ML engineering teams to define deployment strategies, optimize release pipelines, and reduce deployment risk
Evaluate and introduce tooling improvements that enhance reliability, scalability, and developer productivity
Contribute to incident response and post-mortem processes, driving root cause analysis and corrective actions
Build and maintain internal documentation on infrastructure architecture, operational runbooks, and DR procedures
Mentor junior team members and provide technical guidance on cloud and DevOps best practices
Requirements
Degree or equivalent work experience in Computer Science, Systems Engineering, or a related discipline
3–6 years of progressive DevOps, cloud engineering, or site reliability engineering experience
Strong hands-on experience with at least two of: Azure, AWS, GCP — multi-cloud exposure is highly valued
Proven experience building and maintaining CI/CD pipelines in production environments
Proficiency with Infrastructure as Code: Terraform required; Bicep, Pulumi, or CDK are a plus
Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.
Mozilla release engineer optimizing Firefox build, test, and deployment pipelines at global scale. Improving developer experience, maintaining automation, and responding to critical service outages without an on - call rotation.
DevOps Manager overseeing releases, enterprise tooling, and incident response for Delta Controls, a building - automation solutions manufacturer. Establishing standards across global product teams and offices.