DevOps Engineer working with cloud technologies to support engineering environments at AI platform for managing people and money. Collaborating with developers and maintaining high availability of systems.
Responsibilities
Design infrastructure and automated systems to support our distributed architecture
Build and Manage CI/CD pipelines and constantly improve their reliability & speed, and reduce lead time for changes
Forecast and plan for the infrastructure needs of a fast-growing SaaS company and find ways to improve efficiency and cost
Trace performance bottlenecks and identify optimizations and improvements at both the infrastructure and application level
Collaborate with our engineering team to meet high SLO and SLA requirements from customers
Maintain highly available web and backend systems that serve millions of users, and 1000’s of requests per second
Closely collaborating with Developers to setup, configure and plan the necessary cloud services in support of new feature development on AWS
Analyze system performance and capacity and plan for future growth
Securing our infrastructure at both the cloud layer (IAM) and application layer (PKI)
Building and expanding monitoring and alerting systems for both infrastructure and business operations, using internal tools & integrating into established 3rd party SaaS ones
Establishing comprehensive infrastructure-as-code coverage to support our entire platform
Develop tools to enhance and support Developer Productivity
Champion automation of manual processes and reducing operational overhead
Requirements
3+ years DevOps or software development experience
3+ years of experience building, maintaining and scaling database technologies such as Postgresql, MySQL, Redis, and DynamoDB
2+ years experience orchestrating large scale distributed microservice deployments on Kubernetes and EC2
2+ years experience building and managing EKS clusters and strong knowledge of the K8s ecosystem
2+ years of experience with Prometheus/Grafana/Cloudwatch metrics monitoring, ELK/OpenSearch stack for logging and PagerDuty and alerting
Benefits
Workday Bonus Plan or role-specific commission/bonus
Principal SRE leading Cerence’s Site Reliability Engineering function for cloud - native automotive AI systems. Owning reliability strategy, incident escalation, observability, automation, and SLO governance.
Senior Data Scientist developing anomaly detection, dashboards, and alerts for General Motors vehicle reliability. Supporting engineering, quality, warranty, and software teams with production analytics.
DevOps Specialist scaling AWS infrastructure and automating deployments for Portage CyberTech’s digital trust, identity, privacy, and security solutions. Maintaining reliable cloud environments and CI/CD pipelines.
SRE Manager leading platform engineering teams across reliability and developer experience. Improving Lightspeed’s global cloud commerce SaaS platform through scalable infrastructure and automation.
SRE Manager leading platform engineering, reliability, and Developer Experience for Lightspeed’s global cloud POS SaaS platform. Guiding teams, infrastructure automation, and strategic delivery.
Gestionnaire SRE dirigeant deux équipes de plateforme chez Lightspeed. Assurant la fiabilité et l’automatisation de sa plateforme SaaS mondiale de commerce infonuagique.
DevOps Engineer managing AWS, Kubernetes, CI/CD, and observability for StellarTech’s scalable global EdTech products. Improving reliability, security, and production operations across multiple products.
Principal DevOps Engineer leading AWS, Kubernetes, and Terraform infrastructure for Campspot’s campground reservation software and camping marketplace. Driving automation, reliability, security, and cloud - cost optimization.