Site Reliability Engineer responsible for building and deploying satellite operations at Planet. Collaborating with teams to design and implement reliable systems for cloud and on-premises environments.
Responsibilities
Build and deploy computing services and infrastructure in customer environments for a next-generation satellite operations and image processing end-to-end platform
Operate in a high-impact, tight knit team to architect novel systems for air-gapped deployments at scale
Clarify and surface requirements from ambiguous use cases defined by cross-functional stakeholders, including internal users and external customers
Responsible for operations such as deployments, service orchestration, and documentation for cross platform stakeholders
Scale architecture while ensuring availability of services
Improve reliability and scalability by resolving edge cases, studying failure modes, and writing tests
Participate in on-call rotations to ensure operational excellence
Requirements
Bachelor’s degree in Computer Science or similar
10+ years of experience building services that leverage cloud-native infrastructure and tooling
Experience deploying and maintaining bare-metal and cloud kubernetes through tools such as Talos, RKE2, Proxmox, or k3s
Proficiency with Terraform, Ansible, Helm, Kustomize, and/or similar IaC / GitOps tooling
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.
Senior Site Reliability Engineer securing AuthZed’s cloud infrastructure and authorization platform, including SpiceDB. Building Kubernetes guardrails, supply - chain security, vulnerability management, and incident response.
Principal SRE leading Cerence’s Site Reliability Engineering function for cloud - native automotive AI systems. Owning reliability strategy, incident escalation, observability, automation, and SLO governance.
Senior Data Scientist developing anomaly detection, dashboards, and alerts for General Motors vehicle reliability. Supporting engineering, quality, warranty, and software teams with production analytics.
DevOps Specialist scaling AWS infrastructure and automating deployments for Portage CyberTech’s digital trust, identity, privacy, and security solutions. Maintaining reliable cloud environments and CI/CD pipelines.
SRE Manager leading platform engineering teams across reliability and developer experience. Improving Lightspeed’s global cloud commerce SaaS platform through scalable infrastructure and automation.
SRE Manager leading platform engineering, reliability, and Developer Experience for Lightspeed’s global cloud POS SaaS platform. Guiding teams, infrastructure automation, and strategic delivery.
Gestionnaire SRE dirigeant deux équipes de plateforme chez Lightspeed. Assurant la fiabilité et l’automatisation de sa plateforme SaaS mondiale de commerce infonuagique.
DevOps Engineer managing AWS, Kubernetes, CI/CD, and observability for StellarTech’s scalable global EdTech products. Improving reliability, security, and production operations across multiple products.