Release Engineer at Supabase, ensuring safe and observable deployments and operational reliability across systems. Engage in incident management, monitoring, and process documentation for improved deployment efficiency.
Responsibilities
Own the reliability of Supabase's deployment and release systems, and the control plane they run on, against clear SLOs and error budgets
Turn pre-production into a trustworthy signal — standardizing and instrumenting today's fragmented, ad-hoc deployment workflows
Drive disaster-recovery readiness, including making environments reproducibly deployable from scratch (untangling undocumented secrets, unclear configuration ownership, and circular service dependencies)
Build and operate health and SLO monitoring for critical user flows, using synthetic testing to catch regressions before customers do
Reduce mean-time-to-detect and mean-time-to-recover for deploy-related incidents — which account for a large share of our incident load
Participate in on-call, lead blameless postmortems, and turn findings into runbooks, alerting, and automation that remove toil
Improve deployment observability and auditability — a clear record of what shipped where, when, and by whom
Document operational procedures — break-glass paths, access models, and runbooks — so reliability knowledge isn't tribal
Define and track SLAs, SLOs, error budgets, and DORA delivery metrics — with meaningful alerting over noise
Ensure deployments fail fast and safely when health checks degrade
Harden access and break-glass workflows (e.g. scoped self-service) so the right people can act in an incident without unsafe workarounds
Partner with product engineering and platform teams to align release practices with reliability and availability targets
Requirements
Have 5+ years in SRE, production operations, platform engineering, or release engineering
Have operated production systems at scale and carried on-call for them
Are fluent in SLAs, SLOs, error budgets, DORA metrics, and operational KPIs — and the observability tooling behind them (Prometheus, Grafana, Alertmanager, or similar)
Have led incident response with tooling like incident.io (or PagerDuty / Opsgenie), run blameless postmortems, and driven down MTTD/MTTR
Operate confidently on AWS (multiple accounts, IAM, VPC) in production
Are comfortable with infrastructure-as-code (Pulumi, Terraform) and Kubernetes
Script and automate to eliminate toil rather than absorb it
Communicate clearly with both infrastructure specialists and product engineers
Thrive in async, globally distributed teams
Are comfortable navigating ambiguity and iterating toward better systems over time.
DevOps Manager overseeing releases, enterprise tooling, and incident response for Delta Controls, a building - automation solutions manufacturer. Establishing standards across global product teams and offices.
DevOps Engineer building AWS infrastructure and automated systems for S&P Global’s financial data and technology solutions. Supporting resilient applications through Terraform, CI/CD, containerization, monitoring, and cloud operations.
Staff SRE leading GCP reliability, observability, and infrastructure automation for Calix’s broadband communications platform. Building resilient GKE, Kafka, database, and networking systems.