Senior SRE strengthening PostgreSQL, Kubernetes, and observability for Alpaca’s global brokerage infrastructure. Operating production systems, improving reliability, and mentoring engineers across cloud and database operations.
Responsibilities
Operate production day-to-day, including on-call, incident response, postmortems, and follow-up actions
Define and refine SLIs/SLOs and error budgets, helping product teams operate within them
Strengthen observability across metrics, logs, traces, and alerting
Ship cloud resources and Kubernetes workloads through code in a GitOps workflow
Own PostgreSQL reliability through performance tuning, schema and migration review, online migrations on large tables, HA/DR, and CDC pipelines
Mentor engineers on reliability and database fundamentals through code review, design review, and pairing
Requirements
4+ years in SRE, DevOps, Platform/Infrastructure, or backend engineering with significant production operations ownership
Hands-on experience operating production services on Kubernetes and shipping infrastructure as code in a GitOps workflow
Solid working knowledge of PostgreSQL in production, including query plans, pg_stat_*, indexing, schema trade-offs, and safe online migrations
Practiced in incident response, structured debugging, and postmortems
At least working proficiency in Go or Python
Strong written and verbal communication
Genuine interest in databases and growing PostgreSQL/DBA expertise
Nice-to-haves include deeper PostgreSQL experience, typed SQL access layers in Go, large-scale messaging systems, security and compliance in regulated environments, and trading, brokerage, or regulated fintech familiarity
Benefits
Competitive Salary & Stock Options
Health Benefits
New Hire Home-Office Setup: One-time USD $500
Monthly Stipend: USD $150 per month via a Brex Card
DevOps Manager overseeing releases, enterprise tooling, and incident response for Delta Controls, a building - automation solutions manufacturer. Establishing standards across global product teams and offices.
DevOps Engineer building AWS infrastructure and automated systems for S&P Global’s financial data and technology solutions. Supporting resilient applications through Terraform, CI/CD, containerization, monitoring, and cloud operations.
Staff SRE leading GCP reliability, observability, and infrastructure automation for Calix’s broadband communications platform. Building resilient GKE, Kafka, database, and networking systems.
Senior Azure DevOps advisor governing platform evolution for Alithya, a digital transformation consulting firm. Defining standards, optimizing pipelines, dashboards, integrations, and AI capabilities.
Senior DevOps Engineer building Azure DevOps pipelines, Terraform infrastructure, and deployment automation. Supporting PLATO, a Canadian Indigenous - owned software testing and technology services company, across product and data teams.