Senior Site Reliability Engineer ensuring platform reliability at Circle. Managing systems and database infrastructure to support high growth in user engagement and system performance.
Responsibilities
Act as a first responder for system incidents and outages, helping Circle stay highly available and performant
Own and evolve our monitoring, alerting, and log management systems
Manage and optimize our database infrastructure (including MySQL, Postgres, Clickhouse, and Redis)
Maintain and improve our server infrastructure and deployment pipelines
Collaborate closely with engineering teams to build scalable, resilient systems
Contribute to internal SRE tooling and automation efforts
Requirements
Strong alignment with our values (find our values on our career page if you haven’t read up on them yet)
You are proficient in English (spoken, written, and reading) at a CEFR Level C2 / ILR Level 5
Deep expertise with AWS and Kubernetes
5+ years of experience in a Site Reliability, DevOps, or Infrastructure Engineering role
Proven experience scaling production systems in a high-growth environment (startup or similar)
Practical, day-to-day experience using AI tools to improve engineering productivity and outcomes (e.g., copilots, LLM-based debugging, automation, or documentation workflows)
You’ve helped scale an early-stage product to 1M+ monthly active users
Experience managing incident response and production system outages
Hands-on experience with database operations and optimization
Familiarity with observability tooling, monitoring, and logging best practices
Based in North or South America (AMER region) — this is a requirement for timezone alignment with our team
Benefits
Fully remote: work from anywhere in the world!
Autonomy and trust to do your job: we care about outcomes over everything else.
Paid time away: all employees are given 35 days of PTO annually. We also offer a paid sabbatical after 5 years.
Generous U.S. benchmarked compensation and startup equity no matter where you are in the world.*
Awesome medical coverage with 100% coverage for you and your family, or medical reimbursement options where applicable!*
Parental leave for parents expanding their family, or just starting one.
Home office stipend to help you get up and running.
Learning & development stipend to help you level up your professional skills.
Annual bonus potential for roles that don't already receive variable income or commission.
Company retreats: Twice a year, the Circle team gets together for a fully paid company retreat in incredible places around the world! We’ve had past retreats in Colombia, Portugal, and Mexico, with more planned on the horizon.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.