Senior Software Engineer improving reliability and security across Coinbase's services. Design and deliver projects for resilience and safe deployments within a fast-paced remote work environment.
Responsibilities
Own the design and delivery of reliability projects and features that improve resiliency across Coinbase's service environment in partnership with other engineering teams.
Partner with critical T0/T1 services to understand architecture, improve scalability, and reduce operational toil.
Build and enhance systems that securely manage service configurations and secrets at scale.
Improve canary-based release systems and expand deployment capabilities to support thousands of services and hundreds of daily deployments with fewer incidents.
Drive reliability best practices and strengthen reliability culture across engineering teams at Coinbase.
Requirements
5+ years of software engineering experience designing, building, and maintaining production services in service-oriented architectures, including experience with Ruby, Go, Terraform, and cloud platforms (AWS, GCP or Azure).
Demonstrated ability to design and operate reliable, high-throughput, low-latency distributed systems at scale, with a track record of writing well-tested, production-quality code.
Proven experience with observability and monitoring tools (e.g., Kibana, Datadog) to debug complex production issues, tune system performance, and reduce incident frequency.
Experience writing and verbally communicating architecture decisions to cross-functional engineering stakeholders.
Ability to participate in on-call rotations and respond to issues outside normal business hours.
Utilizes generative AI responsibly, maintaining human oversight to deliver business-ready outputs and drive measurable improvements in workflow efficiency, cost, and quality.
Benefits
Total compensation may also include equity and bonus eligibility and benefits (including medical, dental, and vision)
Senior DevOps / Cloud Infrastructure Engineer needed for hybrid role in North York, ON. Requires 10+ years experience with GCP, AWS, Kubernetes, Terraform, and CI/CD.
Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.
Mozilla release engineer optimizing Firefox build, test, and deployment pipelines at global scale. Improving developer experience, maintaining automation, and responding to critical service outages without an on - call rotation.