Site Reliability Engineer embedding with teams for reliability design and leading initiatives. Building frameworks and measurement for operational excellence in software development.
Responsibilities
Embed with development teams from project inception for reliability design
Define production-readiness standards and launch criteria
Establish measurable SLIs, SLOs, and error budgets
Create reliable frameworks and tooling across cloud platforms
Lead incident management and postmortem practices
Build relationships with cloud providers to enhance capabilities
Requirements
Multi-Cloud Fluency across AWS, GCP, and Azure
Comfort supporting TypeScript and Ruby on Rails services
Significant experience as an SRE or platform engineer
Strong software engineering fundamentals
Technical Leadership & Influence without formal authority
Ability to drive high-scope problems to completion
Systems Thinking to identify process and technical debt
Data-Driven Leadership experience
Strong verbal and written English communication skills
Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.
Mozilla release engineer optimizing Firefox build, test, and deployment pipelines at global scale. Improving developer experience, maintaining automation, and responding to critical service outages without an on - call rotation.
DevOps Manager overseeing releases, enterprise tooling, and incident response for Delta Controls, a building - automation solutions manufacturer. Establishing standards across global product teams and offices.