Site Reliability Engineer ensuring high availability, scalability, and performance of Emburse’s systems. Collaborating on distributed systems while mentoring junior engineers.
Responsibilities
Proactively identify, evaluate, and implement preventative measures to reduce customer impact.
Ensure all services are designed and operated with 24/7 availability, scalability, and resilience in mind.
Monitor, troubleshoot, and provide visibility to improve site latency, performance, and uptime.
Design, develop, and automate reliable cloud infrastructure and platform services.
Apply Infrastructure-as-Code (IaC) principles to manage large-scale distributed systems.
Write and maintain scripts, tools, and automation frameworks to support operational efficiency.
Partner with engineering leadership to develop solutions enabling developer productivity and remove cross functional dependencies.
Collaborate with Platform Engineering teams on project definitions, requirements, backlog grooming, and planning processes.
Align operational goals with product and engineering roadmaps to ensure reliability requirements are met early in the lifecycle.
Define non-functional requirements (NFRs) and influence standards for scalability, observability, and fault tolerance.
Lead cross-functional troubleshooting of complex issues spanning applications, infrastructure, databases, and networks.
Serve as a technical mentor to SRE I and II engineers, guiding them in best practices for reliability, automation, and incident management.
Lead root cause analysis and postmortem reviews, driving continuous improvement initiatives.
Support offshore and distributed teams, promoting effective collaboration and communication.
Participate in design and architecture reviews, offering technical recommendations and documentation for key stakeholders
Requirements
Bachelor’s degree in Computer Science or a STEM field
Minimum 6 years of experience in an engineering or operations role with a focus on reliability, scalability, and automation.
Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.
Mozilla release engineer optimizing Firefox build, test, and deployment pipelines at global scale. Improving developer experience, maintaining automation, and responding to critical service outages without an on - call rotation.
DevOps Manager overseeing releases, enterprise tooling, and incident response for Delta Controls, a building - automation solutions manufacturer. Establishing standards across global product teams and offices.