Site Reliability Engineer maintaining Kubernetes UI services for an enterprise cloud software client.
Monitoring production health, troubleshooting incidents, and supporting reliable cloud-native operations.
Responsibilities
Support the deployment, operations, and ongoing maintenance of production services running on Kubernetes
Monitor service health, availability, and performance
Investigate and troubleshoot production incidents using logs, monitoring, and debugging tools
Perform log analysis and incident debugging using Splunk
Identify service issues and collaborate with engineering teams to support timely resolution
Participate in incident response and production support activities
Perform first-level debugging of UI-related issues involving Web Components
Support service reliability and continuous improvement initiatives
Assist with CI/CD pipelines and cloud-native application operations when needed
Work effectively within a client-directed backlog and established priorities
Requirements
4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Support, or a related role
Hands-on experience supporting deployment, operations, and ongoing maintenance of production services running on Kubernetes
Experience monitoring service health, troubleshooting production issues, and supporting service reliability
Proficiency with Splunk for log analysis and incident debugging
Experience participating in production incident response and root-cause analysis
Working knowledge of Web Components and ability to perform first-level debugging of UI-related issues
Strong troubleshooting, analytical, and problem-solving skills
Experience collaborating with software engineering and cross-functional teams
Ability to work independently and effectively within a client-directed backlog
Excellent written and spoken English, at least B2 level
Experience supporting CI/CD pipelines
Familiarity with multi-tenant services
Experience with cloud-native application operations
Experience supporting high-availability enterprise or SaaS platforms
Familiarity with additional monitoring and observability tools
Experience with cloud platforms such as AWS, Azure, or GCP
Familiarity with container and deployment technologies such as Docker and Helm
Benefits
Competitive salary
Laptop
Professional development and training opportunities
Work with cutting-edge cloud and container technologies
Flexible work arrangements and collaborative team environment
Impact on organization-wide digital transformation initiatives
DevOps Manager overseeing releases, enterprise tooling, and incident response for Delta Controls, a building - automation solutions manufacturer. Establishing standards across global product teams and offices.
DevOps Engineer building AWS infrastructure and automated systems for S&P Global’s financial data and technology solutions. Supporting resilient applications through Terraform, CI/CD, containerization, monitoring, and cloud operations.
Staff SRE leading GCP reliability, observability, and infrastructure automation for Calix’s broadband communications platform. Building resilient GKE, Kafka, database, and networking systems.