Senior Site Reliability Engineer supporting the deployment and operations of cloud-native applications on Kubernetes. Focused on maintaining reliability and resolving production incidents for a leading enterprise software company.
Responsibilities
Support deployment, operations, and ongoing maintenance of a production service running on Kubernetes.
Monitor application health, availability, and performance.
Investigate and resolve production incidents using logs, monitoring, and debugging tools.
Perform log analysis using Splunk to identify root causes and troubleshoot service issues.
Collaborate with software engineers to improve service reliability and operational efficiency.
Participate in incident response and production support activities.
Assist with first-level debugging of UI-related issues involving Web Components.
Contribute to continuous improvements in automation, monitoring, and operational processes.
Support CI/CD pipelines and cloud-native deployment practices.
Requirements
5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Production Operations.
Strong hands-on experience with Kubernetes in production environments.
Experience supporting cloud-native applications.
Experience monitoring production systems and troubleshooting complex incidents.
Strong knowledge of Splunk for log analysis and debugging.
Experience working in Linux environments.
Understanding of networking fundamentals and distributed systems.
Strong troubleshooting and root cause analysis skills.
Excellent written and spoken English (B2+).
Benefits
Competitive salary and laptop
Professional development and training opportunities
Work with cutting-edge cloud and container technologies
Flexible work arrangements and collaborative team environment
Impact on organization-wide digital transformation initiatives
We're hiring a Database Reliability Engineer in Brampton, ON (hybrid) to manage SQL Server and Azure SQL, improve performance, and automate deployments.
Seeking an experienced Production Support Specialist for L3 support, incident management, and platform stability on a large - scale enterprise risk platform.
ML Ops/Production Engineer focused on scientific computing and data processing AI systems optimization. Innovating frameworks for fast experimentation, training/validating models, and managing production systems.
Project Engineer coordinating activities within the Production Technical Department at a pharmaceutical company. Leading project design and supporting development and troubleshooting processes throughout project lifecycle.
Staff Software Developer enhancing platform reliability at Wealthsimple. Collaborating across engineering teams to prevent incidents and improve operational standards.
Aarorn Technologies seeks a Production Support pro (6+ yrs) for a hybrid Toronto role, troubleshooting production issues and finding long - term solutions.
Hiring Production Support Engineer for a 12 - month hybrid contract in Toronto. Requires 7+ years of experience with Java Spring Boot, ReactJS, and incident management.