Production Support Engineer / SRE role supporting critical digital applications with SRE practices. Requires 5+ years experience with Ansible, Elasticsearch, MongoDB, Redis, OpenShift, Azure, and Linux/Windows administration.
Responsibilities
This role blends hands-on production support with Site Reliability Engineering (SRE) practices, focusing on automation, infrastructure reliability, and high availability. Key responsibilities include infrastructure and toil reduction (SSL/TLS certificate renewals, server patching, credential rotation, vulnerability remediation), database and cluster management (Elasticsearch, MongoDB, Redis clusters, performance tuning, disaster recovery), automation and SRE (Ansible playbooks, Infrastructure-as-Code, runbooks, monitoring/alerting), and production support (troubleshooting incidents, RCA/postmortems, performance monitoring, team collaboration).
Requirements
Must-Have Skills: 5+ years in Production Support / SRE, strong experience with Ansible (2+ years), Elasticsearch & MongoDB (2+ years), Redis, OpenShift, Azure, Linux (RHEL/CentOS/Ubuntu) & Windows Server Admin, Shell scripting (Bash), incident management & troubleshooting. Nice to Have: Kubernetes/container platforms, Python or Go scripting, CI/CD tools (Jenkins, GitLab, Azure DevOps), monitoring tools (Prometheus, Grafana, ELK, Datadog), Terraform/CloudFormation, certifications (AZ-900, CKA, Elasticsearch). Key Requirements: Experience in financial services (preferred), strong analytical & problem-solving skills, ability to work independently & in teams, excellent documentation & communication.
We're hiring a Database Reliability Engineer in Brampton, ON (hybrid) to manage SQL Server and Azure SQL, improve performance, and automate deployments.
Seeking an experienced Production Support Specialist for L3 support, incident management, and platform stability on a large - scale enterprise risk platform.
ML Ops/Production Engineer focused on scientific computing and data processing AI systems optimization. Innovating frameworks for fast experimentation, training/validating models, and managing production systems.
Senior Site Reliability Engineer supporting the deployment and operations of cloud - native applications on Kubernetes. Focused on maintaining reliability and resolving production incidents for a leading enterprise software company.
Project Engineer coordinating activities within the Production Technical Department at a pharmaceutical company. Leading project design and supporting development and troubleshooting processes throughout project lifecycle.
Staff Software Developer enhancing platform reliability at Wealthsimple. Collaborating across engineering teams to prevent incidents and improve operational standards.
Aarorn Technologies seeks a Production Support pro (6+ yrs) for a hybrid Toronto role, troubleshooting production issues and finding long - term solutions.
Hiring Production Support Engineer for a 12 - month hybrid contract in Toronto. Requires 7+ years of experience with Java Spring Boot, ReactJS, and incident management.