Production Support Engineer / SRE role supporting critical digital applications with SRE practices. Requires 5+ years experience with Ansible, Elasticsearch, MongoDB, Redis, OpenShift, Azure, and Linux/Windows administration.
Responsibilities
This role blends hands-on production support with Site Reliability Engineering (SRE) practices, focusing on automation, infrastructure reliability, and high availability. Key responsibilities include infrastructure and toil reduction (SSL/TLS certificate renewals, server patching, credential rotation, vulnerability remediation), database and cluster management (Elasticsearch, MongoDB, Redis clusters, performance tuning, disaster recovery), automation and SRE (Ansible playbooks, Infrastructure-as-Code, runbooks, monitoring/alerting), and production support (troubleshooting incidents, RCA/postmortems, performance monitoring, team collaboration).
Requirements
Must-Have Skills: 5+ years in Production Support / SRE, strong experience with Ansible (2+ years), Elasticsearch & MongoDB (2+ years), Redis, OpenShift, Azure, Linux (RHEL/CentOS/Ubuntu) & Windows Server Admin, Shell scripting (Bash), incident management & troubleshooting. Nice to Have: Kubernetes/container platforms, Python or Go scripting, CI/CD tools (Jenkins, GitLab, Azure DevOps), monitoring tools (Prometheus, Grafana, ELK, Datadog), Terraform/CloudFormation, certifications (AZ-900, CKA, Elasticsearch). Key Requirements: Experience in financial services (preferred), strong analytical & problem-solving skills, ability to work independently & in teams, excellent documentation & communication.
Technicien en ingénierie de production chez Gentec, mettant en production des produits et optimisant les méthodes manufacturières. Gestion des nomenclatures, gammes de fabrication et projets d’amélioration.
Lean manufacturing specialist leading aerospace production - system improvements at Schaeffler Canada. Teaching lean tools, interpreting KPIs, and facilitating regional initiatives across manufacturing operations.
We're hiring a Database Reliability Engineer in Brampton, ON (hybrid) to manage SQL Server and Azure SQL, improve performance, and automate deployments.
Seeking an experienced Production Support Specialist for L3 support, incident management, and platform stability on a large - scale enterprise risk platform.
ML Ops/Production Engineer focused on scientific computing and data processing AI systems optimization. Innovating frameworks for fast experimentation, training/validating models, and managing production systems.
Senior Site Reliability Engineer supporting the deployment and operations of cloud - native applications on Kubernetes. Focused on maintaining reliability and resolving production incidents for a leading enterprise software company.
Project Engineer coordinating activities within the Production Technical Department at a pharmaceutical company. Leading project design and supporting development and troubleshooting processes throughout project lifecycle.
Staff Software Developer enhancing platform reliability at Wealthsimple. Collaborating across engineering teams to prevent incidents and improve operational standards.