Site Reliability Engineer overseeing cloud solutions for Akur8, focusing on reliability and automation. Collaborating across teams to enhance cloud infrastructure and maintain SLOs.
Responsibilities
Maintain and improve our infrastructure-as-code repositories (Terraform) to ensure the reliability and resilience of Akur8's cloud products.
Contribute to expanding Akur8's product offerings while maintaining our SLOs.
Strengthen automation and orchestration of pipelines to reduce repetitive manual tasks.
Train and support teams on DevOps best practices across the organization.
Contribute to the design of our AWS and Azure platform architectures, in collaboration with product and development teams, to improve performance, reliability, cost control, and to support new product features.
Help continuously improve monitoring and observability, primarily using Datadog.
Contribute to our CI pipelines (GitHub Actions), ensuring best practices are consistently applied when using containers (Docker).
Work closely with our Security team to secure workloads, maintain IT security standards and best practices, and participate in implementing infrastructure scanning.
Contribute to open-source projects where appropriate.
Actively participate in the on-call rotation (1 week every 4 to 6 weeks).
Requirements
Degree in Computer Science, Information Technology, or a related field, or equivalent experience.
At least 5 years of professional experience configuring, monitoring, and maintaining AWS and/or Azure production systems across the full software development lifecycle.
Strong hands-on experience with Terraform in AWS and/or Azure environments.
DevOps Manager overseeing releases, enterprise tooling, and incident response for Delta Controls, a building - automation solutions manufacturer. Establishing standards across global product teams and offices.
DevOps Engineer building AWS infrastructure and automated systems for S&P Global’s financial data and technology solutions. Supporting resilient applications through Terraform, CI/CD, containerization, monitoring, and cloud operations.
Staff SRE leading GCP reliability, observability, and infrastructure automation for Calix’s broadband communications platform. Building resilient GKE, Kafka, database, and networking systems.
Senior Azure DevOps advisor governing platform evolution for Alithya, a digital transformation consulting firm. Defining standards, optimizing pipelines, dashboards, integrations, and AI capabilities.
Senior DevOps Engineer building Azure DevOps pipelines, Terraform infrastructure, and deployment automation. Supporting PLATO, a Canadian Indigenous - owned software testing and technology services company, across product and data teams.
Pilote Azure DevOps pour un organisme public de santé québécois en transformation numérique. Gouvernance, intégrations Power BI/Dynamics 365, administration de plateforme et accompagnement Agile.