DevOps Engineer responsible for maintaining corporate IT systems and cloud infrastructure. Collaborating with business teams to deliver technology-driven solutions.
Responsibilities
Own and maintain corporate IT infrastructure using Terraform, ensuring configurations are versioned, auditable, and secure.
Design, build, and deploy automations using serverless automations in the cloud to streamline operational workflows and reduce manual effort.
Own alerting and notification pipelines using platforms such as incident.io and other incident management tools, ensuring anomalies and critical events surface to the appropriate responders.
Participate in and improve incident response workflows, including maintaining and iterating on runbooks, conducting post-incident reviews, and driving down mean time to resolution.
Package, deploy, and maintain internal tooling using Docker to support IT operations and automation efforts.
Develop targeted scripts and lightweight applications in Bash, Python, and JavaScript/TypeScript to solve operational problems and integrate corporate systems.
Collaborate cross-functionally with Security, Platform/SRE Engineering, and business stakeholders to align IT initiatives with organizational needs.
Maintain and troubleshoot network infrastructure fundamentals, including DNS, VPN, and firewall configurations.
Requirements
5 years of professional experience in IT engineering, systems administration, or a DevOps-adjacent discipline.
Administer Google Workspace at an organizational level, including user lifecycle management, security policies, group management, and audit log review.
Demonstrated experience with Terraform for infrastructure-as-code, specifically managing cloud resources.
Hands-on experience with serverless and managed compute services
Experience building and consuming REST APIs and webhook-based integrations between corporate systems.
Working proficiency in scripting using Bash and Python.
Practical experience with k8s for containerizing and deploying internal tools and services.
Familiarity with monitoring and observability platforms such as Datadog, Grafana, or equivalent.
Familiarity with alerting and incident management platforms (e.g., incident.io, PagerDuty, or equivalent) and the ability to configure, tune, and maintain notification pipelines.
Solid understanding of networking fundamentals, including DNS, VPN, and firewall technologies.
Curiosity, initiative and willingness to learn.
Nice to Have (But Not Required)
Experience working in regulated industries (e.g., financial services, healthcare)
Familiarity with compliance standards (e.g., NIST, ISO 27001, CIS)
Knowledge of additional languages like JavaScript or TypeScript.
Experience with CI/CD pipelines and tooling such as GitHub Actions, Cloud Build, or similar.
Direct participation in incident response processes, including triage, escalation, remediation, and post-incident review.
Familiarity with ITIL or ITSM frameworks and their application to service delivery and incident management.
Experience managing device fleet using MDM and endpoint tooling, ensuring devices are compliant, patched and properly configured.
Strong documentation skills, including the ability to author clear, actionable runbooks, standard operating procedures, and technical reference material.
DevOps Intern supporting CI/CD, cloud infrastructure, and automation for Ludia’s mobile game studio. Improving reliability and developer tools in production game environments.
Manager leading global SRE teams for Akamai's distributed Cloud IAM services. Improving reliability, scalability, security, and usability through cloud - native tooling and software.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.