Site Reliability Engineer ensuring Emburse’s systems are highly available, scalable, and performant. Collaborating across teams to drive automation and operational excellence in cloud infrastructure.
Responsibilities
Develop and maintain infrastructure as code using Terraform, OpenTofu, Ansible, and related automation tooling.
Administer Kubernetes and EKS environments, including installation, networking, security, troubleshooting, monitoring, autoscaling, upgrades, and cluster management.
Build, manage, and support containerized workloads using Kubernetes, EKS, and related cloud-native technologies.
Support GitOps-based deployment workflows using ArgoCD, Kustomize, GitHub Actions, Jenkins, and Kubernetes manifests.
Build and improve self-service platform capabilities that help engineering teams provision infrastructure, onboard services, and deploy applications efficiently.
Monitor site availability, investigate production issues, and provide remediation for incidents.
Create and maintain monitoring, logging, alerting, dashboards, APM configuration, and incident reporting.
Support secure-by-default platform practices, including least-privilege Kubernetes configurations, container vulnerability scanning, static analysis, and infrastructure-as-code security checks.
Troubleshoot infrastructure, application platform, Linux, networking, IAM, and cloud-related issues.
Write SQL, ELK, and other operational queries to diagnose issues and support production investigations.
Create, review, merge, and apply infrastructure pull requests using tools such as Ansible, Terraform, and OpenTofu.
Create and maintain AMIs, support rightsizing, autoscaling, and infrastructure optimization.
Serve as a technical lead on complex platform projects, driving work to successful completion on time and on budget.
Leverage AI-assisted engineering tools, such as Cursor, GitHub Copilot, Claude Code, or similar technologies, to improve development speed, code quality, automation, documentation, and troubleshooting workflows.
Requirements
Experience with infrastructure as code and the full lifecycle of SaaS implementations.
Strong Kubernetes administration experience, including networking, security, troubleshooting, monitoring, and day-2 operations.
Experience with containers, EKS, Kubernetes, GitHub Actions, Jenkins, ArgoCD, Kustomize, and cloud-native deployment practices.
AWS proficiency, including basic IAM management, autoscaling, AMIs, and cloud infrastructure operations.
Intermediate to advanced Linux and Unix skills.
Understanding of TCP/IP, OSI model, stateless architecture, infrastructure, and system architecture.
Ability to write SQL and ELK queries.
Experience with monitoring applications, APM tools, logs, alerts, and incident diagnostics.
Experience with secure delivery practices, including vulnerability scanning, static analysis, and least-privilege infrastructure patterns.
Ability to effectively use AI-assisted engineering tools, such as Cursor, GitHub Copilot, Claude Code, or similar technologies, while applying sound engineering judgment, code review practices, and security awareness.
Ability to merge and apply pull requests for Ansible, Terraform, OpenTofu, or similar infrastructure tooling.
Deep understanding of release cycles, SDLC, infrastructure, and architecture.
Strong analytical, reasoning, troubleshooting, and problem-solving skills.
Excellent written and verbal communication skills in English.
Strong listening, teamwork, time management, and attention to detail.
Minimum of 3 years of direct experience in a similar role with a Bachelor’s degree in Computer Science or related STEM field.
Minimum of 7 years of direct experience in a similar role without a Bachelor’s degree.
Senior Platform Engineer building scalable backend and cloud infrastructure for ExaCare’s AI - powered post - acute care platform. Improving reliability, developer velocity, and healthcare admission workflows.
Principal Platform Engineer leading SkyWatch’s satellite - data platform and AI agent infrastructure. Owning architecture, customer - driven roadmap delivery, and platform engineering leadership.
Platform Engineer securing Just Eat Takeaway.com’s global food - delivery edge infrastructure. Building gateways, automation, and resilient traffic routing across production environments.
Director leading Blackpoint Cyber’s cloud - based Unified Security Posture data platform for cybersecurity solutions. Driving platform roadmap, reliability, APIs, data engineering, and team growth.
Ingénieur logiciel principal intégrant des plateformes, API et solutions IA chez EDC. Gouvernance technique, sécurité, résilience et mentorat dans une société canadienne de financement du commerce.
Senior AI Platform Developer building scalable AI services and LLM workflows for MaintainX’s industrial work execution platform. Improving reliability, observability, performance, and cost efficiency.
Infrastructure team lead building and operating Spare’s GCP and Kubernetes platform for on - demand transit. Leading developers while improving reliability, security, AI SRE, and cloud cost efficiency.
AI Platform Developer building Azure - based agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Creating secure, observable, governed AI services for product teams.
Senior AI Platform Developer building reusable, secure AI - agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Driving Azure - based platform capabilities across orchestration, evaluation, observability, and governance.