Site Reliability Engineer ensuring Emburse’s systems are highly available, scalable, and performant. Collaborating across teams to drive automation and operational excellence in cloud infrastructure.
Responsibilities
Develop and maintain infrastructure as code using Terraform, OpenTofu, Ansible, and related automation tooling.
Administer Kubernetes and EKS environments, including installation, networking, security, troubleshooting, monitoring, autoscaling, upgrades, and cluster management.
Build, manage, and support containerized workloads using Kubernetes, EKS, and related cloud-native technologies.
Support GitOps-based deployment workflows using ArgoCD, Kustomize, GitHub Actions, Jenkins, and Kubernetes manifests.
Build and improve self-service platform capabilities that help engineering teams provision infrastructure, onboard services, and deploy applications efficiently.
Monitor site availability, investigate production issues, and provide remediation for incidents.
Create and maintain monitoring, logging, alerting, dashboards, APM configuration, and incident reporting.
Support secure-by-default platform practices, including least-privilege Kubernetes configurations, container vulnerability scanning, static analysis, and infrastructure-as-code security checks.
Troubleshoot infrastructure, application platform, Linux, networking, IAM, and cloud-related issues.
Write SQL, ELK, and other operational queries to diagnose issues and support production investigations.
Create, review, merge, and apply infrastructure pull requests using tools such as Ansible, Terraform, and OpenTofu.
Create and maintain AMIs, support rightsizing, autoscaling, and infrastructure optimization.
Serve as a technical lead on complex platform projects, driving work to successful completion on time and on budget.
Leverage AI-assisted engineering tools, such as Cursor, GitHub Copilot, Claude Code, or similar technologies, to improve development speed, code quality, automation, documentation, and troubleshooting workflows.
Requirements
Experience with infrastructure as code and the full lifecycle of SaaS implementations.
Strong Kubernetes administration experience, including networking, security, troubleshooting, monitoring, and day-2 operations.
Experience with containers, EKS, Kubernetes, GitHub Actions, Jenkins, ArgoCD, Kustomize, and cloud-native deployment practices.
AWS proficiency, including basic IAM management, autoscaling, AMIs, and cloud infrastructure operations.
Intermediate to advanced Linux and Unix skills.
Understanding of TCP/IP, OSI model, stateless architecture, infrastructure, and system architecture.
Ability to write SQL and ELK queries.
Experience with monitoring applications, APM tools, logs, alerts, and incident diagnostics.
Experience with secure delivery practices, including vulnerability scanning, static analysis, and least-privilege infrastructure patterns.
Ability to effectively use AI-assisted engineering tools, such as Cursor, GitHub Copilot, Claude Code, or similar technologies, while applying sound engineering judgment, code review practices, and security awareness.
Ability to merge and apply pull requests for Ansible, Terraform, OpenTofu, or similar infrastructure tooling.
Deep understanding of release cycles, SDLC, infrastructure, and architecture.
Strong analytical, reasoning, troubleshooting, and problem-solving skills.
Excellent written and verbal communication skills in English.
Strong listening, teamwork, time management, and attention to detail.
Minimum of 3 years of direct experience in a similar role with a Bachelor’s degree in Computer Science or related STEM field.
Minimum of 7 years of direct experience in a similar role without a Bachelor’s degree.
Staff Platform Engineer defining cloud and DevSecOps strategy for Robots & Pencils’ enterprise AI systems. Leading Kubernetes, AI/ML infrastructure, migrations, reliability, security, and platform standards remotely in Canada.
Principal platform developer designing AWS - native integrations, CI/CD pipelines, and developer tooling for Autodesk’s design software. Leading architecture, reliability, and cross - team engineering initiatives.
Senior MLOps Developer operationalizing machine learning models and scalable AI/ML infrastructure for Autodesk’s design and entertainment software. Building deployment, monitoring, governance, and recovery systems.
Senior SRE/Platform Engineer needed for global company. 6+ years SRE experience, AWS/Azure, Kubernetes, Terraform, observability tools. Contract - to - hire in Mississauga.
Senior Full Stack Engineer building GraphQL, React, and TypeScript platforms for PENN Entertainment’s online gaming and sports media products. Improving shared client tooling, server - driven UI, performance, observability, and release workflows.
Senior platform engineering lead shaping compute and virtualization strategy for BMO, a major bank. Driving modernization, architecture standards, automation and hybrid - cloud infrastructure transformation.
Software Engineer building backend AI platform systems for DraftKings’ sports entertainment and gaming technology. Developing retrieval, vector database, automation, and agent infrastructure for scalable AI applications.
Software Engineer building AI platform backend systems, integrations, and retrieval infrastructure for DraftKings’ digital sports entertainment and gaming products. Developing scalable services, automation, and LLM - powered applications.