Platform Engineer owning AWS, Kubernetes, observability, and Linux edge infrastructure. Improving reliability, automation, and developer workflows for Fleetworthy’s fleet-readiness technology.
Responsibilities
Own and evolve AWS infrastructure, including EKS, EC2, ALB/ELB, ACM, IAM, ECR, and Auto Scaling Groups, focusing on uptime, cost efficiency, and deployment safety
Maintain Kubernetes platform health, including deployments, ingress, HPA, secrets, Helm releases, and production incident response
Partner with development teams to improve deployment pipelines, reduce manual steps, and increase reliability
Build and maintain dashboards, alerts, and telemetry pipelines using Grafana, Prometheus/Mimir, Loki, Grafana Alloy, Tempo, OpenTelemetry, Datadog, and CloudWatch
Create actionable metrics, log views, and traces for engineering and operations teams
Write PromQL, LogQL, SQL, and CloudWatch queries and improve alert quality
Support distributed Linux-based field hardware, embedded devices, vendor integrations, and cloud data-forwarding services
Troubleshoot connectivity, routing, ARP, DNS, firewall rules, VPN behavior, and TCP socket data flows
Develop and maintain runbooks, automation, and configuration management practices for field operations
Own and improve Terraform and Terraform Cloud codebases, Ansible playbooks, Azure DevOps and GitLab CI/CD pipelines, and shell/Python automation
Reduce configuration drift, manual toil, and undocumented procedures
Harden and document Linux systems across cloud and edge environments
Build supported platform workflows that help product teams provision services, ship code, and operate software in production
Gather engineering feedback, prioritize platform improvements, and measure adoption and satisfaction
Partner with security on secrets management, policy-as-code, supply chain security, and least-privilege defaults
Requirements
5+ years of experience in platform engineering, site reliability, DevOps, or infrastructure roles working with similar technologies at production scale
Strong Linux fundamentals across Ubuntu and Debian environments, including systemd, nmcli, networking, package management, and embedded or edge hardware
Proven Kubernetes experience, including troubleshooting pods, inspecting events, tuning resource scheduling, and tracing requests through layered routing
Hands-on Terraform, Ansible, and CI/CD experience
Observability fluency with PromQL, LogQL, Databricks SQL, and CloudWatch queries
Solid networking foundation covering DNS, TLS, TCP, ARP, firewalls, VPN behavior, and load balancers in cloud and physical environments
Ability to work with legacy systems, vendor constraints, and incomplete documentation
Ability to write clear runbooks and explain complex systems to non-infrastructure teammates
Staff Platform Engineer defining cloud and DevSecOps strategy for Robots & Pencils’ enterprise AI systems. Leading Kubernetes, AI/ML infrastructure, migrations, reliability, security, and platform standards remotely in Canada.
Principal platform developer designing AWS - native integrations, CI/CD pipelines, and developer tooling for Autodesk’s design software. Leading architecture, reliability, and cross - team engineering initiatives.
Senior MLOps Developer operationalizing machine learning models and scalable AI/ML infrastructure for Autodesk’s design and entertainment software. Building deployment, monitoring, governance, and recovery systems.
Senior SRE/Platform Engineer needed for global company. 6+ years SRE experience, AWS/Azure, Kubernetes, Terraform, observability tools. Contract - to - hire in Mississauga.
Senior Full Stack Engineer building GraphQL, React, and TypeScript platforms for PENN Entertainment’s online gaming and sports media products. Improving shared client tooling, server - driven UI, performance, observability, and release workflows.
Senior platform engineering lead shaping compute and virtualization strategy for BMO, a major bank. Driving modernization, architecture standards, automation and hybrid - cloud infrastructure transformation.
Software Engineer building AI platform backend systems, integrations, and retrieval infrastructure for DraftKings’ digital sports entertainment and gaming products. Developing scalable services, automation, and LLM - powered applications.
Staff Platform Software Engineer building reusable infrastructure for Cantina Labs’ social AI platform. Improving search, AWS operations, Go services, observability, and developer productivity.