Platform Engineer owning AWS, Kubernetes, observability, and Linux edge infrastructure. Improving reliability, automation, and developer workflows for Fleetworthy’s fleet-readiness technology.
Responsibilities
Own and evolve AWS infrastructure, including EKS, EC2, ALB/ELB, ACM, IAM, ECR, and Auto Scaling Groups, focusing on uptime, cost efficiency, and deployment safety
Maintain Kubernetes platform health, including deployments, ingress, HPA, secrets, Helm releases, and production incident response
Partner with development teams to improve deployment pipelines, reduce manual steps, and increase reliability
Build and maintain dashboards, alerts, and telemetry pipelines using Grafana, Prometheus/Mimir, Loki, Grafana Alloy, Tempo, OpenTelemetry, Datadog, and CloudWatch
Create actionable metrics, log views, and traces for engineering and operations teams
Write PromQL, LogQL, SQL, and CloudWatch queries and improve alert quality
Support distributed Linux-based field hardware, embedded devices, vendor integrations, and cloud data-forwarding services
Troubleshoot connectivity, routing, ARP, DNS, firewall rules, VPN behavior, and TCP socket data flows
Develop and maintain runbooks, automation, and configuration management practices for field operations
Own and improve Terraform and Terraform Cloud codebases, Ansible playbooks, Azure DevOps and GitLab CI/CD pipelines, and shell/Python automation
Reduce configuration drift, manual toil, and undocumented procedures
Harden and document Linux systems across cloud and edge environments
Build supported platform workflows that help product teams provision services, ship code, and operate software in production
Gather engineering feedback, prioritize platform improvements, and measure adoption and satisfaction
Partner with security on secrets management, policy-as-code, supply chain security, and least-privilege defaults
Requirements
5+ years of experience in platform engineering, site reliability, DevOps, or infrastructure roles working with similar technologies at production scale
Strong Linux fundamentals across Ubuntu and Debian environments, including systemd, nmcli, networking, package management, and embedded or edge hardware
Proven Kubernetes experience, including troubleshooting pods, inspecting events, tuning resource scheduling, and tracing requests through layered routing
Hands-on Terraform, Ansible, and CI/CD experience
Observability fluency with PromQL, LogQL, Databricks SQL, and CloudWatch queries
Solid networking foundation covering DNS, TLS, TCP, ARP, firewalls, VPN behavior, and load balancers in cloud and physical environments
Ability to work with legacy systems, vendor constraints, and incomplete documentation
Ability to write clear runbooks and explain complex systems to non-infrastructure teammates
Senior Platform Engineer building scalable backend and cloud infrastructure for ExaCare’s AI - powered post - acute care platform. Improving reliability, developer velocity, and healthcare admission workflows.
Principal Platform Engineer leading SkyWatch’s satellite - data platform and AI agent infrastructure. Owning architecture, customer - driven roadmap delivery, and platform engineering leadership.
Platform Engineer securing Just Eat Takeaway.com’s global food - delivery edge infrastructure. Building gateways, automation, and resilient traffic routing across production environments.
Director leading Blackpoint Cyber’s cloud - based Unified Security Posture data platform for cybersecurity solutions. Driving platform roadmap, reliability, APIs, data engineering, and team growth.
Ingénieur logiciel principal intégrant des plateformes, API et solutions IA chez EDC. Gouvernance technique, sécurité, résilience et mentorat dans une société canadienne de financement du commerce.
Senior AI Platform Developer building scalable AI services and LLM workflows for MaintainX’s industrial work execution platform. Improving reliability, observability, performance, and cost efficiency.
Infrastructure team lead building and operating Spare’s GCP and Kubernetes platform for on - demand transit. Leading developers while improving reliability, security, AI SRE, and cloud cost efficiency.
AI Platform Developer building Azure - based agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Creating secure, observable, governed AI services for product teams.
Senior AI Platform Developer building reusable, secure AI - agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Driving Azure - based platform capabilities across orchestration, evaluation, observability, and governance.