Senior Data Platform Engineer developing infrastructure for data processing at Tubi in Toronto. Handling Spark-on-Kubernetes and real-time feature pipelines with responsibilities across data platform components.
Responsibilities
Spark-on-Kubernetes — EKS-based compute platform for Spark workloads: cluster configuration, Pod Identity IAM, job environment setup, Kustomize overlays, and shadow canary validation
Event ingestion — Rust services and Flink jobs processing billions of events per day over Kinesis; throughput, reliability, on-call response, and AI-assisted operational tooling to reduce toil
Platform infrastructure — Terraform modules for environment provisioning, cross-account AWS IAM, ARC runner infrastructure, and CI/CD for data platform changes
Feature store and ML compute — Flink-based real-time feature pipelines feeding a large-scale MemoryDB cluster; GPU capacity governance and Databricks multi-environment operations for ML training workloads
Workflow orchestration and CDC — Airflow-based DAG deployment, change data capture pipeline operations, and data quality monitoring
Requirements
3+ years building and operating production data platform infrastructure at the cluster or platform level, across Spark, Flink, Kinesis, Kubernetes, or equivalent
Deep experience in at least one of: Spark-on-K8s cluster operations, Rust-based data or systems engineering, Kubernetes platform engineering and IaC, or data catalog and governance tooling
Production AWS experience or equivalent: EKS, S3, Kinesis, and multi-account IAM patterns (EKS Pod Identity, KIAM, or IRSA)
You've owned a critical platform component, you wrote the runbooks, tracked the cost, and were on-call for it
Strong in at least one of: Rust, JVM (Java or Scala), or Python for data platform work.
Benefits
This role is also eligible for an annual discretionary bonus
long-term incentive plan
medical/dental/vision
insurance
vacation/paid time off
Flexible Time Off Policy to manage all personal matters
generous Parental Leave Program allowing parents twelve (12) weeks of paid bonding leave
Ingénieur logiciel principal intégrant des plateformes, API et solutions IA chez EDC. Gouvernance technique, sécurité, résilience et mentorat dans une société canadienne de financement du commerce.
Senior AI Platform Developer building scalable AI services and LLM workflows for MaintainX’s industrial work execution platform. Improving reliability, observability, performance, and cost efficiency.
Infrastructure team lead building and operating Spare’s GCP and Kubernetes platform for on - demand transit. Leading developers while improving reliability, security, AI SRE, and cloud cost efficiency.
AI Platform Developer building Azure - based agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Creating secure, observable, governed AI services for product teams.
Senior AI Platform Developer building reusable, secure AI - agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Driving Azure - based platform capabilities across orchestration, evaluation, observability, and governance.
Kubernetes/DevOps Engineer operating Kubernetes infrastructure for a petabyte - scale social media platform. Building distributed systems and machine - learning workloads for Capgemini Engineering’s client.
Lead AI Platform Engineer securing and operating EQ Bank’s enterprise AI platforms. Driving Azure engineering, observability, automation, governance, and production readiness.
Senior Platform Engineer building and operating GitLab Orbit, a Rust - based knowledge graph service for GitLab’s DevSecOps platform. Improving distributed - system reliability, observability, cloud infrastructure, and data workflows.