Data Engineer building scalable pipelines for SumerSports’ football intelligence platform. Supporting deep learning, video, LLM, analytics, and AI-driven products across sports.
Responsibilities
Build and operate robust data pipelines for ingestion, cleaning, and transformation using Databricks, Airflow, or Kubernetes
Develop efficient ETL/ELT workflows in Python and SQL for batch and streaming workloads
Partner with ML/AI teams to make datasets and tools discoverable and safe for autonomous agents, including evaluation and guardrails for AI-generated queries
Develop retrieval pipelines (RAG, vector search) over structured statistics and unstructured sources such as scouting notes and video metadata
Model and maintain structured data assets in Delta, Parquet, and Iceberg for reliability, versioning, and lineage tracking
Implement orchestration and monitoring by scheduling jobs, tracking dependencies, and automating failure recovery
Ensure data quality and compliance through validation frameworks, schema enforcement, and audit logging
Contribute to data platform evolution by evaluating tools, standardizing best practices, and improving developer experience
Support performance and cost optimization across compute, storage, and orchestration systems
Collaborate with MLOps and Sports Data teams to integrate data and AI systems
Requirements
3–8 years of experience as a Data Engineer or ETL Developer in a production environment
Proficiency in Python and SQL
Strong familiarity with Databricks, Spark, or equivalent big-data frameworks
Experience with workflow orchestration tools such as Airflow, Dagster, Luigi, or Prefect
Deep understanding of data modeling, data warehousing, and distributed data processing
Knowledge of modern data lakehouse architectures
Familiarity with CI/CD, GitHub Actions, Infrastructure as Code, and data pipeline testing frameworks
Comfort working cross-functionally with ML, product, and analytics teams
Exposure to LLM-powered data tools, including text-to-SQL, RAG, agent/tool interfaces such as MCP, or natural-language analytics
Previous work with cloud infrastructure such as AWS, GCP, or Azure
Experience with container orchestration using Docker or Kubernetes
Preferred: previous experience with sports, telemetry, or sensor data pipelines
Preferred: familiarity with streaming frameworks and event-driven data processing such as Kafka, Spark Structured Streaming, or Flink
Preferred: general knowledge of American football, the NFL, and college football
Preferred: background in data governance, lineage, and observability tools such as Monte Carlo, Great Expectations, Unity Catalog, or OpenLineage
Preferred: experience designing semantic layers or metric definitions consumed by AI and BI tools
Preferred: exposure to machine-learning model management and MLOps best practices
Benefits
Competitive Salary and Bonus Plan
Comprehensive health insurance plan
Retirement savings plan (401k) with company match
Remote working environment
A flexible, unlimited time off policy
Generous paid holiday schedule - 13 in total including Monday after the Super Bowl
Annual performance bonus
Benefits and/or other applicable incentive compensation plans
Data & Analytics Engineer building scalable data platforms for Ledgebrook, an insurance company. Designing pipelines, governance, and cloud data systems while partnering with technical and insurance teams.
Data Platform Engineer building ingestion pipelines, storage, governance, and observability systems for Movable Ink’s AI - driven marketing personalization platform. Supporting scalable, secure, multi - tenant data services.
Senior Data Engineer building AWS ETL/ELT pipelines and dbt models for Tango’s cloud real - estate and facilities SaaS. Managing databases, data quality, warehousing, and observability.
Senior Data Engineer building AWS and dbt data pipelines for Tango’s cloud real - estate and facilities SaaS platform. Managing databases, warehouses, data quality, and analytics - ready datasets.
Data Engineer building streaming pipelines and analytical databases for Movable Ink’s data - activated marketing personalization platform. Developing Elixir/Python services and reliable event - data products at billions - event scale.
Senior Data Engineer needed for a 1 - year contract with an asset management client in Toronto. Requires 7+ years of experience, Snowflake, Python, SQL, and cloud platform expertise.
Senior Data Engineer owning finance data pipelines and models for Instacart, a grocery delivery platform. Supporting accounting, billing, revenue, and financial reporting at scale.
Senior AI Data Engineer building production AI systems and marketing data infrastructure for Samsara’s connected - operations IoT platform. Automating workflows and enabling data - driven decisions.
Senior Data Engineer modernizing Alberta government regulatory data platforms. Building governed Azure pipelines, models, APIs, and analytics - ready datasets for DRAS.
Senior Data Engineer modernizing Alberta environmental regulatory data on Microsoft Azure. Building governed pipelines, APIs, and analytics - ready datasets for the DRAS platform.