Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Data Engineer building scalable pipelines for SumerSports’ football intelligence platform. Supporting deep learning, video, LLM, analytics, and AI-driven products across sports.

Responsibilities

  • Build and operate robust data pipelines for ingestion, cleaning, and transformation using Databricks, Airflow, or Kubernetes
  • Develop efficient ETL/ELT workflows in Python and SQL for batch and streaming workloads
  • Partner with ML/AI teams to make datasets and tools discoverable and safe for autonomous agents, including evaluation and guardrails for AI-generated queries
  • Develop retrieval pipelines (RAG, vector search) over structured statistics and unstructured sources such as scouting notes and video metadata
  • Model and maintain structured data assets in Delta, Parquet, and Iceberg for reliability, versioning, and lineage tracking
  • Implement orchestration and monitoring by scheduling jobs, tracking dependencies, and automating failure recovery
  • Ensure data quality and compliance through validation frameworks, schema enforcement, and audit logging
  • Contribute to data platform evolution by evaluating tools, standardizing best practices, and improving developer experience
  • Support performance and cost optimization across compute, storage, and orchestration systems
  • Collaborate with MLOps and Sports Data teams to integrate data and AI systems

Requirements

  • 3–8 years of experience as a Data Engineer or ETL Developer in a production environment
  • Proficiency in Python and SQL
  • Strong familiarity with Databricks, Spark, or equivalent big-data frameworks
  • Experience with workflow orchestration tools such as Airflow, Dagster, Luigi, or Prefect
  • Deep understanding of data modeling, data warehousing, and distributed data processing
  • Knowledge of modern data lakehouse architectures
  • Familiarity with CI/CD, GitHub Actions, Infrastructure as Code, and data pipeline testing frameworks
  • Comfort working cross-functionally with ML, product, and analytics teams
  • Exposure to LLM-powered data tools, including text-to-SQL, RAG, agent/tool interfaces such as MCP, or natural-language analytics
  • Previous work with cloud infrastructure such as AWS, GCP, or Azure
  • Experience with container orchestration using Docker or Kubernetes
  • Preferred: previous experience with sports, telemetry, or sensor data pipelines
  • Preferred: familiarity with streaming frameworks and event-driven data processing such as Kafka, Spark Structured Streaming, or Flink
  • Preferred: general knowledge of American football, the NFL, and college football
  • Preferred: background in data governance, lineage, and observability tools such as Monte Carlo, Great Expectations, Unity Catalog, or OpenLineage
  • Preferred: experience designing semantic layers or metric definitions consumed by AI and BI tools
  • Preferred: exposure to machine-learning model management and MLOps best practices

Benefits

  • Competitive Salary and Bonus Plan
  • Comprehensive health insurance plan
  • Retirement savings plan (401k) with company match
  • Remote working environment
  • A flexible, unlimited time off policy
  • Generous paid holiday schedule - 13 in total including Monday after the Super Bowl
  • Annual performance bonus
  • Benefits and/or other applicable incentive compensation plans

Job title

Job type

Full Time

Experience level

Mid levelSenior

Salary

$160,000 - $190,000 per year

Degree requirement

No Education Requirement

Tech skills

AirflowAWSAzureCloudDockerETLGoogle Cloud PlatformKafkaKubernetesPythonSparkSQLUnity

Location requirements

RemoteUnited States

Report this job

Found something wrong with the page? Please let us know by submitting a report below.