Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Staff Data Engineer designing AI-ready data platforms for Robots and Pencils, an applied AI engineering firm. Leading pipelines, governance, modernization, and production AI/ML data systems.

Responsibilities

  • Define data architecture and platform strategy across pipelines, warehouses, and data lakes
  • Build and optimize scalable batch and real-time data pipelines
  • Define and enforce data governance, quality standards, and compliance frameworks
  • Build monitoring, logging, and alerting for data pipelines and services
  • Contribute to CI/CD workflows for data deployment and automation
  • Drive data platform modernization for performance, cost, and scalability
  • Use Claude, Cursor, and other modern AI assistants to ship higher-quality work
  • Design and implement data contracts and event flows with backend, platform, and engineering teams
  • Lead data pipelines for production AI/ML systems, including embeddings, vector stores, RAG data preparation, feature stores, and training/inference data flows
  • Integrate data services with APIs, middleware, and third-party systems
  • Partner with leadership on data strategy
  • Collaborate with engineering, analytics, AI, and product teams
  • Advocate for data quality, governance, and platform best practices
  • Establish data engineering standards across the team
  • Mentor junior and mid-level engineers
  • Make high-stakes architectural decisions with clear ownership and consideration of long-term tradeoffs

Requirements

  • 7+ years of professional data engineering experience, including leading complex data platform initiatives
  • Strong system architecture background and expertise in distributed data systems
  • Expert proficiency in Python, Scala, and SQL
  • Deep expertise with cloud-native data platforms and enterprise data warehousing
  • Strong expertise in data pipeline orchestration and processing
  • Strong experience with streaming platforms and real-time data processing, such as Kafka, Kinesis, or Pub/Sub
  • Strong data modeling expertise and experience with data transformation
  • Strong experience with data quality, governance, and compliance frameworks
  • Strong experience with container orchestration and CI/CD for data systems
  • Strong experience building data pipelines for production AI/ML systems, including embeddings, vector stores, RAG data preparation, feature stores, and training/inference data flows
  • Demonstrated leadership and technical mentoring experience
  • Strong stakeholder communication skills and ability to translate technical depth across audiences
  • Demonstrable day-to-day usage and expert knowledge of AI-forward coding tools such as Claude and Cursor
  • Excellent problem-solving skills and sound judgment in ambiguous technical and business challenges
  • Experience with data mesh or data fabric concepts, lakehouse architectures, or governance framework implementation is a plus
  • Experience handling and modeling healthcare-industry data is a plus
  • AWS certifications, such as Certified Data Engineer – Associate, are strongly preferred
  • Employment may be conditional upon successful completion of a background check in accordance with local legislation

Benefits

  • Potential employment offer conditional upon successful completion of a background check in accordance with local legislation
  • Equal employment opportunity commitment

Job title

Job type

Full Time

Experience level

SeniorLead

Salary

Not specified

Degree requirement

No Education Requirement

Tech skills

AWSCloudKafkaPythonScalaSQL

Location requirements

RemoteUnited States

Report this job

Found something wrong with the page? Please let us know by submitting a report below.