Staff Data Engineer

Posted 8 hours ago

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Staff Data Engineer building scalable pipelines and lakehouse architecture for Sonatype, a software supply chain security company. Driving trusted analytics, ML, and business intelligence data with Databricks, Spark, and modern cloud technologies.

Responsibilities

  • Design, build, and maintain scalable data pipelines and ETL/ELT processes
  • Architect and optimize data models and storage solutions for analytics and operational use
  • Collaborate with data scientists, analysts, and engineers to deliver trusted, high-quality datasets
  • Own and evolve parts of the data platform using Databricks and Spark
  • Implement observability, alerting, and data quality monitoring for critical pipelines
  • Drive data engineering best practices, including documentation, testing, and CI/CD
  • Drive long-term architectural vision and mentor the team on engineering best practices
  • Partner with stakeholders to ensure data solutions support business outcomes
  • Contribute to the design and evolution of the next-generation data lakehouse architecture

Requirements

  • 8+ years of experience as a Data Engineer or similar backend engineering role
  • Bachelor’s degree in Computer Science, Engineering, or related technical field
  • Databricks optimization, including tuning Spark jobs, optimizing joins, and managing Delta Lake architecture
  • Experience with AI-assisted development tools and AI/ML technologies
  • Strong programming skills in Python, Scala, or Java
  • Hands-on experience with distributed data systems such as Spark or Kafka
  • Proficiency writing and optimizing complex SQL and NoSQL queries
  • Experience building and maintaining robust production ETL/ELT pipelines
  • Understanding of data modeling techniques, including star schema and dimensional modeling
  • Familiarity with software supply chain, cybersecurity, or large-scale software ecosystem data
  • Track record improving data platform reliability, scalability, performance, and cost efficiency
  • Familiarity with workflow orchestration tools such as Airflow or Dagster
  • Hands-on experience with cloud data platforms, particularly AWS
  • Familiarity with Delta Lake, Apache Iceberg, or Apache Hudi
  • Experience implementing data observability, lineage, governance, and automated data quality frameworks
  • Experience designing real-time or streaming data architectures using data lake technologies

Benefits

  • Parental leave
  • Diversity and inclusion working groups
  • Flexible working practices
  • Paid Volunteer Time Off (VTO)

Job title

Job type

Full Time

Experience level

Lead

Salary

Not specified

Degree requirement

Bachelor's Degree

Tech skills

AirflowApacheAWSCloudCyber SecurityETLJavaKafkaNoSQLPythonScalaSparkSQL

Location requirements

RemoteCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.