Data Engineer building scalable Databricks ingestion pipelines for Irth Solutions’ infrastructure-protection SaaS. Supporting stakeholder-engagement analytics, LLM/NLP pipelines, and downstream data models.
Responsibilities
Design, build, and maintain ingestion pipelines from high-volume external APIs that run continuously and reliably at scale
Implement ingestion and transformation workflows in Databricks using Spark/PySpark, SQL, and Delta Live Tables
Apply medallion architecture patterns from raw content to clean, structured, analysis-ready data
Build deduplication and relevance-filtering infrastructure in collaboration with the Data Scientist
Implement schema evolution handling and data validation rules
Configure and manage Delta Lake storage structures, tables, partitions, and optimization routines
Design and evolve data schemas balancing query performance, cost, and maintainability
Maintain metadata and table-structure documentation for Data Science and application teams
Ensure pipeline reliability and observability through error handling, retries, monitoring, and alerting
Adapt pipelines to changing external API contracts, rate limits, authentication methods, and new data sources
Troubleshoot pipeline failures, perform recovery, and tune performance
Build, schedule, and monitor workflows using Databricks Workflows, Delta Live Tables, or similar tools
Contribute to CI/CD pipelines for code deployment, versioning, and environment management
Collaborate with the Data Scientist to provide structured data for LLM/NLP pipelines and downstream models
Participate in data-architecture decisions and propose solutions as team needs evolve
Document pipelines, data dictionaries, job schedules, and transformation logic
Support onboarding of new data sources and pipelines as the product expands
Requirements
Strong preference for candidates residing in Quebec
Fluency in French (spoken and written) is a strong asset in addition to English
3 to 5 years of experience in data engineering, with solid experience building and operating production-grade data pipelines
Familiarity with data modeling, data quality, and schema evolution
Solid understanding of data pipeline reliability practices: monitoring, alerting, and handling failures gracefully in a continuously running system
Hands-on experience with Databricks or an equivalent Spark-based environment, including schema design, Delta Lake, performance tuning, and pipeline orchestration
Experience with at least one major cloud provider; Azure preferred, AWS/GCP also beneficial
Experience integrating with external APIs at scale, including authentication, pagination, rate limiting, retries, and error handling
Strong proficiency in Python and SQL
Comfortable working with unstructured/semi-structured text data at scale
LLM prompting experience and/or basic understanding of AI/NLP concepts
Exposure to medallion architecture or lakehouse best practices
Experience with orchestration frameworks such as ADF, Workflows, Airflow, or DBX
Experience with CI/CD tools and version control, such as Git or GitHub Actions
Basic understanding of security practices, including RBAC, encryption, and credential management
Databricks certification (Data Engineer Associate or equivalent)
Benefits
Competitive compensation package based on experience and qualifications
Medical, Dental, and Vision Insurance
401(k) Plan with Company Match
Generous Paid Time Off (PTO)
Company-Paid Holidays
Flexible Work Options / Work-from-home opportunities
On-Call Compensation — Additional pay for eligible on-call shifts
Senior Data Engineer owning NimbleRx’s compliant healthcare data platform. Building scalable pipelines, warehouse models, and AI tooling for pharmacy and patient engagement operations.
Administrateur principal de plateformes de données chez EDC, société d’État aidant les entreprises canadiennes à réussir à l’étranger. Conception, exploitation et gouvernance de plateformes de données infonuagiques.
Senior Data Platform Administrator shaping and operating EDC’s secure enterprise data platforms. Supporting Canadian businesses through Export Development Canada’s trade - finance solutions.
Staff full - stack engineer architecting GitLab’s data products, integrations, and knowledge graph. Building APIs, marketplace data access, and AI - powered developer intelligence.
Data Governance Architect shaping governance strategies and cloud data architectures for Lovelytics’ enterprise data and AI consulting clients. Leading technical delivery, presales, migrations, and governance implementations on Databricks.
Senior Data Engineer owning Databricks lakes, pipelines, and AI data infrastructure. Building Quandri’s AI operating system for insurance agencies and brokerages.
Senior Data Engineer architecting Azure and Databricks data platforms for TTEC Digital’s client experience solutions. Building governed pipelines, APIs, MCP integrations, and agentic AI data services.
Senior Data Engineer building Palantir Foundry data solutions for Unit8, a Swiss AI and data analytics consultancy. Supporting client delivery and establishing its Canadian market presence.
Staff Software Engineer building AWS data infrastructure and integrations for Solink’s cloud video - security platform. Driving scalability, architecture, and engineering mentorship.
Senior Azure Fabric Data Engineer building reliable pipelines and modern data platforms for Data Elephant, a Canadian data, analytics, and AI consultancy. Delivering trusted datasets for reporting, AI, and machine learning.