Data Engineer building scalable Databricks ingestion pipelines for Irth Solutions’ infrastructure-protection SaaS. Supporting stakeholder-engagement analytics, LLM/NLP pipelines, and downstream data models.
Responsibilities
Design, build, and maintain ingestion pipelines from high-volume external APIs that run continuously and reliably at scale
Implement ingestion and transformation workflows in Databricks using Spark/PySpark, SQL, and Delta Live Tables
Apply medallion architecture patterns from raw content to clean, structured, analysis-ready data
Build deduplication and relevance-filtering infrastructure in collaboration with the Data Scientist
Implement schema evolution handling and data validation rules
Configure and manage Delta Lake storage structures, tables, partitions, and optimization routines
Design and evolve data schemas balancing query performance, cost, and maintainability
Maintain metadata and table-structure documentation for Data Science and application teams
Ensure pipeline reliability and observability through error handling, retries, monitoring, and alerting
Adapt pipelines to changing external API contracts, rate limits, authentication methods, and new data sources
Troubleshoot pipeline failures, perform recovery, and tune performance
Build, schedule, and monitor workflows using Databricks Workflows, Delta Live Tables, or similar tools
Contribute to CI/CD pipelines for code deployment, versioning, and environment management
Collaborate with the Data Scientist to provide structured data for LLM/NLP pipelines and downstream models
Participate in data-architecture decisions and propose solutions as team needs evolve
Document pipelines, data dictionaries, job schedules, and transformation logic
Support onboarding of new data sources and pipelines as the product expands
Requirements
Strong preference for candidates residing in Quebec
Fluency in French (spoken and written) is a strong asset in addition to English
3 to 5 years of experience in data engineering, with solid experience building and operating production-grade data pipelines
Familiarity with data modeling, data quality, and schema evolution
Solid understanding of data pipeline reliability practices: monitoring, alerting, and handling failures gracefully in a continuously running system
Hands-on experience with Databricks or an equivalent Spark-based environment, including schema design, Delta Lake, performance tuning, and pipeline orchestration
Experience with at least one major cloud provider; Azure preferred, AWS/GCP also beneficial
Experience integrating with external APIs at scale, including authentication, pagination, rate limiting, retries, and error handling
Strong proficiency in Python and SQL
Comfortable working with unstructured/semi-structured text data at scale
LLM prompting experience and/or basic understanding of AI/NLP concepts
Exposure to medallion architecture or lakehouse best practices
Experience with orchestration frameworks such as ADF, Workflows, Airflow, or DBX
Experience with CI/CD tools and version control, such as Git or GitHub Actions
Basic understanding of security practices, including RBAC, encryption, and credential management
Databricks certification (Data Engineer Associate or equivalent)
Benefits
Competitive compensation package based on experience and qualifications
Medical, Dental, and Vision Insurance
401(k) Plan with Company Match
Generous Paid Time Off (PTO)
Company-Paid Holidays
Flexible Work Options / Work-from-home opportunities
On-Call Compensation — Additional pay for eligible on-call shifts
Senior Databricks Architect needed for contract role in Winnipeg, MB. Must have Azure Databricks, PySpark, and Azure DevOps expertise; onsite mandatory.
Data Engineering Developer building scalable Azure data pipelines for an agile tech development firm. Advising clients, implementing governance, and improving cloud data engineering standards.
Ingénieur(e) données concevant des pipelines et architectures analytiques Azure. Collaboration client et amélioration des standards de gouvernance, qualité et performance des données.
Junior Data Engineer building reliable ETL pipelines and data solutions for PLATO, Canada’s Indigenous - owned technology services company. Improving data quality, accessibility, governance, and pipeline performance.
Senior Data Engineer building scalable, AI - powered data infrastructure for CloudBlue, HostPapa’s cloud commerce platform. Developing real - time pipelines, APIs, and production analytics systems.
Senior Fabric Data Engineer modernizing enterprise data for Canadian IT consulting clients. Building Microsoft Fabric pipelines, curated datasets, security controls, and analytics - ready models.
Staff Data Engineer building scalable pipelines and lakehouse architecture for Sonatype, a software supply chain security company. Driving trusted analytics, ML, and business intelligence data with Databricks, Spark, and modern cloud technologies.
Senior Fabric Data Engineer modernizing learning - platform data for Cornerstone Galaxy. Building Microsoft Fabric pipelines, curated datasets, governance, and analytics - ready integrations.
GCP Data Platform Engineer maintaining and optimizing production data platforms for Innodata, a global data engineering and AI services company. Supporting Airflow pipelines, GCP infrastructure, reliability, and troubleshooting.