Data Engineer responsible for designing and maintaining data pipelines to support analytics and AI initiatives. Collaborating with various stakeholders to ensure high data quality and efficiency.
Responsibilities
Design, build, and maintain batch, incremental, and file-based data pipelines to support analytics, reporting, and operational use cases.
Support data migration initiatives, including the movement of data from legacy systems to modern platforms, ensuring continuity, accuracy, and reconciliation.
Work with enterprise data integration tools such as Talend, including maintaining existing Talend jobs and supporting the migration or re-implementation of Talend pipelines into other platforms where required.
Establish and maintain a robust data engineering environment (e.g., dev/test/prod separation, access controls, naming standards, source control, deployment processes, and monitoring) that enables reliable delivery and safe change management.
Support ingestion, storage, and management of unstructured and semi-structured datasets (e.g., documents, PDFs, text extracts, files, metadata) alongside traditional relational data.
Support and optimize data models in Power BI dashboards and AI-enabled analytics.
Prepare and curate analytics- and AI-ready datasets, including clean, well-defined tables and reference data used in automation and AI-enabled solutions.
Collaborate closely with AI & Automation Specialists to ensure data requirements for AI initiatives are well understood, properly sourced, validated, and production-ready.
Translate business and reporting requirements into scalable data models, transformations, and pipeline designs.
Develop and maintain trusted datasets that support dashboards, scorecards, client reporting, and downstream analytics.
Implement data quality checks, validation logic, reconciliation processes, and monitoring to ensure data reliability and consistency.
Troubleshoot and resolve data issues across ingestion, transformation, and consumption layers, identifying root causes and remediation actions.
Develop and maintain clear documentation, including source-to-target mappings, data definitions, data lineage, and known limitations.
Work in partnership with analytics, IT, data governance, and security stakeholders to ensure data solutions comply with privacy, security, and regulatory expectations.
Contribute to the prioritization, planning, and delivery of multiple data engineering, migration, and enablement initiatives in parallel.
Requirements
Bachelor’s Degree in Computer Science, Information Technology, Data Analytics or related field
4+ year experience working in data engineering, analytics engineering, business intelligence, or data platform roles.
4+ years Experience with data integration and ETL/ELT tools (e.g., Talend, Azure Data Factory, or similar enterprise platforms).
Experience with Microsoft Power Automate (or similar workflow automation tools) to orchestrate process automation, notifications, and data movement/integration.
Proficiency with Python (or similar scripting languages) for data transformations, automation, API integrations, and lightweight tooling.
Understanding of data orchestration and scheduling patterns, including dependency management, retries, and operational runbooks for production pipelines.
Experience implementing data quality and observability practices (e.g., validation rules, monitoring/alerting, SLAs) to ensure reliable and trusted datasets.
Working knowledge of secure data handling and privacy-by-design principles (e.g., least-privilege access, encryption, PHI/PII considerations) when building and operating data pipelines.
Strong SQL experience for data transformation, validation, reconciliation, and performance tuning.
Experience with Power BI (data modeling, DAX fundamentals, and performance considerations) to support scalable dashboards and self-serve analytics.
Experience working with both structured and unstructured datasets, including file-based or document-oriented data sources.
Familiarity with relational and cloud-based data platforms used for analytics and reporting.
Experience supporting data migrations, legacy system consolidation, or platform modernization initiatives.
Understanding of data quality, documentation practices, and foundational data governance principles.
Strong analytical and problem-solving skills, particularly in diagnosing data quality issues and pipeline failures.
Ability to collaborate effectively with both technical and non-technical stakeholders, including analytics and AI delivery teams.
Strong written and verbal communication skills, with the ability to clearly document data flows, dependencies, and assumptions.
GCP Data Platform Engineer maintaining and optimizing production data platforms for Innodata, a global data engineering and AI services company. Supporting Airflow pipelines, GCP infrastructure, reliability, and troubleshooting.
Senior Data Engineer building finance data pipelines and models for Instacart’s grocery delivery platform. Owning financial data infrastructure supporting accounting, billing, invoicing, and reporting.
Senior Data Architect modernizing Alberta justice data marts into an integrated enterprise data warehouse. Designing ETL, Power BI testing, data models, governance, and reporting architecture.
Senior Data Architect integrating Alberta court data marts into an enterprise data warehouse. Designing ETL, data models, governance practices, and Power BI validation reports.
Data & Analytics Engineer building scalable data platforms for Ledgebrook, an insurance company. Designing pipelines, governance, and cloud data systems while partnering with technical and insurance teams.
Data Engineer building scalable pipelines for SumerSports’ football intelligence platform. Supporting deep learning, video, LLM, analytics, and AI - driven products across sports.
Data Platform Engineer building ingestion pipelines, storage, governance, and observability systems for Movable Ink’s AI - driven marketing personalization platform. Supporting scalable, secure, multi - tenant data services.
Data Engineer building streaming pipelines and analytical databases for Movable Ink’s data - activated marketing personalization platform. Developing Elixir/Python services and reliable event - data products at billions - event scale.
Senior Data Engineer building AWS ETL/ELT pipelines and dbt models for Tango’s cloud real - estate and facilities SaaS. Managing databases, data quality, warehousing, and observability.