Senior Applied Research Engineer developing AI model systems and gold data layers for healthcare providers at PointClickCare. Fostering innovation and collaboration with AI experts to enhance healthcare delivery.
Responsibilities
Own the gold data layer. Transform messy, silver tables into curated, semantically rich, clean and documented gold datasets suitable for AI model development, including datasets and features reusable for AI development across projects.
Maintain the data as products and needs evolve. To do this you will
Reverse-engineer data semantics. Talk with product engineers, clinical and workflow experts to learn how the products are used and how data are created in the field.
Understand SQL queries, stored procedures, technical data definitions, and other code to know how products represent and transform data.
Learn how data are ingested into the data lake, what silver tables and columns actually represent and how they behave.
Capture provenance, semantics, clinical event sequencing, cross module record linkage and known quirks.
Bridge semantics with AI needs. Understand researcher data needs to design and build the gold data product, with documentation that evolves, to meet AI applied research needs for a highly efficient AI-first foundation for model R&D.
Curate datasets across modalities. For various AI uses such as generative AI, RAG, predictive and other techniques, support researcher needs for chunked and tagged unstructured content with rich metadata, point-in-time-correct features and clean labels. For classical ML and statistical work, deliver model-ready tables.
Build pipelines for reuse. Develop transformations from silver into gold inside Databricks/Spark as scheduled, observable workloads. Design them so researchers can iterate on new features and data mixes without rebuilding from scratch.
Automate quality, filtering, and synthesis. Support research needs for programmatic labeling, weak supervision, near-duplicate detection, boilerplate and noise removal, and LLM-API-driven synthetic data generation where ground truth is scarce.
Version and hand off. Maintain reproducible dataset snapshots. Define clean lineage and semantic definitions so the downstream team can use and re-use gold datasets in AI R&D.
Requirements
5+ years building production data systems, with at least 2 supporting ML or AI workloads.
Track record of learning complex new data domains quickly, through reading source code, interviewing experts, and building durable artifacts others rely on.
Advanced Python, SQL, and PySpark /Databricks for working with large, messy data.
Expert SQL specifically: comfortable reading complex stored procedures and reverse-engineering business logic from queries.
AI domain literacy: working understanding of embeddings, tokenization, feature engineering, point-in-time correctness, train/validation/test splits, data drift, and the differences between what classical ML and generative models need from data.
Data wrangling across modalities: transforming unstructured content (text, PDFs, transcripts, logs) and structured tabular data into clean, model-ready forms.
AI-friendly data formats (Parquet, Hugging Face datasets) and storage layout decisions — partitioning, sharding, caching, that keep researcher workflows responsive in Azure, AWS or other working environments.
Data quality, filtering, and synthesis pipelines: support for programmatic labeling and weak supervision (e.g. Snorkel or equivalent), near-duplicate detection (MinHash /LSH), content and quality filters, LLM-API-driven synthetic data generation.
Pipeline orchestration (e.g. a la Airflow, Databricks Workflows, Dagster , or Prefect) and dataset versioning including Unity Catalog and feature-store support.
Experience handling regulated or sensitive data under controlled access (HIPAA or equivalent). Familiarity with general de-identification concepts.
Git-based version control and CI/CD for data and code.
Strong written documentation. Skill in eliciting requirements and tacit knowledge from technical and non-technical experts.
Bachelor’s degree in computer science, data science, engineering, statistics, or related field. Equivalent practical experience considered.
Senior Databricks Architect needed for contract role in Winnipeg, MB. Must have Azure Databricks, PySpark, and Azure DevOps expertise; onsite mandatory.
Data Engineering Developer building scalable Azure data pipelines for an agile tech development firm. Advising clients, implementing governance, and improving cloud data engineering standards.
Ingénieur(e) données concevant des pipelines et architectures analytiques Azure. Collaboration client et amélioration des standards de gouvernance, qualité et performance des données.
Junior Data Engineer building reliable ETL pipelines and data solutions for PLATO, Canada’s Indigenous - owned technology services company. Improving data quality, accessibility, governance, and pipeline performance.
Senior Data Engineer building scalable, AI - powered data infrastructure for CloudBlue, HostPapa’s cloud commerce platform. Developing real - time pipelines, APIs, and production analytics systems.
Senior Fabric Data Engineer modernizing enterprise data for Canadian IT consulting clients. Building Microsoft Fabric pipelines, curated datasets, security controls, and analytics - ready models.
Staff Data Engineer building scalable pipelines and lakehouse architecture for Sonatype, a software supply chain security company. Driving trusted analytics, ML, and business intelligence data with Databricks, Spark, and modern cloud technologies.
Senior Fabric Data Engineer modernizing learning - platform data for Cornerstone Galaxy. Building Microsoft Fabric pipelines, curated datasets, governance, and analytics - ready integrations.
GCP Data Platform Engineer maintaining and optimizing production data platforms for Innodata, a global data engineering and AI services company. Supporting Airflow pipelines, GCP infrastructure, reliability, and troubleshooting.