MLOps Data Engineer bridging data science and production systems at Triton Digital. Designing CI/CD pipelines and optimizing data processing with Apache Spark for advertising systems.
Responsibilities
Design, implement, and maintain CI/CD pipelines for machine learning workflows using tools like GitHub Actions, Azure DevOps, or Jenkins.
Build and optimize data processing pipelines in Apache Spark (PySpark and Scala) for large-scale, distributed listener datasets.
Deploy and manage Databricks environments, ensuring efficient cluster usage, job scheduling, and cost optimization.
Collaborate with data scientists to productionize ML models, integrating them into scalable APIs or batch processing systems that feed real-time, machine-readable audience signals.
Implement automated testing, monitoring, and alerting for ML pipelines to ensure the reliability and reproducibility that certified buyers require.
Champion best practices in version control, model registry management, and environment reproducibility.
Help evolve our listener data infrastructure toward agent-compatible supply — live, structured, queryable data feeds that autonomous buying systems can discover and act on without human mediation.
Requirements
Proven experience in Data Engineering, MLOps, and DevOps roles with a focus on automation and scalability.
Strong programming skills in Python, with hands-on experience in Apache Spark.
Scala is a huge plus.
Advanced expertise in Databricks, including Delta Lake, structured streaming, feature engineering.
Solid understanding of CI/CD principles and tools (e.g., GitHub Actions, Jenkins, Azure DevOps, GitLab CI, ArgoCD).
Familiarity with cloud platforms (AWS, Azure, or GCP) for data and ML workloads.
A problem-solving mindset and the ability to work closely with cross-functional teams.
Strong architectural mindset, capable of evaluating trade-offs across cost, performance, scalability, and maintainability when selecting tools and designing systems.
Experience working with containerized and orchestrated environments (Kubernetes / OpenShift), including deployment, scaling, and fault tolerance of data and ML workloads.
Advanced English required.
French is an asset.
Familiarity with IAB data standards, programmatic advertising infrastructure, or AdTech data pipelines is a strong asset.
Benefits
Fully remote position (must be based in ONTARIO or QUEBEC)
4 weeks of vacation + 5 paid personal days annually
Group insurance programs as of your first day, including access to telemedicine and an EAP
Senior Data Engineer building scalable pipelines and AI - ready infrastructure for CINC Systems’ community association management software. Enabling reliable analytics, automation, and intelligent product experiences.
Senior Data Engineer building AWS, Snowflake, Python, PySpark, and SQL data products for Sun Life’s Canadian financial - services business. Enabling analytics, data science, governance, and client insights.
AWS Data Engineer needed for 12+ month contract in Toronto. Lead design of cloud - native data pipelines, mentor engineers, and drive data quality and cost optimization.
Hiring a Test Data Management (TDM) contractor for a long - term hybrid role in Toronto. Requires strong TDM experience, Delphix, and Java; Python is a plus.
Senior Data Engineer building AWS/GCP data infrastructure for Ready’s broadband monitoring and BEAD programs. Developing reliable pipelines, data quality systems, and AI - enabled workflows.
Seeking an experienced Data Architect to design scalable data solutions for enterprise environments. Expertise in Data Architecture and Cloud - Based Data Analytics required.
Databricks Data Engineer building scalable ingestion and ETL pipelines for Avanquest’s global software products. Ensuring data quality, lineage, reliability, and optimized storage.
Administrateur de plateforme de données chez EDC, société d’État aidant les entreprises canadiennes à réussir à l’étranger. Soutien, sécurisation et amélioration de plateformes infonuagiques de données et d’analyse.