Junior Data Scientist developing data infrastructure for CaRMS to support decision-making. Engaging in data engineering and data-driven insights for stakeholders in a remote role.
Responsibilities
Designing, implementing, and operating critical data infrastructure, including systems for updating the corporate Data Warehouse, passing information to and from our matching software, and generating data products (operational reporting, data contracts, match statistics, etc.)
Migrating ETL associated with passing information to and from the matching software from Informatica PowerCenter to new PostgreSQL/Python-based data platform (PostgreSQL, SQLAlcemy / SQLModel, Dagster, and MkDocs)
Developing internal matching platform API (using FastAPI) to run ETL associated with the matching software and help application developers use it
Consolidating overlapping SQL views across data products to ensure consistency
Developing modular Python-based reporting framework for producing data contracts, operational reporting, and custom data requests
Maintaining and extending match simulation software and conducting “what-if” scenario analysis for stakeholders in collaboration with the Lead Data Scientist
Contributing to R markdown/Quarto-based "insight" research pieces for internal and external stakeholders
Helping our stakeholders understand applicant and employer preferences (preference modeling)
Developing better ways to help our clients find their ideal candidates/residency positions (for use in our broader web application)
Requirements
Four-year degree in data science, economics, computer science, engineering, applied mathematics, statistics or equivalent work experience.
Very strong proficiency in Python and advanced SQL skills is required.
3-5 years of experience with Python-based data engineering / data science packages (particularly SQLAlchemy/SQLModel, pandas, Dagster, FastAPI, and LangChain).
Experience using cloud data storage (AWS S3), PostgreSQL-compatible database services (i.e., fully managed through RDS / Aurora, or self-managed on Amazon EC2), and compute (EC2, ECS, Fargate) is very highly valued.
Significant experience with relational database systems (e.g., Oracle, PostgreSQL, etc.).
Deep understanding of data management concepts associated with designing, building, maintaining, and extending an Enterprise Data Warehouse.
Use of version control (Git) and test-based development practices should be strongly engrained in your workflow.
Practical experience with any of the following is valued: Implementing semantic search and Q&A on documents, Computational statistics, particularly resampling techniques, Matching algorithms, Informatica PowerCenter, Using Quarto/R markdown to produce reproducible reporting, Developing and supporting dashboards (e.g., Tableau, MS Power BI, etc.), Jira and Confluence collaboration tools.
Staff Data Scientist developing advanced machine - learning models for MindBridge Analytics’ enterprise datasets. Leading experimentation, production deployment, technical standards and mentorship.
Data conversion manager leading ETL services for Catalis, a government SaaS and integrated payments provider. Scaling migration standards across court, land records, criminal justice, and jury solutions.
Microbiome Data Scientist analyzing microbiome and clinical data for Tiny Health’s precision testing platform. Developing biomarkers, metrics, publications, and new personalized - health products.
Data Scientist building AI fraud detection models for Oscilar’s risk platform. Analyzing large datasets and strengthening fraud prevention for banks, fintechs, and digital organizations.
GTM Analytics Lead owning revenue data models, metrics, and Sales BI. Powering decisions for ButterflyMX’s smartphone - based building access platform.
Data Scientist developing models, dashboards, and analytics to detect transactional fraud at Desjardins, North America’s largest cooperative financial group. Deploying and monitoring mitigation tools.
Senior Data Scientist powering product and operational analytics for Flagler Health’s musculoskeletal care platform. Measuring patient journeys, clinical workflows, engagement, and business performance.
TD digital analytics expert optimizing banking customer journeys through web and mobile data. Delivering dashboards, insights, and recommendations to improve conversion, engagement, and digital performance.
Staff Data Scientist embedded across growth, product, AI, or partnerships at Jerry.ai, an AI - powered platform managing car and home ownership. Defining metrics, running experiments, and driving data - informed business decisions.
Technical Lead governing AI data readiness and API gateways at Coupa, whose platform helps businesses optimize total spend. Defining safeguards for GenAI applications, data quality, compliance, and observability.