Data Scientist designing and deploying AI-powered applications for querying and analyzing scientific data. Integrating large language models into workflows for enhanced data analysis and visualization.
Responsibilities
Design and implement agentic AI systems that allow scientists to query Oracle databases and scientific data platforms using natural language, generating interactive plots and structured reports from preclinical data.
Integrate large language models into scientific data workflows using both cloud-hosted services (Azure OpenAI) and locally deployed open-weight models (Ollama, vLLM, or similar), including prompt engineering, tool/function calling, guardrails, output validation, and structured output parsing.
Design and implement retrieval-augmented generation (RAG) pipelines over scientific documents and database schemas to ground LLM responses in domain-specific context.
Evaluate, benchmark, and select appropriate LLM backends (cloud vs. local, model size, quantization) based on latency, accuracy, cost, and data privacy requirements.
Build scalable data models and ETL pipelines that surface scientific data through web-based applications and GUIs in Python (Plotly Dash, FastAPI).
Use Docker to build, test, and deploy containerized applications across on-premises and Azure environments.
Communicate effectively with scientific and technical stakeholders, including presenting methods, architectures, and results to broader audiences.
Write detailed application and system documentation using GitHub Pages, Sphinx, or similar professional tooling.
Requirements
Bachelor's degree (minimum) in Computer Science, Engineering, Mathematics, or a related quantitative field
Advanced Python programming skills: clean, well-documented, production-quality code with appropriate testing and error handling
Experience with SQL scripting and relational database systems (Oracle preferred), including query optimization and schema design
Demonstrated ability to work with LLMs and AI agent frameworks — prompt engineering, retrieval-augmented generation (RAG), function/tool calling, structured output parsing, or similar orchestration patterns
Hands-on experience deploying and serving LLMs locally using Ollama, vLLM, llama.cpp, or similar inference frameworks, including model selection, quantization trade-offs, and GPU resource management
Proficiency with Python web frameworks for building interactive front-end applications (Plotly Dash and/or FastAPI), including working knowledge of HTML/CSS for UI refinement
Experience with Docker for building and deploying containerized applications
Strong Git workflows (branching, merging, pull requests) and familiarity with CI/CD tooling (GitHub Actions or similar)
Comfortable working in Linux environments (Ubuntu), writing bash scripts, and managing applications on servers or VMs
Excellent written and verbal communication skills with a demonstrated ability to document systems and workflows professionally.
Benefits
Medical
Dental
Vision
Short-& long-term disability
Accidental death & dismemberment
Life insurance programs
Employee Assistance Program
Travel insurance
Retirement savings programs with company matching contributions
GTM Analytics Lead owning revenue data models, metrics, and Sales BI. Powering decisions for ButterflyMX’s smartphone - based building access platform.
Data Scientist developing models, dashboards, and analytics to detect transactional fraud at Desjardins, North America’s largest cooperative financial group. Deploying and monitoring mitigation tools.
Senior Data Scientist powering product and operational analytics for Flagler Health’s musculoskeletal care platform. Measuring patient journeys, clinical workflows, engagement, and business performance.
TD digital analytics expert optimizing banking customer journeys through web and mobile data. Delivering dashboards, insights, and recommendations to improve conversion, engagement, and digital performance.
Staff Data Scientist embedded across growth, product, AI, or partnerships at Jerry.ai, an AI - powered platform managing car and home ownership. Defining metrics, running experiments, and driving data - informed business decisions.
Technical Lead governing AI data readiness and API gateways at Coupa, whose platform helps businesses optimize total spend. Defining safeguards for GenAI applications, data quality, compliance, and observability.
Senior Principal Data Scientist building predictive models and evaluation systems for Autodesk’s agentic AI design platform. Defining telemetry, experimentation, and product intelligence for complex user workflows.
Senior Data Scientist evaluating football ML models for SumerSports’ NFL intelligence platform. Building rigorous testing, comparison, and experimentation tooling for production model quality.
Data Scientist developing fraud detection solutions for Desjardins, North America's largest cooperative financial group. Applying data science expertise to protect financial products and services from fraud.
Data Scientist II applying SQL, Python, machine learning, and AI to Life & Health insurance products at TD Bank. Developing strategies and insights for better decisions.