Lead AI Platform Engineer securing and operating EQ Bank’s enterprise AI platforms. Driving Azure engineering, observability, automation, governance, and production readiness.
Responsibilities
Lead the engineering, configuration, and operation of enterprise AI platforms to ensure availability, performance, resilience, and scalability
Define and implement platform engineering patterns, standards, reusable components, and operational guardrails
Lead platform triage, incident resolution, escalation coordination, and post-incident reviews
Track and report service reliability indicators, incident trends, engineering risks, and operational performance improvements
Enable approved AI use cases in non-production and production environments through environment readiness, dependency validation, release readiness, operational supportability, and service transition planning
Partner with architecture, security, cloud, infrastructure, delivery, and application teams on secure, supportable, scalable platform implementations
Ensure AI platform changes meet engineering, operational, security, and control readiness criteria
Design and improve observability capabilities including telemetry, logging, metrics, traces, dashboards, and alerting
Lead automation initiatives to reduce manual effort, improve reliability, and standardize operational activities
Analyze operational data for anomalies, recurring issues, root-cause patterns, performance bottlenecks, and service improvement opportunities
Implement AI Ops use cases including alert correlation, anomaly detection, forecasting, root-cause support, knowledge retrieval, and repetitive-task automation
Mentor engineers on observability, automation, troubleshooting, and service reliability
Embed governance, security, privacy, auditability, traceability, and human oversight into AI platform engineering
Assess implementation risks, close control gaps, maintain audit and governance evidence, and escalate technical and control risks
Maintain visibility of AI platform assets, validate ownership and configuration integrity, and promote engineering standards and reusable patterns
Requirements
University degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience
7+ years of experience in platform engineering, site reliability engineering, DevOps, cloud operations, enterprise IT operations, or production platform support
Experience leading technical delivery, engineering standards, production readiness, incident response, problem management, service restoration, and operational reporting for enterprise platforms
Advanced experience with cloud platforms, observability, automation, configuration management, and integration patterns
Experience with Azure Automation runbooks, Azure AI, Copilot integrations, AKS, virtual networks, App Service, and supporting Azure services
Expertise with Azure Monitor, Application Insights, Log Analytics, Grafana, dashboards, alerting, and operational telemetry design
Experience with CI/CD, automation, and infrastructure-as-code tools including Azure DevOps, GitHub Actions, Logic Apps, Bicep, Terraform, Azure Policy, and Key Vault
Knowledge of API Management, open-source API tools, Service Bus, Event Grid, and Apache Kafka
Working knowledge of Elastic, Azure AI Search, Cosmos DB, and related data platform capabilities
Knowledge of enterprise network, edge security, identity, access management, DNA, Fortinet, and Akamai is an asset
Working knowledge of AI/ML operational concepts, model lifecycle support, platform telemetry, governance controls, human-in-the-loop practices, responsible AI, and production monitoring
Understanding of ITIL/ITSM processes including change, release, incident, problem, configuration, service reporting, and operational risk practices
Senior AI Platform Developer building scalable AI services and LLM workflows for MaintainX’s industrial work execution platform. Improving reliability, observability, performance, and cost efficiency.
Infrastructure team lead building and operating Spare’s GCP and Kubernetes platform for on - demand transit. Leading developers while improving reliability, security, AI SRE, and cloud cost efficiency.
Senior AI Platform Developer building reusable, secure AI - agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Driving Azure - based platform capabilities across orchestration, evaluation, observability, and governance.
AI Platform Developer building Azure - based agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Creating secure, observable, governed AI services for product teams.
Kubernetes/DevOps Engineer operating Kubernetes infrastructure for a petabyte - scale social media platform. Building distributed systems and machine - learning workloads for Capgemini Engineering’s client.
Senior Platform Engineer building and operating GitLab Orbit, a Rust - based knowledge graph service for GitLab’s DevSecOps platform. Improving distributed - system reliability, observability, cloud infrastructure, and data workflows.
Platform Software Engineer supporting Invoice Simple, EverCommerce’s invoicing SaaS for small businesses. Building cloud infrastructure, CI/CD pipelines, and reliable production systems.
Software Engineer building cloud infrastructure and deployment systems for Invoice Simple, EverCommerce’s invoicing platform for small businesses. Improving reliability, observability, and production operations.
Senior cybersecurity engineer securing L3Harris’s mission - critical platform - management systems for military and cruise ships. Hardening Linux, networks, software, and infrastructure across the product lifecycle.