Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Lead AI Platform Engineer securing and operating EQ Bank’s enterprise AI platforms. Driving Azure engineering, observability, automation, governance, and production readiness.

Responsibilities

  • Lead the engineering, configuration, and operation of enterprise AI platforms to ensure availability, performance, resilience, and scalability
  • Define and implement platform engineering patterns, standards, reusable components, and operational guardrails
  • Lead platform triage, incident resolution, escalation coordination, and post-incident reviews
  • Track and report service reliability indicators, incident trends, engineering risks, and operational performance improvements
  • Enable approved AI use cases in non-production and production environments through environment readiness, dependency validation, release readiness, operational supportability, and service transition planning
  • Partner with architecture, security, cloud, infrastructure, delivery, and application teams on secure, supportable, scalable platform implementations
  • Guide release coordination, change readiness, maintenance planning, capacity planning, and technical risk mitigation
  • Ensure AI platform changes meet engineering, operational, security, and control readiness criteria
  • Design and improve observability capabilities including telemetry, logging, metrics, traces, dashboards, and alerting
  • Lead automation initiatives to reduce manual effort, improve reliability, and standardize operational activities
  • Analyze operational data for anomalies, recurring issues, root-cause patterns, performance bottlenecks, and service improvement opportunities
  • Implement AI Ops use cases including alert correlation, anomaly detection, forecasting, root-cause support, knowledge retrieval, and repetitive-task automation
  • Mentor engineers on observability, automation, troubleshooting, and service reliability
  • Embed governance, security, privacy, auditability, traceability, and human oversight into AI platform engineering
  • Assess implementation risks, close control gaps, maintain audit and governance evidence, and escalate technical and control risks
  • Maintain visibility of AI platform assets, validate ownership and configuration integrity, and promote engineering standards and reusable patterns

Requirements

  • University degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience
  • 7+ years of experience in platform engineering, site reliability engineering, DevOps, cloud operations, enterprise IT operations, or production platform support
  • Experience leading technical delivery, engineering standards, production readiness, incident response, problem management, service restoration, and operational reporting for enterprise platforms
  • Advanced experience with cloud platforms, observability, automation, configuration management, and integration patterns
  • Experience with Azure Automation runbooks, Azure AI, Copilot integrations, AKS, virtual networks, App Service, and supporting Azure services
  • Expertise with Azure Monitor, Application Insights, Log Analytics, Grafana, dashboards, alerting, and operational telemetry design
  • Experience with CI/CD, automation, and infrastructure-as-code tools including Azure DevOps, GitHub Actions, Logic Apps, Bicep, Terraform, Azure Policy, and Key Vault
  • Knowledge of API Management, open-source API tools, Service Bus, Event Grid, and Apache Kafka
  • Working knowledge of Elastic, Azure AI Search, Cosmos DB, and related data platform capabilities
  • Knowledge of enterprise network, edge security, identity, access management, DNA, Fortinet, and Akamai is an asset
  • Working knowledge of AI/ML operational concepts, model lifecycle support, platform telemetry, governance controls, human-in-the-loop practices, responsible AI, and production monitoring
  • Understanding of ITIL/ITSM processes including change, release, incident, problem, configuration, service reporting, and operational risk practices
  • Technical leadership, mentoring, standards development, implementation guidance, and complex cross-functional delivery coordination
  • Advanced troubleshooting, root-cause analysis, prioritization, risk assessment, and continuous improvement skills
  • Experience creating technical documentation, engineering patterns, operational procedures, support playbooks, dashboards, and user guidance materials

Benefits

  • Full-time employment
  • Hybrid work arrangement

Job type

Full Time

Experience level

Senior

Salary

Not specified

Degree requirement

Bachelor's Degree

Tech skills

ApacheAzureCloudGrafanaITSMKafkaTerraformVault

Location requirements

HybridTorontoCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.