Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Azure Platform Engineer operating secure, reliable AKS infrastructure for Smile Digital Health’s FHIR-based healthcare data platform. Owning Kubernetes networking, CI/CD, observability, security, and production operations.

Responsibilities

  • Act as the subject-matter expert (SME) for Kubernetes deployments, troubleshooting, and production issues across all environments
  • Design and deploy Azure Kubernetes Service (AKS) clusters with private cluster configurations, managed identities, and RBAC
  • Own the health, scaling, and lifecycle management of production AKS clusters, including upgrades, node pool management, autoscaling, and capacity planning
  • Configure AKS networking, including Azure CNI, internal load balancers, and ingress controllers (NGINX, Traefik)
  • Design and maintain integrations between AKS and Azure Container Registry, Key Vault via CSI driver, Azure Monitor for containers, Azure SQL, Kafka/Event Hubs, Azure Storage, and client-facing dependencies
  • Manage containerized application deployments using Docker and Helm; maintain reusable chart and templating standards, namespaces, resource quotas, and Azure Policy for AKS
  • Harden AKS environments through policy enforcement, network policies, and image scanning
  • Own container and cluster vulnerability management, including scanning, triage, prioritization, and remediation coordination
  • Contribute to Terraform-based infrastructure as code for provisioning and managing Azure resources
  • Support Azure DevOps or equivalent CI/CD pipelines, including GitOps workflows with Flux/ArgoCD
  • Design, implement, and manage observability across Grafana, Prometheus, Loki, Tempo, Azure Monitor, Log Analytics Workspace, and Application Insights
  • Define SLIs/SLOs and tune alerting to reduce noise
  • Identify, diagnose, and resolve performance bottlenecks across the AKS platform and dependent services
  • Contribute to Disaster Recovery and Business Continuity Planning procedures, including failover drills and RTO/RPO validation
  • Provide escalation support for production Kubernetes and infrastructure incidents; participate in on-call rotation, lead root-cause analysis, and drive preventative actions
  • Document runbooks, post-incident reviews, and operational knowledge; maintain a living knowledge base

Requirements

  • 5+ years of hands-on Kubernetes experience
  • Strong knowledge of Kubernetes internals: scheduling, networking, storage, RBAC
  • Proficiency in Azure CNI networking and AKS private cluster configuration
  • Hands-on experience with Azure PaaS services (ACR, AKV, Azure SQL, Kafka and Managed Identities)
  • Experience troubleshooting network-layer dependencies (DNS, firewall, private endpoints)
  • Experience with Helm
  • Working knowledge of Terraform for infrastructure as code
  • Working knowledge of Azure DevOps or similar CI/CD tooling
  • Practical experience with container/cluster vulnerability management and remediation workflows
  • Scripting skills in Bash and Python
  • Excellent written and verbal communication skills; able to convey technical issues clearly to both technical and non-technical stakeholders
  • Ability to participate in an on-call rotation may be required

Benefits

  • Remote Work Environment
  • Flexible Time Away From Work Policy including PTO, Personal and Sick Days
  • Competitive Salary and Health/Medical Benefits
  • RRSP/TFSA/401K Employee Contribution
  • Life and Disability
  • Employee Assistance Program
  • FHIR Study Program and Skillsoft Learning
  • Super HAPI Fun Club
  • Workplace accommodations during interviews or while working at Smile

Job type

Full Time

Experience level

Mid levelSenior

Salary

CA$115,000 - CA$130,000 per year

Degree requirement

No Education Requirement

Tech skills

AzureDNSDockerFluxGrafanaKafkaKubernetesNGINXNode.jsPrometheusPythonSQLTerraformVault

Location requirements

RemoteCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.