Senior platform engineer building and operating Virtasant’s production Kubernetes infrastructure. Owning cloud reliability, observability, migrations, and developer delivery platforms across remote Americas teams.
Responsibilities
Design, build, and operate production Kubernetes clusters, including networking, workload isolation, and multi-region topologies
Work with Kubernetes internals, resource quotas, scheduling, NetworkPolicy enforcement, and custom controllers or operators
Implement and operate service mesh capabilities, including mTLS, workload authentication and authorization, and traffic management
Optimize containerized workloads for performance, cost, and resource efficiency
Write, refactor, and maintain production Go, Python, or Java services, controllers, and middleware
Build HTTP, REST, and gRPC interfaces for internal engineering teams
Write unit and integration tests
Lead incident response, investigate root causes, and author postmortems
Troubleshoot production systems using logs, metrics, traces, and profiling tools
Define and drive SLOs and actionable alerting
Own infrastructure as code and maintain reusable modules
Build and improve CI/CD and GitOps delivery workflows
Plan and execute cloud migration initiatives while maintaining reliability and minimizing downtime
Build and maintain metrics, dashboards, alerting policies, and distributed tracing
Partner with product, security, and infrastructure teams on requirements and architecture
Contribute to design reviews, set technical direction, and mentor engineers
Requirements
6+ years of professional experience in software, platform, infrastructure, or site reliability engineering
Significant experience operating production distributed systems
Demonstrated experience building and operating production Kubernetes platforms
Production experience writing Go, Python, or Java
Experience taking systems from ambiguous starting points to production
Experience migrating production cloud services between providers or environments
Degree in Computer Science, Engineering, or a related field, or equivalent practical experience
Strong understanding of Kubernetes internals, including CNI networking, NetworkPolicy, resource management, and cluster behavior under load
Hands-on experience with a service mesh such as Istio, Envoy, or Linkerd, plus mTLS and workload identity
Solid Linux fundamentals, including cgroups and resource management
Infrastructure as code experience at scale with Terraform or equivalent
Production experience with GCP, AWS, or Azure
Experience with Prometheus, Grafana, OpenTelemetry, and query languages such as PromQL
Production experience with relational databases, including PostgreSQL or managed Postgres-compatible services
Docker and container tooling experience
Strong debugging and performance profiling skills
Experience with Go testing frameworks such as Ginkgo and Gomega preferred
Experience building Kubernetes controllers or operators preferred
Additional Python strength preferred
Experience with IAM technologies such as SSO, Keycloak, OIDC, SAML, or Vault preferred
Experience designing high availability and disaster recovery across regions or providers preferred
Experience with monorepos and build systems such as Bazel preferred
Ability to read Java preferred
Exposure to SOC 2 or GDPR preferred
Experience with Alibaba Cloud highly preferred
Experience with Apache Spark and Apache Flink highly preferred
Strong analytical and problem-solving ability
Clear written and verbal communication
Ability to work independently within a distributed team
Comfort working in a fast-moving, highly technical environment
Senior Cloud Engineer designing secure Azure and AWS infrastructure for JLL Technologies’ enterprise real estate data platform. Driving CI/CD, Kubernetes, infrastructure - as - code, governance, and reliability across engineering teams.
Cloud Engineer designing and operating secure AKS platforms for Manulife’s Canadian financial services business. Automating infrastructure, improving reliability, and integrating developer platform capabilities.
Cloud Platform Engineer managing Kubernetes and Docker infrastructure for EQ Bank, Canada’s challenger bank. Supporting cloud operations, automation, monitoring, disaster recovery, and production incidents.
Architecte infonuagique chez TEHORA, firme québécoise de services techniques et de gestion de projets. Conception et gouvernance de plateformes cloud, sécurité, conformité et accompagnement des équipes.
Staff Cloud Architect leading cloud, data, and AI/ML architecture for Robots and Pencils’ applied AI engineering engagements. Designing scalable, secure production systems across US and Canadian remote teams.
Senior Cloud Architect designing cloud - native, data, and AI/ML systems for Robots & Pencils, an applied AI engineering firm. Building production - ready enterprise solutions across US and Canada.
Student Cloud Infrastructure Analyst at Sun Life, helping operate enterprise AI and AWS cloud platforms. Supporting Infrastructure as Code, CI/CD, automation, and DevOps initiatives.
Cloud infrastructure manager leading Azure operations, migrations, governance, and team delivery. Supporting Coastal Community Credit Union’s financial services and local member communities.
Cloud Engineer supporting Sun Life’s financial security and wellness technology infrastructure. Assisting cloud migration, automation, security, and multi - cloud initiatives across hybrid environments.
Staff ML Engineer building production PyTorch training pipelines on AWS Trainium and SageMaker. Delivering cost - aware, hardware - optimized AI systems for enterprise clients.