Staff SRE defining reliability for Yuno’s AI-powered payments infrastructure. Owning AWS architecture, messaging, observability, incident response, and resilience at scale.
Responsibilities
Set the technical direction for reliability across Yuno’s infrastructure, beginning with the AWS platform that provisions, deploys, and manages AI agents at scale
Own the platform reliability strategy, including architectural decisions, reliability measurement, and engineering standards
Define SLO culture, error-budget policy, and incident practices across engineering teams
Design and own durable, reliable asynchronous messaging for inter-service communication
Own cloud infrastructure and automate provisioning with Infrastructure as Code
Ensure the platform scales reliably as transaction volume grows
Build monitoring, tracing, and alerting systems for platform health
Serve as senior escalation point for difficult production incidents
Run blameless postmortems and root-cause analyses that produce permanent fixes
Conduct continuous fault injection and resilience experiments
Mentor senior and mid-level engineers and raise organization-wide reliability standards
Requirements
7+ years of experience
Designed and owned event-driven systems using message queues such as Kafka, NATS, or RabbitMQ
Understanding of at-least-once delivery, consumer groups, dead letters, and backpressure
Experience migrating systems from synchronous to asynchronous communication
Deep AWS experience with EC2, VPC, IAM, S3, and RDS
Strong networking fundamentals
Infrastructure as Code experience with Terraform or Pulumi
Kubernetes and Docker production experience, including container lifecycle, resource limits, health checks, and orchestration at scale
Datadog fluency or equivalent experience with dashboards, monitors, APM, and distributed tracing
Track record defining and operating SLOs, SLIs, and error budgets across services
Hands-on fault injection, game day, or chaos experiment experience using Gremlin, Chaos Mesh, AWS FIS, or similar
Distributed systems debugging experience
Comfortable coding automation and tooling in Go, Python, or similar
Solid SQL and PostgreSQL knowledge
NoSQL experience with MongoDB and Redis, including indexing, replication, and performance tuning
Proven technical leadership, architecture influence across teams, and engineering mentorship
Advanced written and spoken English proficiency
Benefits
Competitive Compensation
Remote Work — you can work from everywhere
Home Office Bonus — a one-time allowance to set up your ideal home office
Work Equipment
Stock Options
Health Plan wherever you are
Flexible Days Off
Language, Professional, and Personal Growth courses
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.