Staff SRE defining reliability for Yuno’s AI-powered payments infrastructure. Owning AWS architecture, messaging, observability, incident response, and resilience at scale.
Responsibilities
Set the technical direction for reliability across Yuno’s infrastructure, beginning with the AWS platform that provisions, deploys, and manages AI agents at scale
Own the platform reliability strategy, including architectural decisions, reliability measurement, and engineering standards
Define SLO culture, error-budget policy, and incident practices across engineering teams
Design and own durable, reliable asynchronous messaging for inter-service communication
Own cloud infrastructure and automate provisioning with Infrastructure as Code
Ensure the platform scales reliably as transaction volume grows
Build monitoring, tracing, and alerting systems for platform health
Serve as senior escalation point for difficult production incidents
Run blameless postmortems and root-cause analyses that produce permanent fixes
Conduct continuous fault injection and resilience experiments
Mentor senior and mid-level engineers and raise organization-wide reliability standards
Requirements
7+ years of experience
Designed and owned event-driven systems using message queues such as Kafka, NATS, or RabbitMQ
Understanding of at-least-once delivery, consumer groups, dead letters, and backpressure
Experience migrating systems from synchronous to asynchronous communication
Deep AWS experience with EC2, VPC, IAM, S3, and RDS
Strong networking fundamentals
Infrastructure as Code experience with Terraform or Pulumi
Kubernetes and Docker production experience, including container lifecycle, resource limits, health checks, and orchestration at scale
Datadog fluency or equivalent experience with dashboards, monitors, APM, and distributed tracing
Track record defining and operating SLOs, SLIs, and error budgets across services
Hands-on fault injection, game day, or chaos experiment experience using Gremlin, Chaos Mesh, AWS FIS, or similar
Distributed systems debugging experience
Comfortable coding automation and tooling in Go, Python, or similar
Solid SQL and PostgreSQL knowledge
NoSQL experience with MongoDB and Redis, including indexing, replication, and performance tuning
Proven technical leadership, architecture influence across teams, and engineering mentorship
Advanced written and spoken English proficiency
Benefits
Competitive Compensation
Remote Work — you can work from everywhere
Home Office Bonus — a one-time allowance to set up your ideal home office
Work Equipment
Stock Options
Health Plan wherever you are
Flexible Days Off
Language, Professional, and Personal Growth courses
DevOps Manager overseeing releases, enterprise tooling, and incident response for Delta Controls, a building - automation solutions manufacturer. Establishing standards across global product teams and offices.
DevOps Engineer building AWS infrastructure and automated systems for S&P Global’s financial data and technology solutions. Supporting resilient applications through Terraform, CI/CD, containerization, monitoring, and cloud operations.
Staff SRE leading GCP reliability, observability, and infrastructure automation for Calix’s broadband communications platform. Building resilient GKE, Kafka, database, and networking systems.
Senior Azure DevOps advisor governing platform evolution for Alithya, a digital transformation consulting firm. Defining standards, optimizing pipelines, dashboards, integrations, and AI capabilities.
Senior DevOps Engineer building Azure DevOps pipelines, Terraform infrastructure, and deployment automation. Supporting PLATO, a Canadian Indigenous - owned software testing and technology services company, across product and data teams.
Pilote Azure DevOps pour un organisme public de santé québécois en transformation numérique. Gouvernance, intégrations Power BI/Dynamics 365, administration de plateforme et accompagnement Agile.
AWS Cloud Engineer optimizing Amazon Connect and enterprise cloud infrastructure for Miratech, a global IT services and consulting company. Automating operations, reliability, observability, and deployments across AWS platforms.