Site Reliability Engineer

Posted 2 days ago

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Staff SRE defining reliability for Yuno’s AI-powered payments infrastructure. Owning AWS architecture, messaging, observability, incident response, and resilience at scale.

Responsibilities

  • Set the technical direction for reliability across Yuno’s infrastructure, beginning with the AWS platform that provisions, deploys, and manages AI agents at scale
  • Own the platform reliability strategy, including architectural decisions, reliability measurement, and engineering standards
  • Define SLO culture, error-budget policy, and incident practices across engineering teams
  • Design and own durable, reliable asynchronous messaging for inter-service communication
  • Own cloud infrastructure and automate provisioning with Infrastructure as Code
  • Ensure the platform scales reliably as transaction volume grows
  • Build monitoring, tracing, and alerting systems for platform health
  • Serve as senior escalation point for difficult production incidents
  • Run blameless postmortems and root-cause analyses that produce permanent fixes
  • Conduct continuous fault injection and resilience experiments
  • Mentor senior and mid-level engineers and raise organization-wide reliability standards

Requirements

  • 7+ years of experience
  • Designed and owned event-driven systems using message queues such as Kafka, NATS, or RabbitMQ
  • Understanding of at-least-once delivery, consumer groups, dead letters, and backpressure
  • Experience migrating systems from synchronous to asynchronous communication
  • Deep AWS experience with EC2, VPC, IAM, S3, and RDS
  • Strong networking fundamentals
  • Infrastructure as Code experience with Terraform or Pulumi
  • Kubernetes and Docker production experience, including container lifecycle, resource limits, health checks, and orchestration at scale
  • Datadog fluency or equivalent experience with dashboards, monitors, APM, and distributed tracing
  • Track record defining and operating SLOs, SLIs, and error budgets across services
  • Hands-on fault injection, game day, or chaos experiment experience using Gremlin, Chaos Mesh, AWS FIS, or similar
  • Distributed systems debugging experience
  • Comfortable coding automation and tooling in Go, Python, or similar
  • Solid SQL and PostgreSQL knowledge
  • NoSQL experience with MongoDB and Redis, including indexing, replication, and performance tuning
  • Proven technical leadership, architecture influence across teams, and engineering mentorship
  • Advanced written and spoken English proficiency

Benefits

  • Competitive Compensation
  • Remote Work — you can work from everywhere
  • Home Office Bonus — a one-time allowance to set up your ideal home office
  • Work Equipment
  • Stock Options
  • Health Plan wherever you are
  • Flexible Days Off
  • Language, Professional, and Personal Growth courses

Job type

Full Time

Experience level

SeniorLead

Salary

Not specified

Degree requirement

No Education Requirement

Tech skills

AWSCloudDistributed SystemsDockerEC2KafkaKubernetesMongoDBNoSQLPostgresPythonRabbitMQRedisSQLTerraformGo

Location requirements

RemoteWorldwide

Report this job

Found something wrong with the page? Please let us know by submitting a report below.