Site Reliability Engineer

Posted 1 hour ago

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental-property software products. Leading incident response, observability, automation, and infrastructure hardening.

Responsibilities

  • Act as first responder for production alerts and incidents across services, from triage through resolution
  • Diagnose and fix issues directly in AWS and Kubernetes, including failing pods, resource exhaustion, bad deployments, networking/DNS, database, and cache problems
  • Roll back, scale, reconfigure, or patch infrastructure to restore service quickly
  • Escalate to development teams only when a code change is needed, providing a clear diagnosis
  • Own PagerDuty setup and business-hours incident response while reducing MTTD and MTTR
  • Run blameless post-mortems and drive technical follow-up work
  • Automate runbooks and repetitive operational work, including AI-assisted triage, investigation, and remediation
  • Build and maintain monitoring for Kubernetes workloads and services using Prometheus/Mimir, Loki, Tempo, Grafana, and OpenTelemetry
  • Monitor production releases and identify regressions in latency, errors, or resource usage
  • Create and maintain synthetic checks, smoke tests, health checks, and load/performance tests
  • Partner with engineering teams on performance and reliability issues and define SLOs, SLIs, and error budgets
  • Harden the platform through Terraform changes, Kubernetes resource tuning, autoscaling, CI/CD checks, secrets, and IAM improvements
  • Maintain service documentation and architecture decisions

Requirements

  • 3+ years in a cloud engineering, DevOps, or SRE role supporting production web applications
  • Hands-on production incident response experience, including diagnosing and fixing issues directly
  • Strong hands-on AWS production experience with EKS, EC2, RDS, VPC networking, IAM, and CloudWatch
  • Deep production Kubernetes experience, including troubleshooting, debugging, and monitoring workloads
  • Experience identifying and resolving performance and reliability issues with engineering teams
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Loki, Datadog, or CloudWatch
  • Experience with on-call and alerting tools such as PagerDuty
  • Experience building automated production reliability tests or checks, including synthetic, smoke, health, or load tests
  • Willingness to participate in a future after-hours on-call rotation
  • Infrastructure as code with Terraform and comfort working in CI/CD pipelines
  • Solid Linux, networking, and container fundamentals
  • Scripting/automation in Bash, Python, or similar
  • Calm, clear communication during incidents and across teams
  • Preferred: AI tools for SRE work; Azure and possibly GCP; multiple technology stacks; AWS certification; LGTM or OpenTelemetry at scale; k6, Locust, or JMeter; MySQL/PostgreSQL operations; Redis/Memcached tuning; Cloudflare; cloud cost optimization and capacity planning

Benefits

  • Remote position
  • Preferential consideration may be given to individuals within a reasonable commuting distance of one of the offices
  • Equal opportunity employer
  • Unique accommodations available during the interview process
  • Potential criminal background check in the final interview phase

Job type

Full Time

Experience level

Mid levelSenior

Salary

CA$80,000 - CA$110,000 per year

Degree requirement

No Education Requirement

Tech skills

AWSAzureCloudDNSEC2Google Cloud PlatformGrafanaJMeterKubernetesLinuxMySQLPostgresPrometheusPythonRedisTerraform

Location requirements

RemoteCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.