Senior Site Reliability Engineer, SRE

Posted 2 days ago

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Senior SRE engineering resilient, self-healing systems for Tubi, a free streaming service serving over 100 million monthly users. Automating infrastructure, observability, and incident response.

Responsibilities

  • Design, build, and maintain scalable, highly available, and fault-tolerant distributed systems
  • Partner with development teams as a reliability consultant and influence architectural decisions
  • Write code to automate operational tasks and CI/CD pipelines
  • Build internal tools, libraries, and frameworks for self-service observability
  • Participate in a 24/7 on-call rotation and act as incident commander during critical disruptions
  • Conduct blameless root cause analyses and implement corrective actions
  • Monitor, measure, and optimize system performance, latency, and capacity
  • Forecast capacity needs using usage patterns and historical data
  • Build and integrate AIOps solutions, including automated responses and self-healing systems
  • Use AI-assisted coding tools such as Claude Code and Cursor
  • Develop and document runbooks and procedural guides for the observability knowledge base
  • Analyze telemetry data, build predictive capacity models, and identify bottlenecks and failure modes

Requirements

  • Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience
  • 5+ years of professional experience in a Site Reliability Engineering, DevOps, or Software Engineering role focused on infrastructure and operations
  • Strong programming proficiency in one or more high-level languages such as Rust, Go, Python, or Typescript
  • Comfortable writing, testing, and deploying production-grade code
  • Deep knowledge of AWS services, especially networking, IAM, EKS, ALBs/NLBs, Route 53, and CloudWatch
  • Proven experience with Kubernetes in production, including service exposure, networking, and availability engineering
  • Solid understanding of Linux/Unix operating systems, TCP/IP, DNS, HTTP, and modern distributed systems architecture

Benefits

  • Annual discretionary bonus
  • Long-term incentive plan
  • Medical, dental, and vision benefits
  • Insurance
  • Flexible Time Off Policy
  • Generous Parental Leave Program: twelve (12) weeks of paid bonding leave in Canada within the first year of birth, adoption, surrogacy, or foster placement, in addition to applicable government leave programs and FOX’s short-term disability policy (if applicable)
  • Monthly wellness reimbursement

Job type

Full Time

Experience level

Senior

Salary

CA$116,000 - CA$235,100 per year

Degree requirement

Bachelor's Degree

Tech skills

AWSDistributed SystemsDNSKubernetesLinuxPythonRustTCP/IPTypeScriptUnixGo

Location requirements

HybridTorontoCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.