Senior SRE engineering resilient, self-healing systems for Tubi, a free streaming service serving over 100 million monthly users. Automating infrastructure, observability, and incident response.
Responsibilities
Design, build, and maintain scalable, highly available, and fault-tolerant distributed systems
Partner with development teams as a reliability consultant and influence architectural decisions
Write code to automate operational tasks and CI/CD pipelines
Build internal tools, libraries, and frameworks for self-service observability
Participate in a 24/7 on-call rotation and act as incident commander during critical disruptions
Conduct blameless root cause analyses and implement corrective actions
Monitor, measure, and optimize system performance, latency, and capacity
Forecast capacity needs using usage patterns and historical data
Build and integrate AIOps solutions, including automated responses and self-healing systems
Use AI-assisted coding tools such as Claude Code and Cursor
Develop and document runbooks and procedural guides for the observability knowledge base
Analyze telemetry data, build predictive capacity models, and identify bottlenecks and failure modes
Requirements
Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience
5+ years of professional experience in a Site Reliability Engineering, DevOps, or Software Engineering role focused on infrastructure and operations
Strong programming proficiency in one or more high-level languages such as Rust, Go, Python, or Typescript
Comfortable writing, testing, and deploying production-grade code
Deep knowledge of AWS services, especially networking, IAM, EKS, ALBs/NLBs, Route 53, and CloudWatch
Proven experience with Kubernetes in production, including service exposure, networking, and availability engineering
Solid understanding of Linux/Unix operating systems, TCP/IP, DNS, HTTP, and modern distributed systems architecture
Benefits
Annual discretionary bonus
Long-term incentive plan
Medical, dental, and vision benefits
Insurance
Flexible Time Off Policy
Generous Parental Leave Program: twelve (12) weeks of paid bonding leave in Canada within the first year of birth, adoption, surrogacy, or foster placement, in addition to applicable government leave programs and FOX’s short-term disability policy (if applicable)
DevOps Engineer leading Linux, cloud, and big - data infrastructure delivery for TD Securities’ capital - markets business. Automating deployments, strengthening service levels, and supporting secure technology solutions.
Senior DevOps Engineer scaling AWS infrastructure for Sureify’s insurance SaaS platform. Automating deployments, reliability, monitoring, and security across the Americas.
Build & Release Support Engineer supporting Maven, Docker, Jenkins, Artifactory, and Spinnaker CI/CD systems. Unblocking developers for Virtasant, a global cloud, data, and engineering services company.
Maintenance Reliability Engineer optimizing asset reliability, availability, and maintenance costs at Tenaris, a global manufacturer of advanced tubular energy products. Supporting steel mill investments, reliability strategies, and complex maintenance solutions.
AI Security & DevOps Engineer securing AI systems from prototype to production for Casper Studios, an AI services firm. Owning security standards, platform controls, and enterprise client approvals.
Senior DevOps Engineer automating cloud and on - premise infrastructure for Keyfactor’s cryptographic identity security platform. Improving Kubernetes, CI/CD, and Infrastructure as Code delivery.
Student Site Reliability Analyst supporting Sun Life’s financial - services technology reliability and monitoring. Building dashboards, analyzing performance, and helping prevent outages across global 24x7 services.
DevOps co - op interns implementing software fixes, cloud configurations, and automated deployments for Canadian insurer Intact. Roles based in Toronto, Vancouver, and Montréal.