Senior Site Reliability Engineer II focused on AI Native infrastructure at Life360. Responsible for scaling and maintaining backend services across the platform with 40,000+ cores.
Responsibilities
Scaling and maintaining our infrastructure and services using AI (Claude Code) as a first-class collaborator in your daily development workflow.
Being opinionated on technical direction and strategy (and documenting those opinions for others to be able to follow).
Leading and mentoring other engineers on the team
Owning and resolving the most complex infrastructure failures — Kubernetes scheduling edge cases, networking degradation, cross-service cascading failures, and AWS platform issues that other engineers escalate
Participating in a shared on-call rotation (roughly one week every six to eight weeks on call)
Estimating schedules, breaking tasks down to reasonable 1-3 day tasks.
Driving cloud cost efficiency by identifying over-provisioned resources, rightsizing EC2 and container workloads, and building tooling to surface cost anomalies before they compound
Requirements
Bachelor's in Computer Science, Engineering, related field, or equivalent practical experience
Expert-level experience (5+ years) managing medium to large-scale deployments on AWS (~5000 instances, 50+ accounts), or equivalent.
3+ years of experience programming in Java, Python, or other formal programming languages
Strong Kubernetes experience (3+ years) deploying and managing at scale (100s of Deployments,10k+ containers, 20k+ Cores).
Understanding of container orchestration and microservices
Experience with service discovery/service mesh
Strong Linux administration experience, shell/bash scripting.
Expert-level experience with Infrastructure as code tools: Terraform, CloudFormation; config management/provisioning tools: Ansible, Chef, etc.
Strong Build / Automation / CI/CD experience.
Strong Knowledge/experience with networking and load-balancer technologies.
Experience with existing open-source projects such as Consul, Docker, ArgoCD, Nexus, Jenkins
Experience with large-scale Kafka deployments
Database knowledge is a plus.
Excellent troubleshooting skills, expertise with any monitoring tools, and attention to detail
Excellent interpersonal skills and highly collaborative working style
Hands-on experience with AI coding tools (Claude Code, Cursor, or equivalent) used for infrastructure scripting, incident response automation, or tooling development
Benefits
Competitive pay and benefits
Medical, dental, vision, life and disability insurance plans
RRSP plan with DPSP company matching program
Employee Assistance Program (EAP) for mental well-being
Flexible PTO, several company-wide days off throughout the year
Winter and Summer Week-long Synchronized Company Shutdowns
Learning & Development programs
Equipment, tools, and reimbursement support for a productive remote environment
Free Life360 Platinum Membership for your preferred circle
Senior DevOps / Cloud Infrastructure Engineer needed for hybrid role in North York, ON. Requires 10+ years experience with GCP, AWS, Kubernetes, Terraform, and CI/CD.
Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.
Mozilla release engineer optimizing Firefox build, test, and deployment pipelines at global scale. Improving developer experience, maintaining automation, and responding to critical service outages without an on - call rotation.