Senior Site Reliability Engineer II focused on AI Native infrastructure at Life360. Responsible for scaling and maintaining backend services across the platform with 40,000+ cores.
Responsibilities
Scaling and maintaining our infrastructure and services using AI (Claude Code) as a first-class collaborator in your daily development workflow.
Being opinionated on technical direction and strategy (and documenting those opinions for others to be able to follow).
Leading and mentoring other engineers on the team
Owning and resolving the most complex infrastructure failures — Kubernetes scheduling edge cases, networking degradation, cross-service cascading failures, and AWS platform issues that other engineers escalate
Participating in a shared on-call rotation (roughly one week every six to eight weeks on call)
Estimating schedules, breaking tasks down to reasonable 1-3 day tasks.
Driving cloud cost efficiency by identifying over-provisioned resources, rightsizing EC2 and container workloads, and building tooling to surface cost anomalies before they compound
Requirements
Bachelor's in Computer Science, Engineering, related field, or equivalent practical experience
Expert-level experience (5+ years) managing medium to large-scale deployments on AWS (~5000 instances, 50+ accounts), or equivalent.
3+ years of experience programming in Java, Python, or other formal programming languages
Strong Kubernetes experience (3+ years) deploying and managing at scale (100s of Deployments,10k+ containers, 20k+ Cores).
Understanding of container orchestration and microservices
Experience with service discovery/service mesh
Strong Linux administration experience, shell/bash scripting.
Expert-level experience with Infrastructure as code tools: Terraform, CloudFormation; config management/provisioning tools: Ansible, Chef, etc.
Strong Build / Automation / CI/CD experience.
Strong Knowledge/experience with networking and load-balancer technologies.
Experience with existing open-source projects such as Consul, Docker, ArgoCD, Nexus, Jenkins
Experience with large-scale Kafka deployments
Database knowledge is a plus.
Excellent troubleshooting skills, expertise with any monitoring tools, and attention to detail
Excellent interpersonal skills and highly collaborative working style
Hands-on experience with AI coding tools (Claude Code, Cursor, or equivalent) used for infrastructure scripting, incident response automation, or tooling development
Benefits
Competitive pay and benefits
Medical, dental, vision, life and disability insurance plans
RRSP plan with DPSP company matching program
Employee Assistance Program (EAP) for mental well-being
Flexible PTO, several company-wide days off throughout the year
Winter and Summer Week-long Synchronized Company Shutdowns
Learning & Development programs
Equipment, tools, and reimbursement support for a productive remote environment
Free Life360 Platinum Membership for your preferred circle
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.