Infrastructure team lead building and operating Spare’s GCP and Kubernetes platform for on-demand transit. Leading developers while improving reliability, security, AI SRE, and cloud cost efficiency.
Responsibilities
Lead the Infrastructure Platform Team responsible for Spare’s core cloud infrastructure platform
Own design and development of infrastructure platform capabilities from inception to launch
Build and evolve tooling, automation, and platform services for engineering teams
Architect and implement scalable distributed systems on GCP and Kubernetes
Improve cluster reliability, application resilience, and internal access security
Operate and maintain Redis and PostgreSQL databases, including tuning, scaling, upgrades, backups, and disaster recovery
Drive AI SRE practices for incident detection, alert triage, and operational workflows
Manage and improve the SRE on-call rotation, escalation paths, and blameless post-mortems
Drive FinOps, cloud spend visibility, right-sizing, and cost optimization
Use AI agentic tooling daily and coach the team in its use
Mentor software developers and increase team capacity
Collaborate with product managers, designers, and software developers
Ensure 99.99% uptime and exceptional system performance
Participate in agile rituals and improve software development processes
Split time approximately 50/50 between hands-on technical contribution and people leadership
Travel to up to four customer site visits per year and attend biannual Vancouver hackathons
Manage four direct-report software developers
Requirements
7+ years of software development experience, with at least 2+ years in a people leadership role
Expert backend technology and strong distributed systems experience
Proficiency with AI-assisted and agentic development workflows
Experience operating systems at scale with a strong reliability and uptime mindset
Experience running or managing an SRE on-call rotation, including incident response and post-mortem culture
Deep experience with GCP and Kubernetes
Experience driving cloud cost optimization and FinOps initiatives
Experience with infrastructure-as-code and configuration management tooling, especially Terraform
Understanding of security best practices, internal access control, and container security
Demonstrated success managing software developers and individual/team performance
Demonstrated ability to mentor developers and provide technical leadership
Strong problem-solving, debugging, and system design skills
Excellent communication and collaboration skills
Legally entitled to work in Canada without additional sponsorship
Able to work hybrid three days per week in downtown Vancouver
Nice-to-have: transit or safety-critical domain experience; internal developer platforms; AI/LLM tooling for SRE; CI/CD at scale; PostgreSQL and Redis production operations
Benefits
Equity options
Competitive salary
Opportunity to work on challenging technical problems with real-world impact
Fast-paced, high-impact role in a rapidly growing startup
Ownership of core systems and opportunity to drive innovation
Dynamic, collaborative, and supportive team culture
Participation in biannual software development hackathons in Vancouver
Senior AI Platform Developer building scalable AI services and LLM workflows for MaintainX’s industrial work execution platform. Improving reliability, observability, performance, and cost efficiency.
AI Platform Developer building Azure - based agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Creating secure, observable, governed AI services for product teams.
Senior AI Platform Developer building reusable, secure AI - agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Driving Azure - based platform capabilities across orchestration, evaluation, observability, and governance.
Kubernetes/DevOps Engineer operating Kubernetes infrastructure for a petabyte - scale social media platform. Building distributed systems and machine - learning workloads for Capgemini Engineering’s client.
Lead AI Platform Engineer securing and operating EQ Bank’s enterprise AI platforms. Driving Azure engineering, observability, automation, governance, and production readiness.
Senior Platform Engineer building and operating GitLab Orbit, a Rust - based knowledge graph service for GitLab’s DevSecOps platform. Improving distributed - system reliability, observability, cloud infrastructure, and data workflows.
Platform Software Engineer supporting Invoice Simple, EverCommerce’s invoicing SaaS for small businesses. Building cloud infrastructure, CI/CD pipelines, and reliable production systems.
Software Engineer building cloud infrastructure and deployment systems for Invoice Simple, EverCommerce’s invoicing platform for small businesses. Improving reliability, observability, and production operations.
Senior cybersecurity engineer securing L3Harris’s mission - critical platform - management systems for military and cruise ships. Hardening Linux, networks, software, and infrastructure across the product lifecycle.