Senior Site Reliability Engineer securing AuthZed’s cloud infrastructure and authorization platform, including SpiceDB. Building Kubernetes guardrails, supply-chain security, vulnerability management, and incident response.
Responsibilities
Operate, build, and evolve Kubernetes and cloud infrastructure powering AuthZed's managed services
Advance infrastructure security across Kubernetes, cloud infrastructure, networking, IAM, secrets management, encryption, and service-to-service communication
Strengthen security across the software supply chain, including CI/CD, dependencies, container images, build artifacts, provenance, vulnerability scanning, and deployment controls
Develop automated guardrails, policies, and secure defaults for the engineering platform
Improve vulnerability identification, prioritization, and remediation across containers, dependencies, infrastructure, and cloud services
Expand security telemetry, logging, alerting, detection capabilities, and incident-response playbooks
Participate in the SRE on-call rotation and lead or contribute to production incidents
Collaborate on architecture and infrastructure design and contribute to threat modeling
Significant experience operating production systems as a Site Reliability Engineer, Platform Engineer, Infrastructure Engineer, or Security Engineer
Deep hands-on experience with Kubernetes and containerized production environments
Strong experience with at least one major cloud provider (AWS, GCP, Azure) and its security model
Strong understanding of IAM, workload identity, networking, TLS/PKI, encryption, secrets management, and cloud-native security
Experience with Infrastructure as Code; Pulumi, Terraform, or AWS CDK
Experience designing, operating, and securing CI/CD and software delivery pipelines
Strong understanding of Linux and container security fundamentals
Experience with observability, monitoring, logging, and production incident response
Strong software engineering skills and an inclination to solve operational and security challenges through automation
Strong security judgment and ability to evaluate risk in real-world distributed systems
Ability to operate independently, drive projects across engineering boundaries, and provide technical leadership in a small, highly technical organization
Experience operating distributed databases or other stateful distributed systems in production is an advantage
Experience securing multi-tenant SaaS or cloud infrastructure is an advantage
Experience contributing to security architecture, threat modeling, penetration testing, or security assessments is an advantage
Relevant certifications such as CKA, CKS, AWS Certified Security – Specialty, CISSP, or CCSP are an advantage
Legally able to work in the U.S. or Canada
Benefits
Competitive salary based on experience and location
Stock options at an early-stage startup
Comprehensive benefits including healthcare (US-based) and other insurance
A full remote and flexible schedule to accommodate different time zones
Twice-yearly travel for team off-sites focused on team bonding, collaboration, and having fun!
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.
Principal SRE leading Cerence’s Site Reliability Engineering function for cloud - native automotive AI systems. Owning reliability strategy, incident escalation, observability, automation, and SLO governance.
Senior Data Scientist developing anomaly detection, dashboards, and alerts for General Motors vehicle reliability. Supporting engineering, quality, warranty, and software teams with production analytics.
DevOps Specialist scaling AWS infrastructure and automating deployments for Portage CyberTech’s digital trust, identity, privacy, and security solutions. Maintaining reliable cloud environments and CI/CD pipelines.
SRE Manager leading platform engineering teams across reliability and developer experience. Improving Lightspeed’s global cloud commerce SaaS platform through scalable infrastructure and automation.
SRE Manager leading platform engineering, reliability, and Developer Experience for Lightspeed’s global cloud POS SaaS platform. Guiding teams, infrastructure automation, and strategic delivery.
Gestionnaire SRE dirigeant deux équipes de plateforme chez Lightspeed. Assurant la fiabilité et l’automatisation de sa plateforme SaaS mondiale de commerce infonuagique.
DevOps Engineer managing AWS, Kubernetes, CI/CD, and observability for StellarTech’s scalable global EdTech products. Improving reliability, security, and production operations across multiple products.