Senior Manager leading Incident Response Engineering for Confluent Cloud while ensuring customer-first incident management. Building and evolving a team to handle incidents at scale across cloud platforms.
Responsibilities
Build and Lead the Team
Recruit, hire, and develop a team of senior incident response engineers distributed across AMER and APAC time zones
Design sustainable on-call models with follow-the-sun coverage
Own Incident Response
Provide incident command for high-severity and critical customer-impacting incidents, with your team as the primary rotation and you as the senior escalation point
Set and enforce standards for how incidents are run: communications cadence, directing engagements with stakeholders, domain expert coordination, handoffs
Drive a customer-first posture in every incident to ensure timely, accurate updates and clear ownership from detection through resolution
Drive Postmortem Rigor and Customer RCA Quality
Own postmortem quality end-to-end: facilitation, root cause analysis, corrective action definition, and ensuring follow-through
Manage the Customer Root Cause Analysis (CRCA) program, ensuring timely, technically accurate, clearly written documents that restore customer trust
Coordinate upstream technical inputs from engineering teams; synthesize ambiguity into clear, actionable narratives
Advance Incident Response Through AI and Automation
Drive an AI-centric approach to scaling incident operations using intelligent tooling to improve triage speed, documentation quality, and pattern detection without sacrificing rigor
Partner with observability, supportability, and resiliency sub-functions with CAR to provide critical inputs into our platform evolution
Own and evolve the incident management tooling stack with a bias towards agentic assistance
Analyze incident data to identify recurring patterns and feed learnings back into engineering practices
When incident load allows, direct your team's capacity toward runbook improvements, automation, and operational hygiene
Represent Cross-Functionally
Partner with Legal, PR, and Customer Success on customer-facing communications during and after major incidents
Brief engineering leadership and executives during active incidents with clarity and composure
Be the person engineering teams proactively seek out when operational standards and incident practices need to improve
Requirements
10+ years in SRE, incident management, or reliability engineering, with at least 5 years managing teams in this space
Proven experience as an incident commander in high-severity, customer-impacting outages at scale. You've personally run incidents that mattered
Cloud infrastructure experience across at least one of AWS, GCP, or Azure
Deep understanding of distributed systems failure modes (Kafka/event streaming experience preferred, or demonstrated ability to rapidly master complex systems)
Strong track record with postmortem facilitation and driving corrective actions to completion
Excellent written communication with customers regarding root-cause analysis. You are comfortable stating things with conviction to executive audiences
Experience working with cross-functional stakeholders (legal, PR, customer success) during incident response
Track record of hiring and developing senior technical talent in a globally distributed, remote-first environment
Comfort operating with significant autonomy and making high-stakes decisions under pressure.
Data Security Specialist protecting Sun Life’s financial - services data through DLP, CASB and insider - threat programs. Investigating cyber risks and advancing enterprise data protection.
Senior SaaS Security Manager protecting RBC’s banking platform from third - party cloud risks. Leading controls, vulnerability management, compliance, and security transformation initiatives.
Développeur.euse sécurité cloud protégeant l’infrastructure de nesto, plateforme de financement hypothécaire canadienne. Conception de contrôles cloud, automatisation DevSecOps et réponse aux incidents.
Lead SCADA and cybersecurity engineer designing compliant electric - substation systems for GE Vernova. Coordinating multidisciplinary teams, vendors, testing, estimates, and project risk for decarbonized energy infrastructure.
Join RBC's Application Security Group to develop innovative security solutions, mentor junior staff, and collaborate across teams to enhance decision - making and automate tasks.
Lead SCADA and cybersecurity engineer designing compliant substation systems for GE Vernova. Guiding project teams, vendor designs, estimates, and acceptance testing for cleaner energy infrastructure.
Senior security advisor simulating cyber threats and strengthening defenses for Desjardins, North America's largest cooperative financial group. Leading complex initiatives, methodologies and cybersecurity risk mitigation.
SA&A Lead securing Azure applications and Microsoft platforms for PLATO, Canada’s Indigenous - owned software testing company. Leading authorization, control testing, evidence collection, and risk remediation.
Senior security advisor simulating and mitigating cyberthreats for Desjardins, North America's largest cooperative financial group. Leading offensive security methodologies, tools and strategic initiatives.