Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.
Responsibilities
Sustain current enterprise release orchestration operations, integrations, automation, governance, and support activities
Design, develop, and support technical solutions that automate agent deployments and release orchestration processes
Provide technical support for vendor evaluations, proof-of-concepts, and transition to a future enterprise release management platform
Participate in pipeline planning, design, development, and support of release and orchestration processes
Provide input to automation policies, standards, and processes
Promote solution reuse and adoption across teams through demos and coaching
Ensure automations comply with Sun Life Financial security directives and policies
Create automation playbooks and companion documentation
Identify automation opportunities and guide effective implementation
Educate others on automation tools and best practices
Operate in compliance with security and change management directives
Balance platform reliability, operational support, stakeholder engagement, and incident resolution with an SRE mindset
Requirements
3-5 years' experience in automation development, specifically Ansible
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.
Senior Site Reliability Engineer securing AuthZed’s cloud infrastructure and authorization platform, including SpiceDB. Building Kubernetes guardrails, supply - chain security, vulnerability management, and incident response.
Principal SRE leading Cerence’s Site Reliability Engineering function for cloud - native automotive AI systems. Owning reliability strategy, incident escalation, observability, automation, and SLO governance.
Senior Data Scientist developing anomaly detection, dashboards, and alerts for General Motors vehicle reliability. Supporting engineering, quality, warranty, and software teams with production analytics.
DevOps Specialist scaling AWS infrastructure and automating deployments for Portage CyberTech’s digital trust, identity, privacy, and security solutions. Maintaining reliable cloud environments and CI/CD pipelines.
SRE Manager leading platform engineering teams across reliability and developer experience. Improving Lightspeed’s global cloud commerce SaaS platform through scalable infrastructure and automation.