Senior Site Reliability Engineer ensuring reliability and performance of Vantage’s services while collaborating across teams. Engaging in incident response and driving infrastructure improvements.
Responsibilities
Collaborate with a diverse team of software engineers, engaging in iterative processes and effective task planning to drive our projects forward.
Take ownership of the availability, scalability, and performance of our services, to proactively identify issues, and implement automation to prevent the recurrence of problems.
Participate in the on-call rotation, responding to incidents and working with the team to restore service and prevent recurrence.
Contribute to automating infrastructure provisioning, configuration, and management using IaC principles with tools like Terragrunt and Ansible.
Help design and enhance monitoring, logging, and alerting systems to improve observability and ensure system health.
Participate in blameless post-mortems, documenting issues, and following up on action items to foster a culture of learning and continuous improvement.
Foster collaboration with other engineering teams, promoting the reuse of existing frameworks and gaining insights into their operation.
Stay current with industry trends, emerging technologies, and best practices in SRE, DevOps, and automation.
Requirements
6+ years of experience as a Site Reliability Engineer, DevOps Engineer, or similar role working with software and infrastructure.
Proficiency with either Python or Bash.
Hands-on experience with Azure or AWS.
Familiarity with CI/CD pipelines and infrastructure as code (IaC) and its tooling such as terraform and ansible.
Demonstrated ability to triage and prioritize effectively when troubleshooting incidents.
History of engaging effectively with cross-functional teams during events such as incident-response and post-mortems.
Track-record of proactively tailoring infrastructure to meet the unique needs of the product it supports.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.