SRE Specialist enhancing system service availability and performance for Morgan Stanley's technology. Collaborating with engineering teams and identifying opportunities for automation and reliability improvements in Montreal.
Responsibilities
Working closely with engineering/development teams to design, build, and maintain systems
Troubleshoot issues across the entire stack: hardware, software, application and network
Identifying and drive opportunities to improve automation for our platforms
Proactively identifying and addressing systems reliability risks
Represent the RPE organization in design reviews and operational readiness exercises for new and existing services
Participate in on-call rotation and periodic conference calls with other specialists from other time zones
Requirements
At least 4 years of experience in a SRE role
Background in Computer Science equivalent to a B.Sc.
Automation-related experience is particularly valued using scripting languages such as Python, Bash, Perl
One higher level language is desired
Experience on supporting three tier architecture which includes exposure to UNIX, Linux platforms and databases such DB2, Sybase or relational databases like MongoDB
Experience with source code and binary repositories, build tools, and CI/CD (Git, Artifactory, Jenkins, Docker) etc. and data streaming technologies like Spark, Kafka
Hands on experience on enterprise tools set such as Grafana, Prometheus, Dynatrace, AppDynamics
Awareness of modern software & systems architectures, including load-balancing, queueing, caching, distributed systems failure modes, micro services
Deep understanding of operating system level concepts such as processes, memory allocation, and the network stack; understanding of how applications are affected by the above, and ability to debug same
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.