SRE Specialist enhancing system service availability and performance for Morgan Stanley's technology. Collaborating with engineering teams and identifying opportunities for automation and reliability improvements in Montreal.
Responsibilities
Working closely with engineering/development teams to design, build, and maintain systems
Troubleshoot issues across the entire stack: hardware, software, application and network
Identifying and drive opportunities to improve automation for our platforms
Proactively identifying and addressing systems reliability risks
Represent the RPE organization in design reviews and operational readiness exercises for new and existing services
Participate in on-call rotation and periodic conference calls with other specialists from other time zones
Requirements
At least 4 years of experience in a SRE role
Background in Computer Science equivalent to a B.Sc.
Automation-related experience is particularly valued using scripting languages such as Python, Bash, Perl
One higher level language is desired
Experience on supporting three tier architecture which includes exposure to UNIX, Linux platforms and databases such DB2, Sybase or relational databases like MongoDB
Experience with source code and binary repositories, build tools, and CI/CD (Git, Artifactory, Jenkins, Docker) etc. and data streaming technologies like Spark, Kafka
Hands on experience on enterprise tools set such as Grafana, Prometheus, Dynatrace, AppDynamics
Awareness of modern software & systems architectures, including load-balancing, queueing, caching, distributed systems failure modes, micro services
Deep understanding of operating system level concepts such as processes, memory allocation, and the network stack; understanding of how applications are affected by the above, and ability to debug same
Senior DevOps / Cloud Infrastructure Engineer needed for hybrid role in North York, ON. Requires 10+ years experience with GCP, AWS, Kubernetes, Terraform, and CI/CD.
Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.
Mozilla release engineer optimizing Firefox build, test, and deployment pipelines at global scale. Improving developer experience, maintaining automation, and responding to critical service outages without an on - call rotation.