Site Reliability Engineer automating deployment and operation of software for dental software provider. Collaborating with developers and mentoring peers to streamline processes and improve efficiencies.
Responsibilities
Stewarding Infrastructure as Code (IaC)
Mentoring peers
Shaping the technical direction of the platform
Collaborate closely with developers to support a wide range of applications
Automate repetitive tasks to improve deployment and operation of software
Responding quickly to incidents while prioritizing long-term solutions
Helping create a positive culture that embraces learning and curiosity
Requirements
Bachelor’s degree in Computer Science or equivalent experience
Deep experience with AWS in production environments, including EC2, S3, IAM, EKS, and RDS
Hands-on experience managing Kubernetes clusters and deploying applications
Strong Linux expertise and system-level troubleshooting skills
Solid understanding of system administration, security best practices, and managing mission-critical data
Proven experience monitoring and optimizing large-scale enterprise web applications
Familiarity with infrastructure automation tools such as Ansible, Packer, and AWS CloudFormation
Proficient in at least one programming language and scripting for automation
Strong analytical skills and experience with root cause analysis
Comfortable working in agile environments, including participation in code reviews and testing
Senior DevOps / Cloud Infrastructure Engineer needed for hybrid role in North York, ON. Requires 10+ years experience with GCP, AWS, Kubernetes, Terraform, and CI/CD.
Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.
Mozilla release engineer optimizing Firefox build, test, and deployment pipelines at global scale. Improving developer experience, maintaining automation, and responding to critical service outages without an on - call rotation.