Senior Site Reliability Engineer with Python infra-as-code for Cloud operations at Canonical. Enabling devsecops for applications on OpenStack and Kubernetes in a remote global environment.
Responsibilities
Bring Python software-engineering skills and rigour to the operations domain
Practise devsecops from bare metal to application
Architect and run OpenStack, Kubernetes and software-defined storage
Enable devsecops for applications running on that infrastructure
Gain experience in a broad range of cloud technologies
Requirements
Degree in Software Engineering or Computer Science
Experience with Linux and familiarity with Linux networking and storage
Python software development expertise
Operational experience
Excellent interpersonal skills, curiosity, flexibility, and accountability
Ability to travel internationally twice a year, for company events up to two weeks long
Experience with OpenStack or Kubernetes deployment or operations (nice-to-have)
Benefits
Distributed work environment with twice-yearly team sprints in person
Personal learning and development budget of USD 2,000 per year
Annual compensation review
Recognition rewards
Annual holiday leave
Maternity and paternity leave
Employee Assistance Programme
Opportunity to travel to new locations to meet colleagues
Priority Pass, and travel upgrades for long haul company events
Senior DevOps / Cloud Infrastructure Engineer needed for hybrid role in North York, ON. Requires 10+ years experience with GCP, AWS, Kubernetes, Terraform, and CI/CD.
Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.
Mozilla release engineer optimizing Firefox build, test, and deployment pipelines at global scale. Improving developer experience, maintaining automation, and responding to critical service outages without an on - call rotation.