Site Reliability / Gitops Engineer supporting and maintaining Canonical’s IT production services. Automating operations with Infrastructure as Code for private and public cloud environments.
Responsibilities
Apply your experience of IaC to develop infrastructure as code practice within IS by constantly increasing automation and improving IaC processes
Automate software operations for re-usability and consistency across private and public clouds, taking into consideration the complexities of distributed systems
Develop new features and improve the resilience and scalability of the existing cloud and container portfolio at Canonical
Maintain operational responsibility for all of Canonical’s core services, networks, and infrastructure
Develop skills in troubleshooting, capacity planning, and performance investigation, Setting up, maintaining and using observability tools such as Prometheus, Grafana, and Elasticsearch; design, implement and maintain monitoring and alerting for various systems and services
Collaborate with development teams to design service architecture, documentation, playbooks, policies and operational procedures
Provide assistance and work with globally distributed engineering, operations, and support peers
Be given uninterrupted development time to focus on larger projects and automation of manual tasks
Share your experience, know-how and best practices with other team members in design sessions, mentorship and ‘doing work together’
Carry final responsibility for time-critical escalations.
Requirements
A deep experience of, and knowledge to define operations in code, using version control, peer review and CI/CD to roll out changes both to applications and infrastructure
Strong modern engineering background (peer-review, unit testing, SCM, CI/CD, Agile)
Python software development experience, with large projects
Practical knowledge of Linux networking, routing, and firewalls
Affinity with various forms of Linux storage, from Ceph to Databases
Hands-on experience administering enterprise Linux servers
Extensive knowledge of cloud computing concepts and technologies
Bachelor's degree or greater, preferably in computer science or related engineering field
Able to communicate clearly and effectively in English over email, chat, video or voice calls and in-person
Motivated and able to troubleshoot from kernel to web, and willing to ask others when appropriate
A willingness to be flexible and able to learn new things quickly
Be inspired by the needs of fast-changing environments
Happy to work within distributed teams
Be passionate and familiarized about open-source, especially Ubuntu or Debian
Benefits
Canonical is an equal opportunity employer
We are proud to foster a workplace free from discrimination.
Diversity of experience, perspectives, and background create a better work environment and better products.
Platform DevOps managing the Enterprise Data and AI Platform across AWS and Kubernetes. Implementing Infrastructure as Code with Terraform and maintaining CI/CD pipelines for secure solutions.
Lead DevOps specialized in AWS/GCP Cloud solutions for FinOps team. Driving cross - functional activation and managing cloud environments, data integrations, and automation strategies.
Skilled DevOps Engineer providing expertise in deployment automation for TD's technology solutions team. Engaging in improving development and release processes while ensuring security and system integrity.
Ingénieur fiabilité des infrastructures pour soutenir les services SaaS critiques. Collaborer, innover et optimiser la fiabilité et la performance des systèmes cloud sur AWS et Kubernetes.
DevOps Engineer to help scale cloud and on - prem environments, automating deployments and enhancing security posture for energy - intelligent compute applications.
Reliability Engineering Architect at Carbon60 managing a team to deliver AWS cloud solutions. Focus on mentoring engineers and integrating AI tools into automated systems.
DevOps Specialist taking over build, release, and environments for Sparrow’s product team. Leading DevOps practices while collaborating with CTO and senior developers in an agile setting.
Developer Advocate advocating for security in cloud native infrastructure within a global leader in recruitment. Collaborating with thought leaders and driving awareness through various channels.
Senior Site Reliability Engineer at Rootly embedding with teams to enhance service performance and reliability. Own CI/CD pipelines and drive capacity planning efforts in a fast - paced environment.
Site Reliability Engineer maintaining and optimizing cloud infrastructure for Tecsys. Collaborating with engineering teams to drive reliability and performance in mission - critical SaaS environments.