Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Senior Systems Administrator managing digital research infrastructure at the University of Alberta. Ensuring effective use of high-performance computing resources across various disciplines.

Responsibilities

  • Participate in the deployment and administration of digital research infrastructure. Plan, design, deploy, and provision digital research infrastructure and associated services, ensuring security, reliability, and scalability for future growth.
  • Take a leadership role in establishing policies, best practices, and procedures to address the needs of research communities.
  • Address operational issues affecting systems and researchers, conducting in-depth root cause analysis.
  • Monitor resources to maximize the utilization of infrastructure, collect metrics, and develop key performance indicators (KPIs) to provide regular reports to senior leadership.
  • Assist researchers in resolving issues and effectively using the digital research infrastructure. This includes a range of services such as workload preparation, code optimization, training, and job submission procedure refinement, all aimed at ensuring the successful utilization of high-performance computing resources to meet research goals.
  • Participate in focus groups and committees, providing expert-level knowledge and gathering input.
  • Participate in industry conferences and maintain relationships with vendors to stay informed about emerging technologies and improve service quality.
  • Participate in a 24/7 on-call support rotation.

Requirements

  • Master’s Degree in Science with a focus on High Performance Computing, or a Bachelor of Science in Computing Science or a related field, with relevant work experience.
  • Minimum of 8 years of extensive experience with Linux systems in roles such as system administrator or system analyst.
  • Advanced working knowledge of administering large numbers of systems in an enterprise environment.
  • Advanced working knowledge of automation technologies for provisioning, configuration management, infrastructure as code, and orchestration tools (Warewulf, Ansible, Puppet, Chef, etc.).
  • Advanced working knowledge of programming and/or scripting with languages such as Bash, Python, and others.
  • Working experience in supporting large-scale High Performance Computing environments.
  • Working knowledge of scheduling software, such as Slurm, is a strong asset.
  • Working knowledge of container deployment (Kubernetes or other) in the cloud or on-premises, as well as the administration of containerized applications.
  • Demonstrated ability to participate in a 24/7 on-call support rotation.
  • Effective communication and facilitation skills.
  • Demonstrated ability in problem-solving and process optimization.

Benefits

  • Comprehensive benefits package

Job type

Full Time

Experience level

Senior

Salary

CA$94,748 - CA$134,132 per year

Degree requirement

Postgraduate Degree

Tech skills

AnsibleChefCloudKubernetesLinuxPuppetPython

Location requirements

HybridEdmontonCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.