Technical Lead – GPU Infrastructure

Posted 2 days ago

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Technical Lead owning Tether’s bare-metal GPU, Slurm, Kubernetes, and inference platform architecture. Leading a distributed infrastructure team supporting Tether’s global digital-finance products.

Responsibilities

  • Own the end-to-end platform architecture through architecture proposals, high-level and low-level designs, reviews, and maintenance of the baseline
  • Lead and line-manage a distributed team of about twelve engineers across backend, frontend, DevOps, QA, and documentation
  • Establish engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input
  • Design, build, and operate a managed Slurm service for research users
  • Own controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity, and isolation
  • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal
  • Implement NVIDIA GPU Operator and Network Operator, VM-based GPU isolation with KubeVirt and VFIO, upgrades, backup and recovery, and node replacement
  • Own managed inference architecture, including serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity
  • Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers
  • Lead incident response, post-incident reviews, and development of a sustainable on-call model
  • Serve as the primary technical interface to infrastructure partners and vendors
  • Translate requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing
  • Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity
  • Complete the platform team and set the technical bar for new engineers

Requirements

  • Eight or more years of hands-on engineering experience
  • At least three years leading teams that build and operate infrastructure platforms other teams depend on
  • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience
  • Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system
  • Experience operating an HPC or GPU training cluster for a research population is ideally preferred
  • Experience operating NVIDIA GPU fleets on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance
  • Experience with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems
  • Deep Linux systems knowledge, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning
  • Production Kubernetes operations experience, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design
  • Experience with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets
  • Experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review
  • Working fluency in JavaScript and Node.js sufficient to review control-plane, CLI, and worker services and make architecture decisions
  • Experience shipping a platform with real users, such as a multi-tenant IaaS/PaaS or research computing service
  • People management across time zones, cross-track review, written architecture decisions, and partner/executive communication
  • Excellent written and spoken English
  • Based between UTC and UTC+5:30
  • Desirable: Slurm operators on Kubernetes or Kubernetes-native schedulers
  • Desirable: modern serving stacks such as vLLM, SGLang, and TensorRT-LLM
  • Desirable: VM/container isolation, confidential computing, Cluster API, kubeadm, Cilium, infrastructure as code, GitOps, GPU cloud/HPC/AI lab experience, distributed systems, and hardware-provider partnership experience

Benefits

  • Fully remote work arrangement
  • Occasional travel to partner sites and team events

Job type

Full Time

Experience level

Senior

Salary

Not specified

Degree requirement

Bachelor's Degree

Tech skills

BootstrapCloudDistributed SystemsGrafanaJavaScriptKubernetesLinuxNFSNode.jsPrometheus

Location requirements

RemoteWorldwide

Report this job

Found something wrong with the page? Please let us know by submitting a report below.