Technical Lead owning Tether’s bare-metal GPU, Slurm, Kubernetes, and inference platform architecture. Leading a distributed infrastructure team supporting Tether’s global digital-finance products.
Responsibilities
Own the end-to-end platform architecture through architecture proposals, high-level and low-level designs, reviews, and maintenance of the baseline
Lead and line-manage a distributed team of about twelve engineers across backend, frontend, DevOps, QA, and documentation
Establish engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input
Design, build, and operate a managed Slurm service for research users
Own controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity, and isolation
Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal
Implement NVIDIA GPU Operator and Network Operator, VM-based GPU isolation with KubeVirt and VFIO, upgrades, backup and recovery, and node replacement
Own managed inference architecture, including serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity
Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers
Lead incident response, post-incident reviews, and development of a sustainable on-call model
Serve as the primary technical interface to infrastructure partners and vendors
Translate requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing
Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity
Complete the platform team and set the technical bar for new engineers
Requirements
Eight or more years of hands-on engineering experience
At least three years leading teams that build and operate infrastructure platforms other teams depend on
Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience
Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system
Experience operating an HPC or GPU training cluster for a research population is ideally preferred
Experience operating NVIDIA GPU fleets on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance
Experience with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems
Deep Linux systems knowledge, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning
Production Kubernetes operations experience, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design
Experience with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets
Experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review
Working fluency in JavaScript and Node.js sufficient to review control-plane, CLI, and worker services and make architecture decisions
Experience shipping a platform with real users, such as a multi-tenant IaaS/PaaS or research computing service
People management across time zones, cross-track review, written architecture decisions, and partner/executive communication
Excellent written and spoken English
Based between UTC and UTC+5:30
Desirable: Slurm operators on Kubernetes or Kubernetes-native schedulers
Desirable: modern serving stacks such as vLLM, SGLang, and TensorRT-LLM
Principal Backend Software Developer building scalable Java and AWS services for Autodesk’s design - tool collaboration platform. Mentoring developers and improving secure, reliable cloud infrastructure.
Remote software engineers licensing owned Git repositories for AI training. Earning revenue through non - exclusive code licenses while retaining intellectual property ownership.
Senior iOS Engineer owning Swift architecture and releases for Xsolla’s gaming commerce apps. Building AI - powered features and integrating payments, virtual currency, and mini - apps.
Software engineering co - op developing internal data - management tools for Kardium’s FDA - approved atrial - fibrillation medical device. Coding, testing, and integrating systems in a regulated environment.
Software Developer Co - op building software for Kardium’s atrial - fibrillation medical device systems. Automating data, monitoring systems, testing, and deploying regulated changes.
Staff Software Developer building reliable Windows, Mac, and Linux endpoint security agents for Arctic Wolf. Leading architecture, technical delivery, mentoring, and cross - functional cybersecurity solutions.
Principal Engineer shaping Bluehost’s Python, Linux, MySQL, Perl, and cPanel hosting platform. Driving reliability, modernization, AI delivery standards, and technical direction for small - business infrastructure.
Lead engineer building AI - powered digital products for Sun Life, a global financial services company. Owning architecture, production code, experimentation, and end - to - end product delivery.
Senior Software Engineer II designing Boomi integrations across Salesforce, SAP, and Adobe platforms for Thomson Reuters. Developing APIs, troubleshooting enterprise workflows, and mentoring engineers.