Exceptional generalist engineers for AI inference engine development, optimizing CUDA kernels and designing distributed systems. Fully remote opportunity with a focus on autonomy.
Responsibilities
This is a globally remote opportunity.
We're seeking exceptional generalist engineers who can work across the entire vLLM stack: from low-level GPU kernels to high-level distributed systems.
This role is designed for self-directed, autonomous individuals who can identify the highest-leverage problems and solve them end-to-end without constant guidance.
You'll work asynchronously with our San Francisco headquarters while maintaining full ownership of critical infrastructure.
You might be optimizing CUDA kernels one week, designing distributed orchestration systems the next, and implementing new model architectures the week after.
The work you do will directly impact how the world runs AI inference.
Potential focus areas include:
- Inference Runtime: Push the boundaries of LLM and diffusion model serving.
- Kernel Engineering: Write the low-level kernels and optimizations.
- Performance & Scale: Build distributed systems that power inference at global scale.
- Cloud Orchestration: Build the operational backbone for cluster management, deployment automation, and production monitoring.
Requirements
Bachelor's degree or equivalent experience in computer science, engineering, or similar
Demonstrated ability to work autonomously and drive projects to completion without close supervision
Excellent asynchronous communication skills and ability to collaborate effectively across time zones
Strong track record of shipping high-impact work in complex technical environments
Deep expertise in at least one of: systems programming, GPU/accelerator programming, distributed systems, or ML infrastructure
Technical Depth (strong in at least two):
- CUDA kernels or equivalent (Triton, TileLang, Pallas) with deep understanding of GPU architecture
- High-performance distributed systems in Rust, Go, or C++
- Python with PyTorch internals and LLM inference systems (vLLM, TensorRT-LLM, SGLang)
- Kubernetes, container orchestration, and infrastructure-as-code at scale
- Transformer architectures, KV-cache memory management, and model serving
Preferred Qualifications:
- Contributions to vLLM or other major open-source ML/systems projects
- Experience with multiple accelerator platforms (NVIDIA, AMD, TPU, Intel)
- Knowledge of quantization techniques, ML-specific kernel optimization, or compiler technologies
- Track record of improving system reliability and performance at scale
- Written widely-shared technical blogs or impactful side projects in the ML infrastructure space.
Benefits
Inferact offers competitive benefits appropriate to your location, including health coverage where applicable.
Aircraft systems installation engineer designing and integrating aerospace equipment for SOGECLAIR. Coordinating multidisciplinary engineering, certification, configuration, and production support.
Spécialiste en ingénierie des matériaux de forgeage pour moteurs aéronautiques chez Pratt & Whitney Canada. Conseil technique sur matériaux, traitements thermiques, revêtements et fournisseurs.
Directeur de l’ingénierie dirigeant la conception mécanique et l’équipe d’une organisation québécoise de solutions techniques pour environnements critiques. Pilotage de la conformité, de l’innovation et de la stratégie d’ingénierie.
IBM MQ developer configuring, migrating, and supporting enterprise applications for Genpact’s AI and digital transformation solutions. Handling incidents, releases, disaster recovery, and service performance.
Inference runtime engineer optimizing vLLM, the open - source AI inference engine. Improving LLM and diffusion - model serving across hardware and architectures.
Engineering Specialist designing piping, mechanical systems, and equipment improvements for Tenaris steel manufacturing operations. Supporting engineering execution and project management at the Sault Ste Marie mill.
Senior Data Developer building MaintainX’s Databricks data platform for industrial operations. Developing pipelines, governance, observability, and developer tooling for product and analytics teams.
Desjardins programmer analyst developing, testing, securing, and optimizing complex IT systems. Advising teams and coordinating technology solutions across business and infrastructure operations.
Associate .NET Software Engineer building and modernizing Sun Life applications. Delivering secure, reliable technology supporting clients’ financial security and wellbeing.
Software Engineer building and modernizing .NET applications for Sun Life’s financial security and insurance business. Owning secure, full - cycle solutions across Agile development, testing, and production support.