Inference runtime engineer optimizing vLLM, the open-source AI inference engine. Improving LLM and diffusion-model serving across hardware and architectures.
Responsibilities
Push the boundaries of LLM and diffusion model serving
Work at the core of vLLM to optimize model execution across diverse hardware and architectures
Develop inference runtime innovations for mixture-of-experts, multimodal, and agentic architectures
Implement inference techniques and model architectures from research papers
Contribute performant and maintainable code
Debug complex machine-learning codebases
Directly improve how AI inference is run by making inference cheaper and faster
Requirements
Bachelor's degree or equivalent experience in computer science, engineering, or similar
Deep understanding of transformer architectures and their variants
Strong programming skills in Python with experience in PyTorch internals
Experience with LLM inference systems such as vLLM, TensorRT-LLM, SGLang, or TGI
Ability to read and implement model architectures and inference techniques from research papers
Ability to contribute performant and maintainable code and debug in complex ML codebases
Preferred: deep understanding of KV-cache memory management, prefix caching, and hybrid model serving
Preferred: familiarity with RL frameworks and algorithms for LLMs
Preferred: experience with multimodal inference across audio, image, video, and text
Contributions to open-source ML or system infrastructure projects are preferred
Bonus: core feature implementation in vLLM or other inference engine projects
Bonus: contributions to vLLM integrations such as verl, OpenRLHF, Unsloth, or LlamaFactory
Bonus: widely-shared technical blogs or side projects on vLLM or LLM inference
Required application materials: resume, GitHub handle, and a link to a relevant personal project, open-source contribution, or technical blog post
Benefits
Competitive benefits appropriate to your location, including health coverage where applicable
Equity
Visa sponsorship on a case-by-case basis
Fully remote work
Timezone-flexible schedule with regular overlap with Pacific Time for critical syncs
Spécialiste en ingénierie des matériaux de forgeage pour moteurs aéronautiques chez Pratt & Whitney Canada. Conseil technique sur matériaux, traitements thermiques, revêtements et fournisseurs.
Directeur de l’ingénierie dirigeant la conception mécanique et l’équipe d’une organisation québécoise de solutions techniques pour environnements critiques. Pilotage de la conformité, de l’innovation et de la stratégie d’ingénierie.
IBM MQ developer configuring, migrating, and supporting enterprise applications for Genpact’s AI and digital transformation solutions. Handling incidents, releases, disaster recovery, and service performance.
Engineering Specialist designing piping, mechanical systems, and equipment improvements for Tenaris steel manufacturing operations. Supporting engineering execution and project management at the Sault Ste Marie mill.
Desjardins programmer analyst developing, testing, securing, and optimizing complex IT systems. Advising teams and coordinating technology solutions across business and infrastructure operations.
Senior Data Developer building MaintainX’s Databricks data platform for industrial operations. Developing pipelines, governance, observability, and developer tooling for product and analytics teams.
Associate .NET Software Engineer building and modernizing Sun Life applications. Delivering secure, reliable technology supporting clients’ financial security and wellbeing.
Software Engineer building and modernizing .NET applications for Sun Life’s financial security and insurance business. Owning secure, full - cycle solutions across Agile development, testing, and production support.
Software Engineer building and modernizing .NET applications at Sun Life, a financial services and insurance company. Driving secure, high - quality solutions across Agile development, testing, and production support.
Firmware Developer building C/C++ embedded systems for ASSA ABLOY’s intelligent hotel lighting and guest room controls. Integrating protocols, debugging hardware, and ensuring safety and reliability.