Member of Technical Staff, Inference

Posted 4 days ago

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Inference runtime engineer optimizing vLLM, the open-source AI inference engine. Improving LLM and diffusion-model serving across hardware and architectures.

Responsibilities

  • Push the boundaries of LLM and diffusion model serving
  • Work at the core of vLLM to optimize model execution across diverse hardware and architectures
  • Develop inference runtime innovations for mixture-of-experts, multimodal, and agentic architectures
  • Implement inference techniques and model architectures from research papers
  • Contribute performant and maintainable code
  • Debug complex machine-learning codebases
  • Directly improve how AI inference is run by making inference cheaper and faster

Requirements

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar
  • Deep understanding of transformer architectures and their variants
  • Strong programming skills in Python with experience in PyTorch internals
  • Experience with LLM inference systems such as vLLM, TensorRT-LLM, SGLang, or TGI
  • Ability to read and implement model architectures and inference techniques from research papers
  • Ability to contribute performant and maintainable code and debug in complex ML codebases
  • Preferred: deep understanding of KV-cache memory management, prefix caching, and hybrid model serving
  • Preferred: familiarity with RL frameworks and algorithms for LLMs
  • Preferred: experience with multimodal inference across audio, image, video, and text
  • Contributions to open-source ML or system infrastructure projects are preferred
  • Bonus: core feature implementation in vLLM or other inference engine projects
  • Bonus: contributions to vLLM integrations such as verl, OpenRLHF, Unsloth, or LlamaFactory
  • Bonus: widely-shared technical blogs or side projects on vLLM or LLM inference
  • Required application materials: resume, GitHub handle, and a link to a relevant personal project, open-source contribution, or technical blog post

Benefits

  • Competitive benefits appropriate to your location, including health coverage where applicable
  • Equity
  • Visa sponsorship on a case-by-case basis
  • Fully remote work
  • Timezone-flexible schedule with regular overlap with Pacific Time for critical syncs

Job type

Full Time

Experience level

Lead

Salary

Not specified

Degree requirement

Bachelor's Degree

Tech skills

PythonPyTorch

Location requirements

RemoteWorldwide

Report this job

Found something wrong with the page? Please let us know by submitting a report below.