Staff Software Engineer, Inference API

Posted 2 days ago

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Staff Software Engineer building production inference APIs for Cerebras Systems’ AI-chip platform. Integrating models, runtimes, and distributed serving systems for fast, reliable inference.

Responsibilities

  • Design, implement, and maintain production ML inference APIs for chat completions, text generation, streaming, model configuration, tool calling, structured outputs, and multimodal inputs
  • Create consistent request and response semantics across GPU prefill, Cerebras decode, and other heterogeneous inference backends
  • Integrate emerging foundation models, tokenizers, prompt formats, sampling methods, attention variants, multimodal inputs, and model-specific features
  • Maintain inference-interface compatibility while designing Cerebras-specific extensions, versioning, deprecation, validation, and backward-compatibility practices
  • Extend and integrate custom inference services with vLLM, PyTorch, Hugging Face libraries, AMD ROCm, and Cerebras runtime components
  • Build control and data paths coordinating GPU prefill with Cerebras decode, including request routing, state transfer, error handling, retries, and lifecycle management
  • Optimize streaming, time to first token, latency, throughput, batching, serialization, tokenization, scheduling, and component communication
  • Build validation systems for tokenization, sampling, logits, generated outputs, precision changes, model upgrades, determinism, and backend compatibility
  • Define service indicators and build structured logging, tracing, metrics, dashboards, health checks, and diagnostic tooling
  • Create conformance tests, workload-replay tools, model-validation suites, performance benchmarks, integration tests, and release gates
  • Build configuration, SDKs, documentation, examples, debugging tools, and self-service workflows
  • Partner with compiler, runtime, kernel, cloud, product, and solutions teams to translate requirements into scalable serving capabilities

Requirements

  • 5+ years of software engineering experience, including substantial individual-contributor ownership of production software or distributed systems
  • Strong programming ability in Python
  • Experience developing performance-sensitive or highly concurrent services in C++, Go, or a similar systems language
  • Hands-on experience with a model-serving framework such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face Text Generation Inference, or an equivalent platform
  • Understanding of modern LLM inference concepts, including tokenization, prompt formatting, sampling, streaming generation, continuous batching, KV-cache management, and model configuration
  • Experience integrating software across service, framework, runtime, and infrastructure boundaries
  • Experience building stable APIs with validation, error handling, observability, compatibility, and versioning practices
  • Experience with Linux, containers, Kubernetes or comparable orchestration systems, CI/CD, and operating latency-sensitive services in production
  • Ability to diagnose correctness, reliability, and performance issues across distributed serving systems
  • Bachelor’s degree in computer science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience
  • Preferred: experience with OpenAI-compatible, gRPC, REST, or streaming inference APIs
  • Preferred: experience modifying or contributing to vLLM, SGLang, PyTorch, Hugging Face Transformers, Triton, TensorRT-LLM, or another open-source ML systems project
  • Preferred: experience enabling transformer, Mixture-of-Experts, diffusion, embedding, reranking, or multimodal model architectures
  • Preferred: understanding of model-specific tokenization, chat templates, generation configuration, logits processing, stopping criteria, tool calling, structured generation, and constrained decoding
  • Preferred: experience with disaggregated prefill/decode architectures, KV-cache transfer, prefix caching, chunked prefill, memory-aware admission control, or request scheduling
  • Preferred: experience designing, building, or operating production APIs and services for machine learning, large language models, or other data-intensive applications

Benefits

  • Build a breakthrough AI platform beyond the constraints of the GPU
  • Publish and open source cutting-edge AI research
  • Work on one of the fastest AI supercomputers in the world
  • Job stability with startup vitality
  • Simple, non-corporate work culture that respects individual beliefs
  • Continuous learning, growth and support
  • Equal and diverse work environment

Job type

Full Time

Experience level

Lead

Salary

Not specified

Degree requirement

Bachelor's Degree

Tech skills

CloudDistributed SystemsGRPCKubernetesLinuxPythonPyTorchC++Go

Location requirements

HybridTorontoCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.