Staff Software Engineer, Inference API

Posted 4 hours ago

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Staff Software Engineer building Cerebras’s production ML inference APIs. Enabling reliable model serving across GPU and Cerebras AI accelerator backends.

Responsibilities

  • Design, implement, and maintain production ML inference APIs for chat completions, text generation, streaming, model configuration, tool calling, structured outputs, multimodal inputs, and emerging inference capabilities
  • Create consistent request and response semantics across GPU prefill, Cerebras decode, and other heterogeneous inference backends
  • Integrate emerging foundation models, tokenizers, prompt formats, sampling methods, attention variants, multimodal inputs, and model-specific features
  • Maintain compatibility with widely adopted inference interfaces while designing Cerebras-specific extensions
  • Establish versioning, deprecation, validation, and backward compatibility practices
  • Extend and integrate custom inference services with vLLM, PyTorch, Hugging Face libraries, AMD ROCm, and Cerebras runtime components
  • Build control and data paths coordinating GPU prefill with Cerebras decode, including request routing, state transfer, error handling, retries, and lifecycle management
  • Optimize streaming, time to first token, latency, throughput, batching, serialization, tokenization, scheduling, and communication between serving components
  • Build validation systems for tokenization, sampling, logits, generated outputs, precision changes, model upgrades, determinism, and backend compatibility
  • Define service indicators and build logging, tracing, metrics, dashboards, health checks, and diagnostic tooling
  • Create conformance tests, workload-replay tools, model-validation suites, performance benchmarks, integration tests, and release gates
  • Build configuration, SDKs, documentation, examples, debugging tools, and self-service workflows
  • Partner with compiler, runtime, kernel, cloud, product, and solutions teams and work directly with customers

Requirements

  • 5+ years of software engineering experience, including substantial individual-contributor ownership of production software or distributed systems
  • Strong programming ability in Python and Go
  • Experience developing performance-sensitive or highly concurrent services in C++, Rust, or a similar systems language
  • Experience building stable APIs with validation, error handling, observability, compatibility, and versioning practices
  • Experience integrating software across service, framework, runtime, and infrastructure boundaries
  • Experience designing or maintaining OpenAI-compatible, gRPC, REST, or streaming inference APIs
  • Experience with Linux, containers, Kubernetes or comparable orchestration systems, CI/CD, and operating latency-sensitive production services
  • Ability to diagnose correctness, reliability, and performance issues across distributed serving systems
  • Strong communication and cross-functional execution skills
  • Bachelor's degree in computer science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience
  • Preferred: experience with vLLM, SGLang, PyTorch, Hugging Face Transformers, Triton, TensorRT-LLM, or another open-source ML systems project
  • Preferred: experience creating API conformance, model-quality, numerical-comparison, determinism, or performance-regression test systems
  • Preferred: experience building SDKs, developer tools, model registries, configuration systems, or self-service ML platforms
  • Preferred: experience with multi-model or multi-tenant inference platforms, routing, admission control, fairness, quotas, rate limiting, and capacity-aware scheduling
  • Preferred: understanding of model-specific tokenization, chat templates, generation configuration, logits processing, stopping criteria, tool calling, structured generation, and constrained decoding
  • Preferred: experience with disaggregated prefill/decode architectures, KV-cache transfer, prefix caching, chunked prefill, memory-aware admission control, or request scheduling
  • Preferred: experience designing, building, or operating production APIs and services for machine learning, large language models, or other data-intensive applications
  • Preferred: familiarity with BF16, FP8, FP4, INT8, or INT4

Benefits

  • Build a breakthrough AI platform beyond the constraints of the GPU
  • Publish and open source cutting-edge AI research
  • Work on one of the fastest AI supercomputers in the world
  • Job stability with startup vitality
  • Simple, non-corporate work culture that respects individual beliefs
  • Continuous learning, growth and support

Job type

Full Time

Experience level

Lead

Salary

Not specified

Degree requirement

Bachelor's Degree

Tech skills

CloudDistributed SystemsGRPCKubernetesLinuxPythonPyTorchRustC++Go

Location requirements

HybridTorontoCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.