Staff Software Engineer building production inference APIs for Cerebras Systems’ AI-chip platform. Integrating models, runtimes, and distributed serving systems for fast, reliable inference.
Responsibilities
Design, implement, and maintain production ML inference APIs for chat completions, text generation, streaming, model configuration, tool calling, structured outputs, and multimodal inputs
Create consistent request and response semantics across GPU prefill, Cerebras decode, and other heterogeneous inference backends
Integrate emerging foundation models, tokenizers, prompt formats, sampling methods, attention variants, multimodal inputs, and model-specific features
Maintain inference-interface compatibility while designing Cerebras-specific extensions, versioning, deprecation, validation, and backward-compatibility practices
Extend and integrate custom inference services with vLLM, PyTorch, Hugging Face libraries, AMD ROCm, and Cerebras runtime components
Build control and data paths coordinating GPU prefill with Cerebras decode, including request routing, state transfer, error handling, retries, and lifecycle management
Optimize streaming, time to first token, latency, throughput, batching, serialization, tokenization, scheduling, and component communication
Build validation systems for tokenization, sampling, logits, generated outputs, precision changes, model upgrades, determinism, and backend compatibility
Define service indicators and build structured logging, tracing, metrics, dashboards, health checks, and diagnostic tooling
Build configuration, SDKs, documentation, examples, debugging tools, and self-service workflows
Partner with compiler, runtime, kernel, cloud, product, and solutions teams to translate requirements into scalable serving capabilities
Requirements
5+ years of software engineering experience, including substantial individual-contributor ownership of production software or distributed systems
Strong programming ability in Python
Experience developing performance-sensitive or highly concurrent services in C++, Go, or a similar systems language
Hands-on experience with a model-serving framework such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face Text Generation Inference, or an equivalent platform
Understanding of modern LLM inference concepts, including tokenization, prompt formatting, sampling, streaming generation, continuous batching, KV-cache management, and model configuration
Experience integrating software across service, framework, runtime, and infrastructure boundaries
Experience building stable APIs with validation, error handling, observability, compatibility, and versioning practices
Experience with Linux, containers, Kubernetes or comparable orchestration systems, CI/CD, and operating latency-sensitive services in production
Ability to diagnose correctness, reliability, and performance issues across distributed serving systems
Bachelor’s degree in computer science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience
Preferred: experience with OpenAI-compatible, gRPC, REST, or streaming inference APIs
Preferred: experience modifying or contributing to vLLM, SGLang, PyTorch, Hugging Face Transformers, Triton, TensorRT-LLM, or another open-source ML systems project
Preferred: experience enabling transformer, Mixture-of-Experts, diffusion, embedding, reranking, or multimodal model architectures
Preferred: understanding of model-specific tokenization, chat templates, generation configuration, logits processing, stopping criteria, tool calling, structured generation, and constrained decoding
Preferred: experience with disaggregated prefill/decode architectures, KV-cache transfer, prefix caching, chunked prefill, memory-aware admission control, or request scheduling
Preferred: experience designing, building, or operating production APIs and services for machine learning, large language models, or other data-intensive applications
Benefits
Build a breakthrough AI platform beyond the constraints of the GPU
Publish and open source cutting-edge AI research
Work on one of the fastest AI supercomputers in the world
Job stability with startup vitality
Simple, non-corporate work culture that respects individual beliefs
Full Stack Developer building automation tools, APIs, and intelligence platforms for PwC Canada. Supporting cybersecurity analysis and globally scaling client - facing services.
Backend Software Engineer building identity decisioning systems and KYC - compliant services for Affirm’s buy - now - pay - later platform. Developing distributed systems with Python or Kotlin across financial product teams.
GTM Engineer building automation, data pipelines, and AI workflows for LumiQ’s accounting and finance professional education platform. Bridging technical systems with Sales and Marketing execution.
Senior Software Engineer building distributed clearing and settlement systems for Alpaca, a global brokerage infrastructure provider. Developing integrations and infrastructure for equities, options, and securities finance products.
Staff Software Engineer designing mission, behavior, and motion planners for AeroVect’s autonomous airport ground equipment. Driving production autonomy through trajectory optimization and robotics planning.
Senior Software Developer leading scalable software application development, mentoring engineers, and aligning technical solutions with business needs. Improving delivery through coding standards, DevOps, automation, and CI/CD.
Tech Lead directing streaming, storage, and training systems for DataVisor’s AI - powered fraud and risk platform. Leading distributed systems execution and AI - assisted engineering practices.
Fullstack Developer criando landing pages, integrações e tracking para campanhas de performance da IPmedia, startup de encontros online. Foco em conversão, velocidade, CRO e escala.