Staff Software Engineer building Cerebras’s production ML inference APIs. Enabling reliable model serving across GPU and Cerebras AI accelerator backends.
Responsibilities
Design, implement, and maintain production ML inference APIs for chat completions, text generation, streaming, model configuration, tool calling, structured outputs, multimodal inputs, and emerging inference capabilities
Create consistent request and response semantics across GPU prefill, Cerebras decode, and other heterogeneous inference backends
Integrate emerging foundation models, tokenizers, prompt formats, sampling methods, attention variants, multimodal inputs, and model-specific features
Maintain compatibility with widely adopted inference interfaces while designing Cerebras-specific extensions
Establish versioning, deprecation, validation, and backward compatibility practices
Extend and integrate custom inference services with vLLM, PyTorch, Hugging Face libraries, AMD ROCm, and Cerebras runtime components
Build control and data paths coordinating GPU prefill with Cerebras decode, including request routing, state transfer, error handling, retries, and lifecycle management
Optimize streaming, time to first token, latency, throughput, batching, serialization, tokenization, scheduling, and communication between serving components
Build validation systems for tokenization, sampling, logits, generated outputs, precision changes, model upgrades, determinism, and backend compatibility
Define service indicators and build logging, tracing, metrics, dashboards, health checks, and diagnostic tooling
Build configuration, SDKs, documentation, examples, debugging tools, and self-service workflows
Partner with compiler, runtime, kernel, cloud, product, and solutions teams and work directly with customers
Requirements
5+ years of software engineering experience, including substantial individual-contributor ownership of production software or distributed systems
Strong programming ability in Python and Go
Experience developing performance-sensitive or highly concurrent services in C++, Rust, or a similar systems language
Experience building stable APIs with validation, error handling, observability, compatibility, and versioning practices
Experience integrating software across service, framework, runtime, and infrastructure boundaries
Experience designing or maintaining OpenAI-compatible, gRPC, REST, or streaming inference APIs
Experience with Linux, containers, Kubernetes or comparable orchestration systems, CI/CD, and operating latency-sensitive production services
Ability to diagnose correctness, reliability, and performance issues across distributed serving systems
Strong communication and cross-functional execution skills
Bachelor's degree in computer science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience
Preferred: experience with vLLM, SGLang, PyTorch, Hugging Face Transformers, Triton, TensorRT-LLM, or another open-source ML systems project
Preferred: experience creating API conformance, model-quality, numerical-comparison, determinism, or performance-regression test systems
Preferred: experience building SDKs, developer tools, model registries, configuration systems, or self-service ML platforms
Preferred: experience with multi-model or multi-tenant inference platforms, routing, admission control, fairness, quotas, rate limiting, and capacity-aware scheduling
Preferred: understanding of model-specific tokenization, chat templates, generation configuration, logits processing, stopping criteria, tool calling, structured generation, and constrained decoding
Preferred: experience with disaggregated prefill/decode architectures, KV-cache transfer, prefix caching, chunked prefill, memory-aware admission control, or request scheduling
Preferred: experience designing, building, or operating production APIs and services for machine learning, large language models, or other data-intensive applications
Preferred: familiarity with BF16, FP8, FP4, INT8, or INT4
Benefits
Build a breakthrough AI platform beyond the constraints of the GPU
Publish and open source cutting-edge AI research
Work on one of the fastest AI supercomputers in the world
Job stability with startup vitality
Simple, non-corporate work culture that respects individual beliefs
Full Stack Liferay Developer building secure B2B customer portals for McKesson Canada’s healthcare business. Developing Liferay, Java, APIs, integrations, and production support solutions.
Sr. Full Stack Liferay Developer building secure Liferay customer portals for McKesson Canada. Developing Java full - stack solutions, integrations, and production support.
Senior Mobile Software Engineer building React Native apps and Node.js services for Manulife’s retirement technology organization. Driving architecture, observability, CI/CD, and full - stack delivery.
Cybersecurity and software engineering specialists debugging code, assessing vulnerabilities, and optimizing backend systems. Improving high - quality datasets for Gramian Consultancy’s AI training project.
Senior software developer building SaaS applications for service businesses. Leading architecture, automated testing, AWS/Terraform development, and end - to - end delivery.
Software Developer building secure Dynamics 365, Power Platform, and Azure solutions for Genetec’s digital ecosystem. Developing integrations, APIs, Copilot agents, and DevOps automation.
Embedded Software Developer building high - performance C/Linux networking systems at Fortinet, which secures enterprise, service provider, and government organizations. Leading engineers and driving reliable system architecture.
Full - Stack Software Engineer building conversational AI and cloud - native applications for Manulife’s financial services customers. Developing C#, .NET, React, Azure, and microservices solutions.
Senior Software Developer building AI - enabled microservices, APIs, and agentic coding tools at GoTo. Supporting a remote - work technology platform used by millions worldwide.
Senior II Software Engineer building Akamai Cloud infrastructure and testing automation. Developing Python backends, cloud environments, and observability tools for distributed engineering teams.