ML Performance Benchmarking Engineer at Cerebras Systems, transforming AI inference with cutting-edge technology. Collaborating with teams to optimize performance across the software stack.
Responsibilities
Design and implement end-to-end telemetry systems across the software stack, providing deep visibility into inference performance.
Architect, build, and scale the automation that generates, analyzes, and visualizes performance data.
Dive deep into system behavior, dissect performance bottlenecks, and deliver actionable insights that influence features.
Partner closely with Core Platform teams to define rigorous testing methodologies that validate inference features for peak performance.
Requirements
Bachelor’s or Master’s degree in Computer Engineering, Systems Engineering, or a related field.
Proficiency in Python and/or C++ programming.
Proven experience in building and scaling automated infrastructure.
Strong background in throughput and performance optimization techniques, especially in complex, large-scale systems.
Excellent problem-solving skills and a strong analytical mindset.
Demonstrated ability to dive deep into new domains.
Ability to work in a fast-paced, ambiguous, and collaborative environment.
Familiarity with problem-solving at the intersection of hardware and software.
Hands-on experience with AI workloads and architectures is a plus.
Benefits
Build a breakthrough AI platform beyond the constraints of the GPU.
Publish and open source their cutting-edge AI research.
Work on one of the fastest AI supercomputers in the world.
Enjoy job stability with startup vitality.
Our simple, non-corporate work culture that respects individual beliefs.
New - grad software engineer building Quora’s distributed ML platform, model serving, and developer tooling. Supporting Quora’s global knowledge - sharing product with scalable GPU infrastructure.
Applied Machine Learning Scientist developing Generative AI and predictive ML solutions at TD, a major North American bank. Supporting model evaluation, deployment, monitoring, and responsible AI governance.
Applied Machine Learning Scientist developing Generative AI and predictive ML solutions for TD banking. Evaluating models, managing AI lifecycles, and supporting responsible implementation.
Senior Machine Learning Engineer building conversational AI agents and production ML systems. Helping Numa automate automotive dealership service and sales through evaluation - first tooling and infrastructure.
Senior ML Engineer building production ML, RAG, and agentic AI capabilities for SailPoint’s cloud identity security platform. Driving scalable, customer - focused AI solutions from research to production.
Senior ML Engineer developing debiased pCTR and conversion models for Instacart’s grocery advertising ecosystem. Advancing ranking, retrieval, and sequence modeling across ads surfaces.
Senior AI/ML Engineer building Generative AI, RAG, and agentic solutions for pharmaceutical Statistical Programming. Deploying secure, validated, production - ready AI applications with Python and AWS.
Senior Machine Learning Engineer building LLM - powered lab interpretation and clinical decision - support tools for Fullscript’s healthcare platform. Owning AI systems from prototyping through production.
Staff ML engineer building models, evaluations, and agentic systems for Sourcegraph’s code - understanding products. Improving enterprise code search quality, latency, cost, and reliability.