Logotype for Cerebras Systems Inc

Cerebras Systems (CBRS) Supernova 2026 summary

Event summary combining transcript, slides, and related documents.

Logotype for Cerebras Systems Inc

Supernova 2026 summary

19 Aug, 2026

Key announcements and product roadmap

  • Announced CS-4, a next-generation system delivering up to 2x faster tokens and 10x more tokens per watt, with a modular Nexus rack-scale platform for rapid deployment and future scalability.

  • Roadmap targets doubling speed every year and achieving 20x more throughput by 2027, with CS-5 and CS-6 planned on the same modular platform.

  • CS-4 features a new Wafer-Scale Engine (WSE-3 Turbo), three wafers per system, and 43 PBps memory bandwidth per wafer, enabling up to 30x faster inference than GPUs.

  • Disaggregated inference partnerships with AMD and AWS allow optimal hardware use for prefill and decode, maximizing speed and throughput.

  • Early access for CS-4 is available now, with general availability later this quarter.

Industry partnerships and customer impact

  • Collaboration with OpenAI enables GPT-5.6 Sol Ultrafast mode, delivering up to 14x faster inference and removing the speed-intelligence trade-off.

  • Partnerships with Figma and Cognition showcase agentic applications in design and coding, leveraging Cerebras hardware for real-time, high-quality results.

  • CrowdStrike partnership highlights the critical role of inference speed in cybersecurity, enabling faster, more effective threat detection and response.

  • Arista partnership focuses on networking innovations to support hyperscale AI clusters, emphasizing the need for high bandwidth, low latency, and robust recovery.

  • AMD partnership enables disaggregated solutions, combining GPU throughput with Cerebras decode speed for optimal economics and user experience.

Technology and performance insights

  • Wafer-Scale Engine architecture eliminates memory bandwidth bottlenecks, allowing all model weights to reside on-chip and enabling ultra-fast inference.

  • CS-4’s modular design supports rapid manufacturing, deployment, and future upgrades in power, compute, and I/O.

  • New wafer I/O module delivers 2x bandwidth, 2x lower latency, and direct wafer-to-wafer links, supporting models up to 10 trillion parameters at 1,000 tokens/sec.

  • Disaggregated inference splits prefill (GPU) and decode (Cerebras), reducing latency and maximizing throughput for large models.

  • Real-world benchmarks show CS-4 achieving up to 30x faster inference than GPUs across a range of models, transforming user experience and agentic workflows.

Partial view of Summaries dataset, powered by Quartr API
AI can get things wrong. Verify important information.
All investor relations material. One API.
Learn more