Scaling low-latency inference with Gimlet Cloud and Cerebras
- Published on
- Authors

- Name
- Zain Asgar
Scaling low-latency inference with Gimlet Cloud and Cerebras
Bringing wafer-scale compute to Gimlet Cloud to deliver AI inference at thousands of tokens per second
The next generation of AI applications will be defined by speed.
As agents take on longer and more complex workflows, inference latency compounds across every model call. A multi-step task that takes 10 minutes at 100 tokens per second may take just 20 seconds at 3,000 tokens per second. That difference changes how much work an AI system can accomplish, how quickly it can iterate, and which experiences developers can build.
Today, we are announcing a strategic partnership with Cerebras to bring wafer-scale compute into Gimlet Cloud.

Figure 1. The Cerebras Wafer-Scale Engine, WSE-3, combines compute with large amounts of on-wafer SRAM and exceptional memory bandwidth, accelerating memory-intensive inference operations.
Together, Gimlet and Cerebras plan to deploy 100 megawatts of Cerebras-powered inference capacity. The first Cerebras-powered Gimlet Cloud datacenter is expected to come online later this year, delivering speeds of up to 3,000 tokens per second for demanding agentic and real-time applications.
This partnership builds on joint customer engagements underway since last year and an integrated solution already serving inference traffic in private deployments. Gimlet will also serve as a launch partner for Cerebras CS-4, with Gimlet Cloud customers expected to gain direct access to Cerebras’ next-generation technology in 2027.
Beyond homogeneous infrastructure
Delivering ultrafast inference at production scale requires a new approach to infrastructure.
Most inference clouds today are built on homogeneous GPU systems. These systems provide world-class aggregate throughput, but applications face a tradeoff between total system throughput and the speed experienced by each user.
Heterogeneous inference accelerators expand that performance frontier. The Cerebras Wafer Scale Engine combines compute with large amounts of on-wafer SRAM and exceptional memory bandwidth. This architecture is particularly well suited to memory-intensive inference operations, including the token-generation phase of large language models, where latency and responsiveness are critical.

Figure 2. The Gimlet multisilicon inference cloud combines GPUs with accelerators including the Cerebras CS-3 / CS-4 systems.
GPUs and wafer-scale systems have different strengths. The opportunity is to combine them within a single inference workload, using GPUs and Cerebras wafer-scale systems where each performs best, and in the process maximize both throughput and latency.
That is what we are building with Gimlet Cloud.
The first multisilicon inference cloud
Gimlet Cloud is a vertically integrated inference platform spanning datacenter infrastructure, systems software, workload orchestration, and developer APIs. Instead of requiring every workload to conform to a single type of processor, Gimlet decomposes inference and maps each phase to the silicon architecture best suited to it.
AI inference is not a single, uniform workload. Prefill, decode, attention, feed-forward layers, and speculative decoding have different compute and memory characteristics. No single architecture is optimal for every phase.
Gimlet uses disaggregation techniques including Prefill-Decode, Attention-FFN, and Speculative Decoding disaggregation to run these phases across different types of silicon. Our software orchestrates the workload across nodes, while presenting developers with a standard inference API.

Figure 3. Gimlet’s multisilicon architecture expands the inference performance frontier across throughput and interactivity. The disaggregation techniques shown are illustrative, composable, and not to scale.
The result is better performance and efficiency than homogeneous infrastructure can provide. Across frontier workloads, Gimlet’s heterogeneous disaggregation techniques can deliver 3-10x higher interactivity at a given throughput-efficiency target. Conversely, they can increase throughput per kilowatt by 3-10x at a given interactivity target.

Figure 4. Homogeneous infrastructure confines applications to a narrower throughput-interactivity tradeoff. Gimlet’s multisilicon disaggregation expands that frontier, delivering greater interactivity across multiple efficiency and throughput targets.
Through this partnership, Cerebras-powered ultrafast inference will become a native, fully supported accelerator within Gimlet Cloud. Combined with Gimlet’s workload orchestration and vertically integrated infrastructure, Cerebras’ performance will be available as part of a broader platform designed to optimize across heterogeneous architectures.
A faster foundation for AI
AI systems are becoming more interactive, more agentic, and more deeply embedded in everyday work. They will generate vastly more tokens, execute longer sequences of actions, and increasingly operate on the critical path of software.
The answer to these challenges isn't just more infrastructure; it's a better architecture. We believe the future of inference is vertically integrated and multisilicon: many architectures, orchestrated as one system and available through one developer experience.
Our partnership with Cerebras is a major step toward that future, and toward making inference at thousands of tokens per second available at production scale.
Learn more or request access here.