Announcing Gimlet's Series B

Published on
Authors

Announcing Gimlet Labs’ Series B

Faster, More Efficient AI Inference at Scale

Today, we are announcing our $300M Series B raise, led by Andreessen Horowitz and joined by Sapphire Ventures, Menlo Ventures, 645 Ventures, Arm, Eclipse, Emergence, Factory, Hudson River Trading, M12, OnePrime Capital, Prosperity7, QuantumLight, Samsung Ventures, Tiger Global Management, Triatomic, Wing Ventures, and XTX Markets.

Since March, we’ve added billions in contracted revenue, gigawatts of datacenter pipeline, and are quickly scaling to hundreds of megawatts in managed capacity.

The Need for Throughput and Speed

Since our Series A just over 5 months ago, the demand for tokens has continued to increase at a breakneck pace. Monthly token generation has increased 6X in 12 months, with some projections stating another 20X increase by 2030.

To support this demand, the industry has embarked on perhaps the most ambitious investment project in modern history, already close to $1T per year and expected to reach a cumulative $7T by 2030. Power is quickly becoming the most critical resource bottleneck. AI data centers consumed approximately 18 GW of capacity in 2025, with that number expected to triple by 2030. Supporting this growth requires new power generation, grid capacity, interconnections, and behind-the-meter energy production. We’ll eventually hit the limits of the resources we can deliver at scale. Maximizing throughput per kW enables the industry to continue scaling while reducing the impact on physical resources.

In addition to raw throughput, we’re also seeing an increasing need for very fast inference tiers and low-latency tokens, driven by:

  • Agentic workloads. Agents execute multi-step inference loops, often requiring dozens of sequential model calls and tool interactions. Latency compounds across these calls.
  • Larger models. Models expanded from a few 100B parameters just a few years ago, reaching 3-10T parameters today.
  • Larger context windows. In 2023, context windows were limited to about 100K tokens. Today, context windows can go to 1M tokens and more.

The industry is able to deliver high throughput or low latency individually. The problem is when these requirements compound. The throughput-interactivity Pareto chart has been widely circulated, and illustrates the tradeoffs that teams make between throughput and latency. The challenge is not only to maximize throughput, or to maximize speed, but it’s to deliver fast tokens at high throughput.

Traditional homogeneous inference infrastructure limits the user to a high-throughput, low interactivity optimum. With heterogeneous disaggregation, we’re able to achieve 3-10X faster performance for frontier workloads in our inference cloud.

Traditional homogeneous inference infrastructure limits the user to a high-throughput, low interactivity optimum. With heterogeneous disaggregation, we’re able to achieve 3-10X faster performance for frontier workloads in our inference cloud.

Delivering Faster, More Efficient Inference with Gimlet Cloud

When we founded Gimlet, we took the following positions:

  1. AI inference will become the majority workload in software (not just in AI).
  2. AI inference is significantly - as in multiple orders of magnitude - less efficient than it can/should be.
  3. Therefore, AI infrastructure will need to be reimagined and rebuilt from the ground up, with inference in mind, to meet the demand of these workloads.

As the industry has moved from training to inference as the dominant workload, the fastest way to serve models has been to repurpose the same infrastructure originally built for training (and in some cases use cases like crypto mining) for inference. However, inference is fundamentally different from training, and even within inference there is a significant variety in the requirements from the underlying hardware.

Fortunately, chip designers have been actively building highly optimized AI accelerators. We now have access to vastly differentiated architectures, each coming with their own unique strengths.

Different chip architectures excel at different tasks and offer performance and efficiency tradeoffs.

Different chip architectures excel at different tasks and offer performance/efficiency tradeoffs.

At Gimlet Labs, we are building the first multisilicon cloud, built from the ground up for inference performance. Gimlet is built on heterogeneous hardware to take advantage of different types of accelerators, including GPUs, near-memory compute, dataflow architectures, and CPUs. And our software stack intelligently disaggregates workloads across these chips to run each phase of the inference workload on its most optimal silicon architecture.

By breaking up the model and running it across different types of accelerators, we are able to achieve 5-10X speedups for the same power footprint, or similar throughput improvements for the same latency. This is critical, because power is the limiting factor in AI deployments today.

Leveraging Heterogeneity with Granular Disaggregation

There are multiple ways of breaking up models and running them across heterogeneous hardware, and different splits offer different performance characteristics. In our software stack, models are traced and decomposed into different parts and scheduled on accelerators based on the workload SLAs and available hardware.

We dynamically rebalance the workload based on the available hardware, so that compute is not stranded or unused. Even if one type of hardware is fully utilized for a given task (such as decode), if the workload needs more decode instances, those can be spun up on other available hardware as needed in order to provide the necessary capacity.

Inference workloads can be disaggregated in multiple ways, each with different performance tradeoffs.

Inference workloads can be disaggregated in multiple ways. Here we show a few common disaggregation schemes, each of which offers different performance tradeoffs.

The most well-known type of disaggregation is prefill/decode, where the prefill phase (which is very compute-bound) runs on one set of devices and the decode phase (which is very memory-bandwidth bound) runs on a different set. However, there are multiple other techniques as well, such as speculative-decode disaggregation and attention-ffn disaggregation. Prefill/decode disaggregation offers the lowest latency, whereas the other types provide a speedup over homogeneous deployments with higher throughput. There are also a few new types that we are working on which we will cover in future technical content.

We’re Hiring!

We’re expanding our team. We’d love to speak with you If you’re interested in joining a highly technical team of researchers, engineers, and operators, working across the full AI infrastructure stack spanning systems, datacenters, compilers, distributed systems, networking, and performance engineering.

We are building Gimlet around small teams, high ownership, and a flat organization. We want people to stay close to the work, collaborate directly, and have the freedom to solve important problems without unnecessary layers or processes.

You can check out our open positions here.

The Future of Inference

Frontier model sizes, token demand, token speed, cache lengths, and chip architectures are all growing rapidly. Software as we know it is changing, and we are thrilled to work with our customers on new product experiences that can be unlocked with a fundamental step change in both total token throughput and model speed.

Today, we are working with frontier labs and other large-scale consumers of inference, but we are ramping quickly to add more capacity to a broader audience. If you would like to learn more about getting access to low-latency inference, please reach out here. Thank you to our investors again for partnering with us on this journey.

Announcing Gimlet's Series B | Gimlet Blog