Published onSeptember 28, 2026Scaling low-latency inference with Gimlet Cloud and CerebrasAnnouncementGimlet LabsCerebrasInferenceGimlet brings Cerebras’ wafer-scale compute into Gimlet Cloud, enabling inference at up to 3,000 tokens per second while expanding the performance and efficiency available for agentic and real-time AI applications.