
  <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
    <channel>
      <title>Gimlet Blog</title>
      <link>https://gimletlabs.ai/blog</link>
      <description>A blog about research on high performance AI systems.</description>
      <language>en-US</language>
      <managingEditor>hello@gimletlabs.ai (Gimlet Labs Inc.)</managingEditor>
      <webMaster>hello@gimletlabs.ai (Gimlet Labs Inc.)</webMaster>
      <lastBuildDate>Mon, 20 Oct 2025 00:00:00 GMT</lastBuildDate>
      <atom:link href="https://gimletlabs.ai/blog/tags/efficiency/feed.xml" rel="self" type="application/rss+xml"/>
      
  <item>
    <guid>https://gimletlabs.ai/blog/heterogeneous-ai-infrastructure</guid>
    <title>Designing infrastructure for running efficient AI workloads</title>
    <link>https://gimletlabs.ai/blog/heterogeneous-ai-infrastructure</link>
    <description>AI workloads are shifting from simple LLM inference to complex, multi-model workflows. To run them efficiently at scale, we need a system that can dynamically decompose workloads, plan and schedule them, and map execution to the right hardware.</description>
    <pubDate>Mon, 20 Oct 2025 00:00:00 GMT</pubDate>
    <author>hello@gimletlabs.ai (Gimlet Labs Inc.)</author>
    <category>Inference</category><category>Performance</category><category>Efficiency</category>
  </item>

  <item>
    <guid>https://gimletlabs.ai/blog/multivendor-prefill-decode-disaggregation</guid>
    <title>Splitting LLM inference across different hardware platforms</title>
    <link>https://gimletlabs.ai/blog/multivendor-prefill-decode-disaggregation</link>
    <description>Separating prefill and decode stages of LLM inference improves token throughput because their resource needs differ. Although most deployments use NVIDIA hardware for both stages, multivendor disaggregation can actually improve efficiency while maintaining SLAs. Based on our models using NVIDIA B200s and Intel Gaudi 3, common workloads can see 1.7X TCO improvement compared to single-vendor disaggregation.</description>
    <pubDate>Mon, 13 Oct 2025 00:00:00 GMT</pubDate>
    <author>hello@gimletlabs.ai (Gimlet Labs Inc.)</author>
    <category>Inference</category><category>Performance</category><category>Efficiency</category><category>Hardware</category>
  </item>

    </channel>
  </rss>
