<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Continuous-Batching on DevOpsTales</title>
    <link>https://devopstales.github.io/tags/continuous-batching/</link>
    <description>Recent content in Continuous-Batching on DevOpsTales</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en-US</language>
    <lastBuildDate>Fri, 02 Oct 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://devopstales.github.io/tags/continuous-batching/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>vLLM Scheduling: How Throughput and Latency Trade Off in Shared Inference</title>
      <link>https://devopstales.github.io/ai/vllm-scheduling-throughput-latency/</link>
      <pubDate>Fri, 02 Oct 2026 00:00:00 +0000</pubDate>
      
      <guid>https://devopstales.github.io/ai/vllm-scheduling-throughput-latency/</guid>
      <description>At 9 in the morning Alice asks the office&rsquo;s local model to summarize yesterday&rsquo;s meeting, and the answer streams back faster than she can read it. By 10 the whole team has found the thing. Same GPU, same model — and now everyone&rsquo;s replies stutter and stall. Something ran out. The hard part is telling what.
The short version: the GPU is being shared, and vLLM&rsquo;s scheduler is good at the sharing part. What usually decides how many people fit is memory space, and long conversations eat most of it. That also tells you which targets to pick for a personal assistant, an office tool, or a customer-facing app.
</description>
      <enclosure url="https://devopstales.github.io/img/vllm-scheduling.webp" length="10336" type="image/png" />
    </item>
    
  </channel>
</rss>
