Vllm

vLLM Scheduling: How Throughput and Latency Trade Off in Shared Inference

vLLM Scheduling: How Throughput and Latency Trade Off in Shared Inference

11 min read

At 9 in the morning Alice asks the office’s local model to summarize yesterday’s meeting, and the answer streams back faster than she can read it. By 10 the whole team has found the thing. Same GPU, same model — and now everyone’s replies stutter and stall. Something ran out. The hard part is telling what.

The short version: the GPU is being shared, and vLLM’s scheduler is good at the sharing part. What usually decides how many people fit is memory space, and long conversations eat most of it. That also tells you which targets to pick for a personal assistant, an office tool, or a customer-facing app.

Ollama to vLLM: When Your Local Server Needs to Grow Up

Ollama to vLLM: When Your Local Server Needs to Grow Up

6 min read

Ollama is the easiest way to run a local model, until the experiment becomes a shared service. Requests queue, latency wobbles, prefixes recompute, and one GPU stops being enough. vLLM answers exactly those problems — but migration is a trade, not an upgrade: simplicity for scheduling, memory control, parallelism, and production operations.