vLLM Scheduling: How Throughput and Latency Trade Off in Shared Inference

vLLM Scheduling: How Throughput and Latency Trade Off in Shared Inference

11 min read

At 9 in the morning Alice asks the office’s local model to summarize yesterday’s meeting, and the answer streams back faster than she can read it. By 10 the whole team has found the thing. Same GPU, same model — and now everyone’s replies stutter and stall. Something ran out. The hard part is telling what.

The short version: the GPU is being shared, and vLLM’s scheduler is good at the sharing part. What usually decides how many people fit is memory space, and long conversations eat most of it. That also tells you which targets to pick for a personal assistant, an office tool, or a customer-facing app.

Model Routing: Stop Using One Model for Everything

Model Routing: Stop Using One Model for Everything

4 min read

A 70B model summarizing an email burns money; a 3B model reviewing production code courts disaster. Routing matches task complexity to model capability across four strategies — capability, cost, latency, hybrid — each optimizing what hurts most. Production systems converge on hybrid; start there only after one model genuinely fails.

LLM Guardrails in Practice: Control the Risk, Not the Model

LLM Guardrails in Practice: Control the Risk, Not the Model

3 min read

Models hallucinate, leak data, emit harmful content, and refuse legitimate requests. Guardrails constrain behavior without removing capability — but only the guardrails that match real risk earn their latency. The discipline is knowing which ones matter and which are noise.

Multi-Model System Design: The Right Model for the Right Task

Multi-Model System Design: The Right Model for the Right Task

3 min read

Running a 70B model to summarize an email wastes money; running a 3B model to review production code courts disaster. Multi-model design isn’t about collecting models — it’s about architecture that puts the right model on the right task at the right time, with the simplest pattern that survives the constraints.

What Is Spec-Driven Development? The Spec as Source of Truth

What Is Spec-Driven Development? The Spec as Source of Truth

4 min read

For most of software history the spec was a temporary planning artifact and the code was ground truth. Spec-driven development inverts that: the versioned specification becomes the primary artifact, and code is generated or verified against it. The idea is old — formal methods, design-by-contract, BDD all contain versions. What’s new is the motive: AI agents need explicit, durable intent because prompts evaporate and sessions reset.

Fixing Proxmox SDN: missing 'source /etc/network/interfaces.d/sdn' directive

Fixing Proxmox SDN: missing 'source /etc/network/interfaces.d/sdn' directive

7 min read

If you have ever applied a Proxmox VE Software-Defined Network configuration and seen this warning pop up in the task view, you are not alone:

WARN: missing 'source /etc/network/interfaces.d/sdn' directive for SDN support!
Created symlink /etc/systemd/system/multi-user.target.wants/dnsmasq@testzone.service -> /lib/systemd/system/dnsmasq@.service.

TASK WARNINGS: 1

It looks harmless, and in most cases it is. But if you want SDN-generated networks to actually come up on the node, that one line is the difference between a pending configuration and a running one. This post explains where the warning comes from in the Proxmox source code, why the fix is a single source directive, and how to verify the whole pipeline works afterwards.

Ceph 19.2.6, the CephX key rotation, and getting through it without losing sleep

Ceph 19.2.6, the CephX key rotation, and getting through it without losing sleep

17 min read

Ceph 19.2.6 is a security patch. What makes it different from every other Ceph patch you have shipped in the last five years is that the fix for the headline CVE introduces the first new CephX key type in the project’s history, aes256k. That turns “update the packages” into “migrate every credential in the cluster, on every node, for every client, without breaking the ones that are in use right now.”

This post walks through the whole thing: what the four CVEs actually are, how CephX authentication works well enough to understand why the rotation is hard, where the keys live on a Proxmox node, the preflight checks that save you during the upgrade, the Proxmox migration helper step by step, and the specific things that went wrong for people who did not follow the procedure.

Multi-Agent Orchestration Patterns: Six Topologies and How Each Breaks

Multi-Agent Orchestration Patterns: Six Topologies and How Each Breaks

6 min read

Single-agent systems peaked with bounded tasks and one prompt. Production moved on: organizations now run double-digit agent counts, and a large share of multi-agent pilots die within months of deployment. The killer isn’t the idea — it’s picking the wrong coordination topology, or the right one without knowing its failure modes.

Open WebUI: Self-Hosted ChatGPT Alternative for Local LLMs

Open WebUI: Self-Hosted ChatGPT Alternative for Local LLMs

3 min read

Open WebUI brings the ChatGPT experience onto your own infrastructure: Ollama plus any OpenAI-compatible backend behind a modern chat UI with RAG document chat, multi-user auth, voice, prompt tooling, and model management. Privacy-first, offline-capable, and scalable from laptop Docker runs to Kubernetes fleets.

Ollama to vLLM: When Your Local Server Needs to Grow Up

Ollama to vLLM: When Your Local Server Needs to Grow Up

6 min read

Ollama is the easiest way to run a local model, until the experiment becomes a shared service. Requests queue, latency wobbles, prefixes recompute, and one GPU stops being enough. vLLM answers exactly those problems — but migration is a trade, not an upgrade: simplicity for scheduling, memory control, parallelism, and production operations.