Ais

vLLM Scheduling: How Throughput and Latency Trade Off in Shared Inference

vLLM Scheduling: How Throughput and Latency Trade Off in Shared Inference

11 min read

At 9 in the morning Alice asks the office’s local model to summarize yesterday’s meeting, and the answer streams back faster than she can read it. By 10 the whole team has found the thing. Same GPU, same model — and now everyone’s replies stutter and stall. Something ran out. The hard part is telling what.

The short version: the GPU is being shared, and vLLM’s scheduler is good at the sharing part. What usually decides how many people fit is memory space, and long conversations eat most of it. That also tells you which targets to pick for a personal assistant, an office tool, or a customer-facing app.

Model Routing: Stop Using One Model for Everything

Model Routing: Stop Using One Model for Everything

4 min read

A 70B model summarizing an email burns money; a 3B model reviewing production code courts disaster. Routing matches task complexity to model capability across four strategies — capability, cost, latency, hybrid — each optimizing what hurts most. Production systems converge on hybrid; start there only after one model genuinely fails.

LLM Guardrails in Practice: Control the Risk, Not the Model

LLM Guardrails in Practice: Control the Risk, Not the Model

3 min read

Models hallucinate, leak data, emit harmful content, and refuse legitimate requests. Guardrails constrain behavior without removing capability — but only the guardrails that match real risk earn their latency. The discipline is knowing which ones matter and which are noise.

Multi-Model System Design: The Right Model for the Right Task

Multi-Model System Design: The Right Model for the Right Task

3 min read

Running a 70B model to summarize an email wastes money; running a 3B model to review production code courts disaster. Multi-model design isn’t about collecting models — it’s about architecture that puts the right model on the right task at the right time, with the simplest pattern that survives the constraints.

What Is Spec-Driven Development? The Spec as Source of Truth

What Is Spec-Driven Development? The Spec as Source of Truth

4 min read

For most of software history the spec was a temporary planning artifact and the code was ground truth. Spec-driven development inverts that: the versioned specification becomes the primary artifact, and code is generated or verified against it. The idea is old — formal methods, design-by-contract, BDD all contain versions. What’s new is the motive: AI agents need explicit, durable intent because prompts evaporate and sessions reset.

Multi-Agent Orchestration Patterns: Six Topologies and How Each Breaks

Multi-Agent Orchestration Patterns: Six Topologies and How Each Breaks

6 min read

Single-agent systems peaked with bounded tasks and one prompt. Production moved on: organizations now run double-digit agent counts, and a large share of multi-agent pilots die within months of deployment. The killer isn’t the idea — it’s picking the wrong coordination topology, or the right one without knowing its failure modes.

Open WebUI: Self-Hosted ChatGPT Alternative for Local LLMs

Open WebUI: Self-Hosted ChatGPT Alternative for Local LLMs

3 min read

Open WebUI brings the ChatGPT experience onto your own infrastructure: Ollama plus any OpenAI-compatible backend behind a modern chat UI with RAG document chat, multi-user auth, voice, prompt tooling, and model management. Privacy-first, offline-capable, and scalable from laptop Docker runs to Kubernetes fleets.

Ollama to vLLM: When Your Local Server Needs to Grow Up

Ollama to vLLM: When Your Local Server Needs to Grow Up

6 min read

Ollama is the easiest way to run a local model, until the experiment becomes a shared service. Requests queue, latency wobbles, prefixes recompute, and one GPU stops being enough. vLLM answers exactly those problems — but migration is a trade, not an upgrade: simplicity for scheduling, memory control, parallelism, and production operations.

The SDD Workflow: Five Phases from Requirements to Verified Code

The SDD Workflow: Five Phases from Requirements to Verified Code

6 min read

Spec-driven development works when the specification is a workflow, not a document filed away after kickoff. The point is a sequence of reviewable artifacts — requirements, design, tasks, implementation, validation — each reducing ambiguity before anyone, human or agent, changes production code.