Multi-Model System Design: The Right Model for the Right Task
Running a 70B model to summarize an email wastes money; running a 3B model to review production code courts disaster. Multi-model design isn’t about collecting models — it’s about architecture that puts the right model on the right task at the right time, with the simplest pattern that survives the constraints.
Sequential, parallel, hierarchical, ensemble: five wirings from one model to many
Five Patterns, One Rule
| Pattern | Complexity | Use when | Tradeoff |
|---|---|---|---|
| Single model | Lowest | Prototyping, uniform tasks | Capability ceiling |
| Sequential | Low | Multi-step workflows | Latency stacks |
| Parallel | Medium | Independent tasks | Cost multiplies |
| Hierarchical | High | Complex reasoning | Orchestration burden |
| Ensemble | Highest | Critical decisions | Highest cost |
Pick the simplest that works — complexity compounds. Note the boundary: this is task-to-model routing, not autonomous agents coordinating; for that see multi-agent orchestration.
Sequential: Pipeline and Router
Pipeline chains specialists, each output feeding the next:
class ModelPipeline:
def __init__(self):
self.models = [
{"model": "qwen3-1.7b", "task": "classify"},
{"model": "qwen3-8b", "task": "extract"},
{"model": "qwen3-32b", "task": "reason"},
]
def process(self, input: str) -> str:
current = input
for model_config in self.models:
current = self.call_model(
model_config["model"],
self.create_prompt(model_config["task"], current)
)
return current
Latency adds per stage — three models, roughly triple the wait. Only chain steps that genuinely need different models. Router instead classifies once and dispatches to one specialist (code, math, creative, general). The classifier is the weak link: misclassify code review as summarization and quality degrades silently. Keep categories crisp enough for even a small classifier.
Parallel: Fan-Out and Voting
Fan-out runs one prompt through several models concurrently — comparison, A/B testing, best-output selection — at full per-model cost:
import asyncio
class ModelFanOut:
def __init__(self):
self.models = ["qwen3-8b", "qwen3-32b", "claude-sonnet-4"]
async def process(self, prompt: str) -> list[str]:
tasks = [self.call_model(model, prompt) for model in self.models]
return await asyncio.gather(*tasks)
Voting adds consensus on top — exact-match majority works for classification; generation needs semantic similarity, not string equality. Worth it where quality beats spend.
Hierarchical: Planner-Executor and Supervisor-Worker
A strong planner decomposes; cheap executors run steps; the planner synthesizes. Planning is expensive, execution cheap — the economics work when decomposition dominates:
class PlannerExecutor:
def __init__(self):
self.planner = "qwen3-32b"
self.executors = {"code": "qwen2.5-coder-7b", "search": "qwen3-8b"}
def process(self, task: str) -> str:
plan = self.call_model(self.planner, f"Plan: {task}")
results = [self.call_model(self.executors.get(s["type"], "qwen3-8b"), s["prompt"])
for s in self.parse_plan(plan)]
return self.call_model(self.planner, f"Synthesize: {results}")
Supervisor-worker adds review to delegation: assign, execute, reviewer pass. The supervisor plans, delegates, and judges — keep it fast or everything queues behind it.
Ensemble: Weighted and Consensus
Weighted ensembles score each output and take the max, weights tracking measured performance rather than benchmark hope. Consensus ensembles require threshold agreement (say 0.7) and escalate to the strongest model below it — strictness dials speed against confidence. Reserve both for decisions where wrongness costs more than tokens.
When It Earns Its Keep
Mixed workloads, critical-decision quality bars, cost or latency walls. Skip it for uniform complexity (one good model wins), prototypes (optimize after the bill arrives), and simplicity-first systems. Start single; add models at measured constraints. The companion practices live nearby: model routing strategies for capability, cost, and latency routing with fallbacks; cost optimization for budgets, caching, and break-even math; guardrails so cheaper models stay inside safety bounds.
Tradeoff Matrix
| Pattern | Cost | Latency | Quality | Complexity |
|---|---|---|---|---|
| Single model | Lowest | Lowest | Variable | Lowest |
| Sequential | Medium | High | High | Medium |
| Parallel | High | Low | High | Medium |
| Hierarchical | High | High | Highest | High |
| Ensemble | Highest | Medium | Highest | Highest |
Summary
Right model, right task, right time — with architecture chosen by constraint, not enthusiasm. Capability gaps call for hierarchy, latency walls for parallelism, critical calls for ensembles, and everything else for the simplest wiring that holds.
Which multi-model pattern carries your production load — and what broke first? Share the topology in the comments below!