Multi-Model System Design: The Right Model for the Right Task

Page content

Running a 70B model to summarize an email wastes money; running a 3B model to review production code courts disaster. Multi-model design isn’t about collecting models — it’s about architecture that puts the right model on the right task at the right time, with the simplest pattern that survives the constraints.

Multi-model LLM system design patterns Sequential, parallel, hierarchical, ensemble: five wirings from one model to many

Five Patterns, One Rule

Pattern Complexity Use when Tradeoff
Single model Lowest Prototyping, uniform tasks Capability ceiling
Sequential Low Multi-step workflows Latency stacks
Parallel Medium Independent tasks Cost multiplies
Hierarchical High Complex reasoning Orchestration burden
Ensemble Highest Critical decisions Highest cost

Pick the simplest that works — complexity compounds. Note the boundary: this is task-to-model routing, not autonomous agents coordinating; for that see multi-agent orchestration.

Sequential: Pipeline and Router

Pipeline chains specialists, each output feeding the next:

class ModelPipeline:
    def __init__(self):
        self.models = [
            {"model": "qwen3-1.7b", "task": "classify"},
            {"model": "qwen3-8b", "task": "extract"},
            {"model": "qwen3-32b", "task": "reason"},
        ]

    def process(self, input: str) -> str:
        current = input
        for model_config in self.models:
            current = self.call_model(
                model_config["model"],
                self.create_prompt(model_config["task"], current)
            )
        return current

Latency adds per stage — three models, roughly triple the wait. Only chain steps that genuinely need different models. Router instead classifies once and dispatches to one specialist (code, math, creative, general). The classifier is the weak link: misclassify code review as summarization and quality degrades silently. Keep categories crisp enough for even a small classifier.

Parallel: Fan-Out and Voting

Fan-out runs one prompt through several models concurrently — comparison, A/B testing, best-output selection — at full per-model cost:

import asyncio

class ModelFanOut:
    def __init__(self):
        self.models = ["qwen3-8b", "qwen3-32b", "claude-sonnet-4"]

    async def process(self, prompt: str) -> list[str]:
        tasks = [self.call_model(model, prompt) for model in self.models]
        return await asyncio.gather(*tasks)

Voting adds consensus on top — exact-match majority works for classification; generation needs semantic similarity, not string equality. Worth it where quality beats spend.

Hierarchical: Planner-Executor and Supervisor-Worker

A strong planner decomposes; cheap executors run steps; the planner synthesizes. Planning is expensive, execution cheap — the economics work when decomposition dominates:

class PlannerExecutor:
    def __init__(self):
        self.planner = "qwen3-32b"
        self.executors = {"code": "qwen2.5-coder-7b", "search": "qwen3-8b"}

    def process(self, task: str) -> str:
        plan = self.call_model(self.planner, f"Plan: {task}")
        results = [self.call_model(self.executors.get(s["type"], "qwen3-8b"), s["prompt"])
                   for s in self.parse_plan(plan)]
        return self.call_model(self.planner, f"Synthesize: {results}")

Supervisor-worker adds review to delegation: assign, execute, reviewer pass. The supervisor plans, delegates, and judges — keep it fast or everything queues behind it.

Ensemble: Weighted and Consensus

Weighted ensembles score each output and take the max, weights tracking measured performance rather than benchmark hope. Consensus ensembles require threshold agreement (say 0.7) and escalate to the strongest model below it — strictness dials speed against confidence. Reserve both for decisions where wrongness costs more than tokens.

When It Earns Its Keep

Mixed workloads, critical-decision quality bars, cost or latency walls. Skip it for uniform complexity (one good model wins), prototypes (optimize after the bill arrives), and simplicity-first systems. Start single; add models at measured constraints. The companion practices live nearby: model routing strategies for capability, cost, and latency routing with fallbacks; cost optimization for budgets, caching, and break-even math; guardrails so cheaper models stay inside safety bounds.

Tradeoff Matrix

Pattern Cost Latency Quality Complexity
Single model Lowest Lowest Variable Lowest
Sequential Medium High High Medium
Parallel High Low High Medium
Hierarchical High High Highest High
Ensemble Highest Medium Highest Highest

Summary

Right model, right task, right time — with architecture chosen by constraint, not enthusiasm. Capability gaps call for hierarchy, latency walls for parallelism, critical calls for ensembles, and everything else for the simplest wiring that holds.

Which multi-model pattern carries your production load — and what broke first? Share the topology in the comments below!