Model Routing: Stop Using One Model for Everything

Page content

A 70B model summarizing an email burns money; a 3B model reviewing production code courts disaster. Routing matches task complexity to model capability across four strategies — capability, cost, latency, hybrid — each optimizing what hurts most. Production systems converge on hybrid; start there only after one model genuinely fails.

LLM model routing across capability, cost, and latency Classify the task, then spend the right model on it — not the biggest one

Capability Routing: Task to Model

Classify, then dispatch. Sentiment goes to 1-3B models, summarization to 3-7B, code to 7-14B coders, hard reasoning to 14-32B, creative and analytical peaks to 32B and flagship APIs:

ROUTING_RULES = {
    "classify": {"model": "qwen3-1.7B", "max_tokens": 100},
    "summarize": {"model": "qwen3-8B", "max_tokens": 500},
    "code_review": {"model": "qwen2.5-coder-7b", "max_tokens": 2000},
    "reason": {"model": "qwen3-32b", "max_tokens": 4000},
    "creative": {"model": "claude-sonnet-4", "max_tokens": 8000},
}

def route_request(task_type: str) -> dict:
    return ROUTING_RULES.get(task_type, ROUTING_RULES["reason"])

The classifier is the weak link — code review filed as “summarization” loses quality silently. Keep categories crisp enough for even a small classifier, and default unknown tasks upward to reasoning models. Architectural sibling: multi-model system design covers pipeline, parallel, hierarchical, and ensemble wirings around the same routing idea.

Cost Routing: Spend Downhill

Local inference amortizes fast — a mid-range GPU pays for itself in months at moderate API volume, after which local tokens cost electricity against dollars per million. Budget-aware routers fall back as sessions spend:

class CostAwareRouter:
    def __init__(self, budget_per_session: float = 0.10):
        self.budget = budget_per_session
        self.spent = 0.0
        self.models = {
            "cheap": {"model": "qwen3-8B", "cost": 0.0},
            "medium": {"model": "qwen3-32b", "cost": 0.0},
            "expensive": {"model": "claude-sonnet-4", "cost": 0.000015},
        }

    def route(self, task: str) -> str:
        ratio = self.spent / self.budget
        if ratio < 0.5:
            return self.models["expensive"]["model"]
        elif ratio < 0.8:
            return self.models["medium"]["model"]
        return self.models["cheap"]["model"]

Quality visibly degrades down the slope — flagship to mid to small — which is acceptable only where later-session work tolerates it. Routing routine traffic locally while reserving APIs for edge cases doubles as vendor-independence hygiene. Full break-even math lives in cost optimization.

Latency Routing: First Token Is the Product

Streaming users feel first-token latency: sub-200ms real-time chat caps models under ~7B, sub-500ms interactive tools under ~14B, batch and research tolerate anything. Route by measured completion budget, fastest qualifying model wins:

class LatencyAwareRouter:
    def __init__(self):
        self.model_latencies = {
            "qwen3-1.7b": {"first_token": 0.05, "complete": 0.5},
            "qwen3-8B": {"first_token": 0.15, "complete": 2.0},
            "qwen3-32b": {"first_token": 0.5, "complete": 10.0},
            "claude-sonnet-4": {"first_token": 0.3, "complete": 5.0},
        }

    def route(self, target_latency: float) -> str:
        for model, latencies in sorted(
            self.model_latencies.items(),
            key=lambda x: x[1]["complete"]
        ):
            if latencies["complete"] <= target_latency:
                return model
        return "qwen3-1.7b"

Published latencies are rough — hardware, quantization, and batch size move them. Measure on the deployment, not the datasheet.

Fallbacks: Best to Most Reliable

Models fail, APIs rate-limit, timeouts fire. Chain best-first down to a local last resort — local never fails on network or keys, only on speed:

class FallbackRouter:
    def __init__(self):
        self.fallback_chain = [
            {"model": "claude-sonnet-4", "timeout": 30},
            {"model": "qwen2.5-72b", "timeout": 60},
            {"model": "qwen3-32b", "timeout": 120},
            {"model": "qwen3-8b", "timeout": 300},
        ]

    def route_with_fallback(self, prompt: str) -> str:
        for config in self.fallback_chain:
            try:
                return self.call_model(
                    config["model"], prompt,
                    timeout=config["timeout"]
                )
            except (TimeoutError, APIError) as e:
                log.warning(f"Model {config['model']} failed: {e}")
                continue
        raise RuntimeError("All fallback models failed")

When Routing Pays

Mixed workloads with classification, summarization, and reasoning in one system — measurable money and latency wins. Skip it for uniform complexity (one good model, no router tax) and early prototypes (one model until cost or latency actually bites). Strategy tradeoffs at a glance: single model simplest and priciest with consistent quality; capability routing cheaper per task at moderate complexity; cost routing cheapest with variable quality; latency routing fastest possibly at quality cost; hybrid best of all at highest implementation cost. Convergence path: capability first, cost when bills arrive, latency when users complain. Keep routed models inside guardrails — cheaper models need the same safety bounds, not weaker ones.

Summary

One model everywhere is either wasteful or reckless. Classify tasks, route by capability, fall back by budget, constrain by latency, and always end the chain locally. Add routing when measurement demands it — never before.

How do you split work across models — rules, classifiers, or budgets? Share the routing table in the comments below!