Model Routing Playground
Interactive demo — TNB Newsletter
Section 01

Live Router Demo

Type any query. The classifier runs in your browser and shows which model tier it would route to — and why.

Try: "What are your hours?" · "Summarize this support ticket" · "Write a migration plan for PostgreSQL"
Cost:
Latency:
Show cost if routed to GPT-4o instead
If you sent this to GPT-4o:

Section 02

Cost Savings Calculator

Adjust your volume, current model, and traffic mix. See projected savings from a three-tier routing strategy.

100,000 queries/mo
Query Mix
Simple queries65%
Medium queries25%
Complex queries10%
Sliders must sum to 100%
Simple
Medium
Complex
Current Monthly Cost
$1,500
GPT-4o at $0.015/query
With Routing
$148
Three-tier routing strategy
Monthly Savings
$1,352
90.1% reduction
Annual Savings
$16,224
Projected 12-month savings
Current$1,500/mo
With routing$148/mo

Section 03

Routing Architecture

How queries flow through a three-tier router — and what the silent failure mode looks like.

INCOMING QUERY User query arrives ROUTER Lightweight classifier TIER 1 — SMALL LM Phi-4 mini / Gemma 3 ~$0.0001/query · ~180ms 65% of queries TIER 2 — MOE Mixtral 8x22B / DeepSeek ~$0.0008/query · ~420ms 25% of queries TIER 3 — LARGE DENSE GPT-4o / Claude Opus ~$0.015/query · ~1800ms 10% of queries RESPONSE Returned to user 95% cost reduction vs. routing everything to Tier 3
The failure mode: a miscalibrated router
Correct routing
Complex query arrives
Router classifies correctly → Tier 3
GPT-4o / Claude Opus processes
Correct, high-quality response
Miscalibrated router
Complex query arrives
Miscalibrated router → Tier 1
Phi-4 mini processes instead
Plausible but wrong response
Nobody catches it for weeks
The model responded. The response looked fine. It was wrong. This is why router monitoring matters more than router accuracy. A 98% accurate router still misfires on 1-in-50 queries — at 100k queries/month, that's 2,000 silent failures every month.