The Brutus Router · catalog v2026-07-20.2

Routing with receipts.

Most model routers are black boxes. Ours shows its work: every decision is grounded in independent public evaluations — the Artificial Analysis Intelligence Index and LMArena's human-preference data — under a versioned, auditable catalog. And every routed request logs what it actually cost against the frontier-on-everything counterfactual. Measured, not marketed.

The catalog the router runs on

The production routing table, verbatim. Sourced scores carry their authority and as-of date; models the authorities haven't scored are marked unsourced and never dressed up as data.

ModelAA Intelligence v4.1LMArena (2026-07-16)SWE-bench Verified$ / 1M in · out
claude-opus-4-8anthropic56148388.6%$5 · $25
gpt-5.6-solopenai591514unsourced$5 · $30
gpt-5.6-terraopenai55unsourcedunsourced$2.50 · $15
claude-sonnet-5anthropic53147985.2%$2 · $10
gpt-5.6-lunaopenai51unsourcedunsourced$1 · $6
gpt-5.4-miniopenai40unsourcedunsourced$0.75 · $4.50
gpt-5.4-nanoopenai38unsourcedunsourced$0.20 · $1.25
claude-haiku-4-5anthropicunsourcedunsourcedunsourced$1 · $5
deepseek-chatopenServed via your connected keyunsourcedunsourcedunsourced$0.14 · $0.28
openai/gpt-oss-120bopenServed via your connected keyunsourcedunsourcedunsourced$0.15 · $0.60
gemini-3.1-progoogleunsourcedunsourcedunsourced$2.50 · $15
gemini-3.1-flashgoogleunsourcedunsourcedunsourced$0.30 · $2.50
grok-4.1xaiunsourcedunsourcedunsourced$3 · $15
grok-4.1-minixaiunsourcedunsourcedunsourced$0.30 · $1.50

Data policy: OpenAI and Anthropic models in this catalog are served under provider API terms that do not train on your inputs by default, with zero-data-retention available to qualifying enterprise agreements — and with an org inference key connected, your own DPA governs directly.

Open-weights entries are served under their host's own terms — verify data policy with the host before routing sensitive work.

Sources: Artificial Analysis Intelligence Index v4.1 (artificialanalysis.ai) · LMArena Arena Score, CC BY 4.0, snapshot 2026-07-16 (style control off) · SWE-bench Verified (swebench.com) · provider list prices. Catalog v2026-07-20.2; refreshed and human-reviewed on every version bump.

What the router picks — and why

Selection rule: models within a complexity-based tolerance of the best task quality qualify; the cheapest qualifier wins. Simple, deterministic, and auditable — no learned black box.

Quick answers & extraction

“What does this clause mean?” · “Pull every deadline into a list”

Routine
gpt-5.4-nano
Complex
gpt-5.4-nano

Signal: AA Intelligence Index v4.1

Drafting & summarization

“Draft the supplier email” · “Summarize these 40 pages”

Routine
gpt-5.6-luna
Complex
gpt-5.6-sol

Routine work runs 5× cheaper than frontier-on-everything — at quality the evals say you won't miss.

Analysis & decisions

“Compare these two vendor bids and recommend one”

Routine
gpt-5.6-luna
Complex
gpt-5.6-sol

Routine work runs 5× cheaper than frontier-on-everything — at quality the evals say you won't miss.

Code & formulas

“Write the SQL / the Excel formula / the script”

Routine
claude-sonnet-5
Complex
claude-opus-4-8

Routine work runs 3× cheaper than frontier-on-everything — at quality the evals say you won't miss.

The receipt

Every routed request in Brutus logs its actual cost — across every step of the execution strategy — against what the same request would have cost running a frontier model over the full input. The savings figure your team sees is computed from your own ledger, not a marketing benchmark. Beyond model choice, the router executes your organization's curated techniques (map-reduce over long material, draft→self-critique for high-stakes writing) and inherits Brutus's guardrails: approved-tool enforcement, data-sensitivity routing, and a full audit trail on routing credentials.

See the whole platform →