Frontier Try it free
Fugal · Frontier Scoreboard SNAPSHOT 2026-07-04 MATRIX 320q · 17 models

Every model's accuracy, against what it costs.

Every model below is scored on the same questions and charged at its real price, so the trade-off is visible rather than asserted. The headline it gives up: the flagship isn't the top of the board. A model costing 4.5× less matches it to within three questions out of 320 — parity, not an edge, and the price gap is the real result. Further down, 91% accuracy costs 75× less than the flagship. No single row is right for every question, which is what a router is for.

Fugal itself is not a row on this board — it was measured separately, on 110 fresh questions: 0.945 at $0.00285 a query, parity with gpt-5.5 (−1.8 pts, McNemar p=0.63) at 5.9× lower cost. How that was measured →

flagship · gpt-5.5
$0.0197/query
0.959 acc · 307 of 320
NOT ON THE FRONTIER
same accuracy · qwen3.7-plus
4.5×cheaper
0.969 acc · 310 of 320 · $0.0044
PARITY · +3 QUESTIONS
cheapest on the board
75×cheaper
0.906 acc · $0.00026 · deepseek-v4-flash
−5.3 PTS · STILL 91%
the board
320questions
17 models · 5 domains · one grader
5,440 CELLS · NO GAPS

Accuracy vs price

17 models · 320-question matrix · cost on a log scale

Score by
flagship / baseline on Pareto frontier dominated model — frontier = no model is both more accurate and cheaper

Full standings

Ranked by matrix accuracy. Heat strip = per-domain accuracy (redder = weaker).
ModelAccuracy$ / queryFrontier gsm8kmath500aime25humanevalmmlupro
domain accuracy 0.3  →  1.0

No lab can publish this board

The most accurate model here is Qwen's, the cheapest is DeepSeek's, and the frontier spans two vendors. A router inside a model lab picks from one catalogue and will never route you to a competitor — so the only place this comparison can be made honestly is outside all of them.

Routers rot — so this has to keep moving

Prices and capabilities shift monthly, so a frozen router decays and a one-off board goes stale with it. This is meant to be re-measured on a cadence, and eventually not by us: the same refresh loop becomes the validator for a public market in routing, where competitors are paid to route better and a stale router simply loses. The design →

Honest by design

This board scores the models, not us — which is why it stays true whatever we ship, and why Fugal is not plotted on these axes. The router was measured separately, on 110 fresh questions rather than these 320; putting a differently-sized experiment on the same chart would invite a comparison the data doesn't support. The result is quoted above and derived on About.

Traceable

Every point regenerates offline from gate_matrix_2026-07.json via figs/emit_frontier.py. No number here lacks an archived, re-gradeable artifact.

schema fugal.frontier/1 · generated 2026-07-04 · 320-question matrix · 17 models · 5 domains · route-once e2e: 110 fresh questions, 2026-07-27