Research Open the demo
Fugal · Frontier Scoreboard SNAPSHOT 2026-07-04 MATRIX 320q · 17 models

Flagship parity at half the price.

The measured system: a router weighs each model's odds of getting the query right against its price; an LLM verifier accepts or escalates. (The shipping product routes once, verifier-free — its own e2e numbers are queued.) Plotted below: every model's measured accuracy against its real cost per query. Fugal lands at the flagship's accuracy for half the spend — and the honest read of the statistics is right here, not buried.

Fugal · live e2e
0.973acc
$0.0079 / query · n=110 fresh
vs gpt-5.5 flagship
2.1×cheaper
COST WIN · ROBUST
Accuracy vs flagship
+0.9pts
PARITY · NOT SIGNIFICANT
Verifier false-accept
3.4%
of accepts · trustworthy gate

Accuracy vs price

17 models · 320-question matrix · cost on a log scale

Score by
Fugal (product) flagship / baseline on Pareto frontier dominated model — frontier = no model is both more accurate and cheaper

Full standings

Ranked by matrix accuracy. Heat strip = per-domain accuracy (redder = weaker).
ModelAccuracy$ / queryFrontier gsm8kmath500aime25humanevalmmlupro
domain accuracy 0.3  →  1.0

Routers rot

Model prices and capabilities shift monthly, so a frozen router decays. This scoreboard is meant to be re-measured on a cadence — the same refresh loop becomes the reference validator for the subnet.

Honest by design

The accuracy edge over gpt-5.5 is not statistically significant at n=110 (McNemar p=1.0). The claim is parity accuracy at 2.1× lower cost — the cost win is what survives the stats.

Traceable

Every point regenerates offline from gate_matrix_2026-07.json and e2e_cache_2026-07.json via figs/emit_frontier.py. No number here lacks an archived, re-gradeable artifact.

schema fugal.frontier/1 · generated 2026-07-04 · 320-question matrix + 110-question live e2e