Every model below is scored on the same questions and charged at its
real price, so the trade-off is visible rather than asserted. The headline it gives
up: the flagship isn't the top of the board. A model costing 4.5× less matches
it to within three questions out of 320 — parity, not an edge, and the price gap is the real
result. Further down, 91% accuracy costs 75× less than the flagship. No single row is
right for every question, which is what a router is for.
Fugal itself is not a row on this board — it
was measured separately, on 110 fresh questions: 0.945 at $0.00285 a query, parity
with gpt-5.5 (−1.8 pts, McNemar p=0.63) at 5.9× lower cost.
How that was measured →
flagship · gpt-5.5
$0.0197/query
0.959 acc · 307 of 320
NOT ON THE FRONTIER
same accuracy · qwen3.7-plus
4.5×cheaper
0.969 acc · 310 of 320 · $0.0044
PARITY · +3 QUESTIONS
cheapest on the board
75×cheaper
0.906 acc · $0.00026 · deepseek-v4-flash
−5.3 PTS · STILL 91%
the board
320questions
17 models · 5 domains · one grader
5,440 CELLS · NO GAPS
Accuracy vs price
17 models · 320-question matrix · cost on a log scale
Score by
flagship / baseline on Pareto frontier dominated model— frontier = no model is both more accurate and cheaper