Not a single number: the leaderboard splits every recorded run by industry and difficulty, so you can see whether Claude owns marketing or GLM owns dashboards. All scores come from the live run log, not editorial opinions.
3
models tracked
88%
top overall
99%
most stable
Model leaderboard — sorted by run volume, broken down by the industry each model wins and loses in
Model
Runs
Overall
Wins in
Wins avg
Loses in
Loses avg
Error rate
01GLM-4.6 (CN)
293
88%
Agency
92%
Local business
81%
8 of 293
02Claude 4.6 Sonnet
184
90%
A11yStrict
94%
Utility
83%
1 of 184
03Codex
184
88%
AI
95%
Food
85%
8 of 184
By difficulty — the hard question: which model do you trust when the brief gets gnarly
Model
Starter
Intermediate
Advanced
Runs
GLM-4.6 (CN)
88%
86%
90%
293
Claude 4.6 Sonnet
90%
89%
94%
184
Codex
88%
87%
91%
184
Briefs that fit
Pick the model you actually use, the industry you're shipping in, and the difficulty of the brief — this shows which prompts from the catalog that combo will handle best.
184 briefs match this combo · sorted by measured fidelity
Scores are a report card, not a benchmark. Fidelity here is an editorial judgement recorded per run — the same numbers a /prompts card shows, just re-projected to answer the “where should I spend my time” question. Every run log on this site says so too.
This site uses Google AdSense to show ads. It sets cookies and may collect browsing data. Learn more.