Skip to content

Model leaderboard

Which model wins, where

Not a single number: the leaderboard splits every recorded run by industry and difficulty, so you can see whether Claude owns marketing or GLM owns dashboards. All scores come from the live run log, not editorial opinions.

3
models tracked
88%
top overall
99%
most stable
Model leaderboard — sorted by run volume, broken down by the industry each model wins and loses in
ModelRunsOverallWins inWins avgLoses inLoses avgError rate
01GLM-4.6 (CN)
29388%Agency92%Local business81%8 of 293
02Claude 4.6 Sonnet
18490%A11yStrict94%Utility83%1 of 184
03Codex
18488%AI95%Food85%8 of 184
By difficulty — the hard question: which model do you trust when the brief gets gnarly
ModelStarterIntermediateAdvancedRuns
GLM-4.6 (CN)88%86%90%293
Claude 4.6 Sonnet90%89%94%184
Codex88%87%91%184

Briefs that fit

Pick the model you actually use, the industry you're shipping in, and the difficulty of the brief — this shows which prompts from the catalog that combo will handle best.

184 briefs match this combo · sorted by measured fidelity

technical, confident

avg 93%best Codex4 runs

dark, premium, minimal

avg 92%best Claude 4.6 Sonnet4 runs

earnest, hopeful, concrete

avg 92%best Claude 4.6 Sonnet4 runs

spec-exact, keyboard-clean, audit-visible

avg 92%best Claude 4.6 Sonnet4 runs

editorial, confident

avg 91%best GLM-4.6 (CN)4 runs

credible, calm, plain-spoken

avg 91%best GLM-4.6 (CN)4 runs

editorial, tactile, playful

avg 91%best GLM-4.6 (CN)4 runs

developer-calm, precise, searchable

avg 91%best Claude 4.6 Sonnet4 runs

linen-calm, honest, neighbourhood

avg 91%best Claude 4.6 Sonnet4 runs

Scores are a report card, not a benchmark. Fidelity here is an editorial judgement recorded per run — the same numbers a /prompts card shows, just re-projected to answer the “where should I spend my time” question. Every run log on this site says so too.