Skip to content
Prompt engineering

Fidelity is a claim

A fidelity score without a rubric is a vibe. How Motif scores runs, why screenshots are not proof, and the scoring discipline that turns 'this looks good' into numbers you can iterate against.

Intermediate2 min read (computed · recorded 10)updated 2026-09-10promptsfidelityscoringprocess

What a fidelity score should mean

Every Motif prompt ships with per-model fidelity scores and run dates, and the score is only useful because it is defined. The rubric scores three axes: structure (are the brief's blocks present, in order, with nothing extra), content (is the copy specific — named products, real figures, no filler), and craft (palette discipline, spacing rhythm, restraint). A score of 92 means 'structure and content held, craft slightly under the brief's ceiling' — not 'I liked it'.

  • Structure is binary-ish: blocks present? In order? Extras? A brief with a block list makes structure machine-checkable — which is why the block list is the backbone of every Motif prompt.
  • Content is specificity: 'the copy names real things' scores; 'compelling copy that elevates the brand' is the model talking about itself. Score the nouns, not the adjectives.
  • Craft is restraint: palette obeyed, one accent, spacing from a scale. The brief's never-list makes craft checkable too — every violated never is a docked point.
  • Screenshots are not proof: a beautiful screenshot of a page that ignored half the brief is a beautiful failure. Score against the brief, then admire the screenshot.

The scoring discipline

scorecard.md
Prompt: bike-shop-service-tiers   Model: Claude 4.6 Sonnet

Structure  10/10  all 4 blocks, in order, no extras
Content     9/10  tiers named + priced; one generic line
Craft       9/10  palette held; spacing 2px off on one card
                     ----------
Fidelity    91    (weighted: structure .4, content .3, craft .3)

Notes: "reads like a menu a mechanic would stand behind"
       -> that is the brief's own line; it held.

Why the claim matters

A score you can defend turns iteration into science: run 1 at 84, change one thing, run 2 at 91 — the delta is attributable. It also keeps the library honest: a component or prompt marked 'verified' has a dated score behind it, and the score's breakdown tells the next user where the weaknesses live. Fidelity is a claim — a claim with evidence attached — and evidence is the only thing that survives contact with a new model version.

Words this guide uses that the glossary defines: Fidelity score

Practise the lesson

Theory sticks when you ship it. These original Motif assets put this guide's lesson to work — open one and copy it into your own page.