prompt-regress

Prompt v13 can fix the select and break the delete confirm. prompt-regress compares two recorded runs of one suite: which cases rose, which fell, and what the grader note says. The published diff is 16 fixture cases, not a live model call.

Why

A new system prompt is rarely better everywhere. It lifts some UI tasks and drops others: the model starts using Select and forgets the danger variant on delete. An average score hides that.

prompt-regress versions the prompt and shows a per-case diff. The diff reads a recorded run, the same kind of fixture ds-eval already publishes.

Prompts

prompts/v12.md asks for speed and allows a native control when a system component is unclear.

You build the screen from the design system.
Prefer speed. If a system component is unclear, use a native control.

prompts/v13.md names the components and bans a raw button, a hex, and a custom modal. When a case conflicts with the system, the prompt says to follow the system.

You build the screen from the design system.
Use Button, Input, Select, and Dialog. Never emit a raw button, a hardcoded hex color, or a custom modal.
If a case conflicts with the system, follow the system and say so.

Run

python3 prompt-regress/prompt_regress.py diff \
  prompt-regress/prompts/v12.md \
  prompt-regress/prompts/v13.md

The prompt filename finds its JSON: prompts/v12.md reads runs/v12.json. A case stores pass, DS score, tokens, and a note. This fixture has no screenshot. The field can be filled when a run saves one.

Diff

A case is improved or regressed when pass/fail flips, or the DS score moves by 10 points or more. Otherwise it is unchanged.

ConditionLabel
pass flips to fail, or DS drops by 10+regressed
fail flips to pass, or DS rises by 10+improved
pass stays and DS moves by less than 10unchanged

On 16 recorded cases the command prints:

v12 → v13
6 improved / 2 regressed / 8 unchanged
tokens  26270 → 26230

The token column is the sum stored on the recorded run, not an API bill.

Every case

Notes are copied from runs/v13.json. Eight cases did not move enough to enter the diff. An average would hide that.

CasePassDSVerdictNote
select-001fail → pass38 → 91improvedSelect with a label.
settings-002fail → pass44 → 93improvedButton and color token.
dialog-003fail → pass41 → 90improvedDialog instead of a div.
empty-004pass86 → 87unchangedStill the empty-state pattern.
table-005pass80 → 82unchangedTable unchanged in substance.
form-006fail → pass52 → 89improvedEvery field has a label.
delete-007pass → fail84 → 61regressedConfirm button lost its danger variant.
tabs-008pass79 → 80unchangedTabs still from the system.
toast-009fail → pass47 → 85improvedToast uses the system component.
billing-010pass83 → 84unchangedBilling stays on tokens.
permissions-011fail → pass55 → 86improvedCheckbox group from the system.
profile-012pass → fail88 → 58regressedAvatar control became a raw button.
bulk-013pass76 → 78unchangedBulk bar still acceptable.
filters-014fail49 → 51unchangedChips still hardcoded.
search-015pass81 → 83unchangedSearch field still labeled.
nav-016pass77 → 79unchangedSide nav unchanged.

Record

One case in both runs. The record shape is the same: id, pass, ds, tokens, note.

{
  "v12": {
    "id": "select-001",
    "pass": false,
    "ds": 38,
    "tokens": 1640,
    "note": "Native select, no label."
  },
  "v13": {
    "id": "select-001",
    "pass": true,
    "ds": 91,
    "tokens": 1720,
    "note": "Select with a label."
  }
}

Where it sits

ds-context records what the model was allowed to see. ds-eval scores a case. prompt-regress compares two of those runs when the prompt, the context pack, or the model changes. Regression here is the list of cases that flipped, not a drop in one headline number.

Limit

The command does not call a model and does not render UI. Notes and scores live in JSON that was recorded earlier. Until a run stores screenshots, the explanation is one line per case. 6 / 2 / 8 belongs to these 16 fixtures, not to a 100-case suite.