Regression
prompt-regress
Prompt v13 can fix the select and break the delete confirm. prompt-regress compares two recorded runs of one suite: which cases rose, which fell, and what the grader note says. The published diff is 16 fixture cases, not a live model call.
Why
A new system prompt is rarely better everywhere. It lifts some UI tasks and drops others: the model starts using Select and forgets the danger variant on delete. An average score hides that.
prompt-regress versions the prompt and shows a per-case diff. The diff reads a recorded run, the same kind of fixture ds-eval already publishes.
Prompts
prompts/v12.md asks for speed and allows a native control when a system component is unclear.
You build the screen from the design system.
Prefer speed. If a system component is unclear, use a native control.prompts/v13.md names the components and bans a raw button, a hex, and a custom modal. When a case conflicts with the system, the prompt says to follow the system.
You build the screen from the design system.
Use Button, Input, Select, and Dialog. Never emit a raw button, a hardcoded hex color, or a custom modal.
If a case conflicts with the system, follow the system and say so.Run
python3 prompt-regress/prompt_regress.py diff \
prompt-regress/prompts/v12.md \
prompt-regress/prompts/v13.mdThe prompt filename finds its JSON: prompts/v12.md reads runs/v12.json. A case stores pass, DS score, tokens, and a note. This fixture has no screenshot. The field can be filled when a run saves one.
Diff
A case is improved or regressed when pass/fail flips, or the DS score moves by 10 points or more. Otherwise it is unchanged.
| Condition | Label |
|---|---|
| pass flips to fail, or DS drops by 10+ | regressed |
| fail flips to pass, or DS rises by 10+ | improved |
| pass stays and DS moves by less than 10 | unchanged |
On 16 recorded cases the command prints:
v12 → v13
6 improved / 2 regressed / 8 unchanged
tokens 26270 → 26230The token column is the sum stored on the recorded run, not an API bill.
Every case
Notes are copied from runs/v13.json. Eight cases did not move enough to enter the diff. An average would hide that.
| Case | Pass | DS | Verdict | Note |
|---|---|---|---|---|
select-001 | fail → pass | 38 → 91 | improved | Select with a label. |
settings-002 | fail → pass | 44 → 93 | improved | Button and color token. |
dialog-003 | fail → pass | 41 → 90 | improved | Dialog instead of a div. |
empty-004 | pass | 86 → 87 | unchanged | Still the empty-state pattern. |
table-005 | pass | 80 → 82 | unchanged | Table unchanged in substance. |
form-006 | fail → pass | 52 → 89 | improved | Every field has a label. |
delete-007 | pass → fail | 84 → 61 | regressed | Confirm button lost its danger variant. |
tabs-008 | pass | 79 → 80 | unchanged | Tabs still from the system. |
toast-009 | fail → pass | 47 → 85 | improved | Toast uses the system component. |
billing-010 | pass | 83 → 84 | unchanged | Billing stays on tokens. |
permissions-011 | fail → pass | 55 → 86 | improved | Checkbox group from the system. |
profile-012 | pass → fail | 88 → 58 | regressed | Avatar control became a raw button. |
bulk-013 | pass | 76 → 78 | unchanged | Bulk bar still acceptable. |
filters-014 | fail | 49 → 51 | unchanged | Chips still hardcoded. |
search-015 | pass | 81 → 83 | unchanged | Search field still labeled. |
nav-016 | pass | 77 → 79 | unchanged | Side nav unchanged. |
Record
One case in both runs. The record shape is the same: id, pass, ds, tokens, note.
{
"v12": {
"id": "select-001",
"pass": false,
"ds": 38,
"tokens": 1640,
"note": "Native select, no label."
},
"v13": {
"id": "select-001",
"pass": true,
"ds": 91,
"tokens": 1720,
"note": "Select with a label."
}
}Where it sits
ds-context records what the model was allowed to see. ds-eval scores a case. prompt-regress compares two of those runs when the prompt, the context pack, or the model changes. Regression here is the list of cases that flipped, not a drop in one headline number.
Limit
The command does not call a model and does not render UI. Notes and scores live in JSON that was recorded earlier. Until a run stores screenshots, the explanation is one line per case. 6 / 2 / 8 belongs to these 16 fixtures, not to a 100-case suite.