Benchmark
ds-eval
Can AI coding agents actually use your design system correctly? ds-eval is an open-source benchmark that runs the same UI tasks through different models and scores the generated interfaces.
Fixture87.8
Naive59.3
| Axis | Fixture | Naive |
|---|---|---|
| DS | 87 | 6 |
| Functional | 70 | 70 |
| A11y | 100 | 78 |
| Code | 98 | 99 |
| Build | 100 | 100 |
DS −81.3 · when the model ignores the design system · 10 regressions
Why
Coding agents can produce a settings page in seconds. They are much worse at using the design system you already have: they invent a Button, hardcode #8B5CF6, skip tokens, and ignore composition rules.
ds-eval turns that into a measurable question. Same tasks, same design system, different models or prompts — then a score you can compare and regress on.
- Compare modelsClaude, Codex, Gemini, Grok, or a local model — one suite, one report
- Catch prompt regressionsPrompt v12 vs v13: what improved, what broke, which cases flipped
- Use your own DSTokens, React components, Storybook, docs, and UX patterns as the source of truth
How It Works
Each eval case starts from an isolated starter app. The model receives the design system, the task prompt, and must generate working UI. The harness then builds, screenshots, and grades the result.
Design System
↓
Eval Task
↓
AI Coding Agent
↓
Generated UI
↓
────────────────────────
AST / DS Compliance
Runtime + Playwright
Accessibility (axe)
Visual Judge
UX Judge
Code Quality
────────────────────────
↓
Benchmark + RegressionRun Locally
Python eval engine. Fixture models run offline. Live models read keys from env — never from config.
cd ds-eval
python3 -m venv .venv
source .venv/bin/activate
pip install -e .
ds-eval run --model fixture --suite smoke
ds-eval run --model fixture-naive --suite smoke
ds-eval compare RUN_FIXTURE RUN_NAIVE
ds-eval dashboard --port 8001
With API keys:
export ANTHROPIC_API_KEY=…
export OPENAI_API_KEY=…
ds-eval run --model claude --suite smoke
ds-eval run --model codex --suite smokeCLI
ds-eval init scaffold a DS manifest
ds-eval list print the dataset
ds-eval run --model claude --limit 20
ds-eval compare RUN_1 RUN_2
ds-eval report ./runs/run-id
report writes report.json, report.html, and summary.md from the run artifacts. The table below is the offline fixture smoke run, not Claude or Codex.
| Model | Overall | DS | A11y | Code |
|---|---|---|---|---|
| fixture | 87.8 | 87 | 100 | 98 |
| fixture-naive | 59.3 | 5.7 | 78 | 99 |
Dataset
100 cases, written as real product tasks — not 100 copies of “make a button”. Easy / medium / hard, plus cases that deliberately conflict with DS rules.
| Category | Count | What it tests |
|---|---|---|
| Component | 25 | Correct use of Button, Select, Tabs, Dialog, Toast… |
| Patterns | 20 | Destructive confirm, empty states, bulk actions, multi-step |
| Pages | 25 | Settings, billing, permissions, employee profile |
| Accessibility | 10 | Keyboard, labels, ARIA, contrast, focus |
| Edge cases | 10 | Long strings, empty data, 100+ rows, mobile, i18n |
| Adversarial | 10 | Prompt pushes a raw color or extra actions the DS forbids |
Cases are YAML with a Pydantic schema: required components, forbidden patterns, assertions, rubric, viewport.
id: component-select-001
category: component
prompt: |
Create a country selector containing 10 countries.
It should have a label, helper text and error state.
requirements:
required_components: [Select]
forbidden: [native-select]
accessibility:
keyboard_navigation: true
label_required: trueGraders
The overall score is not a single LLM opinion. Deterministic graders and LLM judges are tracked separately on the dashboard.
- Static / AST deterministicDS imports vs raw
<button>, hardcoded colors, spacing tokens, duplicate components - Component compliance deterministicRequired / forbidden components, variants, props, composition rules
- Runtime deterministicBuild, Playwright, console errors, dialogs, tabs, form submit
- Accessibility deterministicaxe-core plus labels, keyboard, focus, semantics, ARIA
- Visual LLM judgeLayout, hierarchy, spacing, alignment, consistency, polish — each 0–5 with reasoning
- Product UX LLM judgeTask completion, pattern correctness, error prevention, feedback, IA
- Code quality LLM judgeStructure, duplication, types, React practices, unnecessary abstraction
Scoring
Weights are configurable. Default mix for the overall score:
| Axis | Weight | Source |
|---|---|---|
| DS Compliance | 25% | AST + component grader |
| Functional | 20% | Playwright interactions |
| Visual | 15% | Vision LLM, rubric 0–5 |
| Product UX | 15% | UX judge |
| Accessibility | 10% | axe + assertions |
| Code Quality | 10% | Static + LLM |
| Build Reliability | 5% | Install / build / boot |
Every generation and judge call stores input/output tokens, estimated cost, and latency. The dashboard shows benchmark cost and cost per eval.
Regression
Compare two runs — model versions or prompt versions — and see what moved.
ds-eval compare runs/sonnet-v1 runs/sonnet-v2
overall +2.7
DS compliance +4.1
Visual +1.2
Accessibility −3.4 ⚠
7 improved · 3 regressions · 90 unchanged
Each flipped case keeps before/after screenshots, grader explanations, and the generated source.
Dashboard
A local research UI over SQLite run history. Overview, models, runs, cases, and a comparison view for any two outputs of the same task.
- OverviewOverall score, categories, pass rate, cost, tokens, latency
- CasesFilter by model, category, difficulty, pass/fail, grader
- Case detailPrompt, rendered UI, screenshot, source, rubric, logs, cost
Custom Design System
Ships with a demo DS (Button, Input, Select, Dialog, Table, tokens, a few UX patterns) so the repo runs from clone. Point it at yours with a manifest:
name: Demo DS
version: 1.0
components:
path: ./src/components
tokens:
path: ./src/tokens
documentation:
path: ./docs
storybook:
url: http://localhost:6006
ds-eval init ./my-design-system
ds-eval run --system ./my-design-system --suite smokeLimitations
LLM judges are not ground truth. Visual and UX scores are subjective, even with a rubric and optional multi-judge setup. Deterministic graders (imports, tokens, axe, Playwright) are the reliable core; treat the rest as signal, not a verdict.
A high score means the model followed this design system on this suite. It does not mean the generated UI is production-ready for your product.
Published run
The 100-case table above is the shape of the dataset, not the size of this run. The published smoke lives in ds-eval/runs/fixture-good: model fixture, 10 cases, overall 87.8, cost $0. It is not a Claude or Codex call. The compare block under Regression shows the report shape, not this run.
| Case | Category | Overall | DS | A11y | Code |
|---|---|---|---|---|---|
| a11y-labels-001 | accessibility | 88.3 | 87.5 | 100.0 | 100.0 |
| adversarial-actions-001 | adversarial | 88.3 | 87.5 | 100.0 | 100.0 |
| adversarial-color-001 | adversarial | 88.3 | 87.5 | 100.0 | 100.0 |
| component-button-001 | component | 88.3 | 87.5 | 100.0 | 100.0 |
| component-input-001 | component | 88.3 | 87.5 | 100.0 | 100.0 |
| component-select-001 | component | 86.7 | 87.5 | 100.0 | 90.0 |
| edge-long-001 | edge-cases | 85.2 | 82.5 | 100.0 | 90.0 |
| page-settings-001 | pages | 88.3 | 87.5 | 100.0 | 100.0 |
| pattern-confirm-001 | patterns | 88.3 | 87.5 | 100.0 | 100.0 |
| pattern-empty-001 | patterns | 88.3 | 87.5 | 100.0 | 100.0 |
One case in full: component-select-001. The task is a country select with a label, helper, and error. The fixture answer is source/Task.tsx.
import { Select } from "@ds";
const COUNTRIES = ["Austria","Belgium","Canada","Denmark","Estonia","Finland","Germany","Hungary","Ireland","Japan"];
export function Task() {
return (
<Select label="Country" helper="Used for billing" error="Select a country to continue" defaultValue="">
<option value="">Choose one</option>
{COUNTRIES.map((name) => (
<option key={name} value={name}>{name}</option>
))}
</Select>
);
}Axes for this case. The visual and UX judges are off: a zero in the record means skipped, not a judgment of the screen. Their weight moves onto the other axes, which is why the overall is not zero.
| Axis | Score | |
|---|---|---|
| DS | 87.5 | |
| Functional | 70.0 | |
| Visual | 0.0 | skipped, weight moved to the other axes |
| UX | 0.0 | skipped, weight moved to the other axes |
| A11y | 100.0 | |
| Code | 90.0 | |
| Build | 100.0 | |
| Overall | 86.7 |
Pipeline
ds-eval scores the UI after the model has a context pack and, if needed, a repair pass:
ds-context prepares the pack
model writes UI
ui-repair patches violations
ds-eval scores the result ← you are here
prompt-regress diffs the next prompt