ds-eval

Can AI coding agents actually use your design system correctly? ds-eval is an open-source benchmark that runs the same UI tasks through different models and scores the generated interfaces.

Example · fixture models, not a live API runDS-aware vs naive · smoke suite10 evals · $0 · offline

Fixture87.8

Naive59.3

AxisFixtureNaive
DS876
Functional7070
A11y10078
Code9899
Build100100

DS −81.3 · when the model ignores the design system · 10 regressions

Why

Coding agents can produce a settings page in seconds. They are much worse at using the design system you already have: they invent a Button, hardcode #8B5CF6, skip tokens, and ignore composition rules.

ds-eval turns that into a measurable question. Same tasks, same design system, different models or prompts — then a score you can compare and regress on.

  • Compare modelsClaude, Codex, Gemini, Grok, or a local model — one suite, one report
  • Catch prompt regressionsPrompt v12 vs v13: what improved, what broke, which cases flipped
  • Use your own DSTokens, React components, Storybook, docs, and UX patterns as the source of truth

How It Works

Each eval case starts from an isolated starter app. The model receives the design system, the task prompt, and must generate working UI. The harness then builds, screenshots, and grades the result.

Design System
↓
Eval Task
↓
AI Coding Agent
↓
Generated UI
↓
────────────────────────
AST / DS Compliance
Runtime + Playwright
Accessibility (axe)
Visual Judge
UX Judge
Code Quality
────────────────────────
↓
Benchmark + Regression

Run Locally

Python eval engine. Fixture models run offline. Live models read keys from env — never from config.

cd ds-eval
python3 -m venv .venv
source .venv/bin/activate
pip install -e .

ds-eval run --model fixture --suite smoke
ds-eval run --model fixture-naive --suite smoke
ds-eval compare RUN_FIXTURE RUN_NAIVE
ds-eval dashboard --port 8001

With API keys:

export ANTHROPIC_API_KEY=…
export OPENAI_API_KEY=…
ds-eval run --model claude --suite smoke
ds-eval run --model codex --suite smoke

CLI

ds-eval init                         scaffold a DS manifest
ds-eval list                         print the dataset
ds-eval run --model claude --limit 20
ds-eval compare RUN_1 RUN_2
ds-eval report ./runs/run-id

report writes report.json, report.html, and summary.md from the run artifacts. The table below is the offline fixture smoke run, not Claude or Codex.

ModelOverallDSA11yCode
fixture87.88710098
fixture-naive59.35.77899

Dataset

100 cases, written as real product tasks — not 100 copies of “make a button”. Easy / medium / hard, plus cases that deliberately conflict with DS rules.

CategoryCountWhat it tests
Component25Correct use of Button, Select, Tabs, Dialog, Toast…
Patterns20Destructive confirm, empty states, bulk actions, multi-step
Pages25Settings, billing, permissions, employee profile
Accessibility10Keyboard, labels, ARIA, contrast, focus
Edge cases10Long strings, empty data, 100+ rows, mobile, i18n
Adversarial10Prompt pushes a raw color or extra actions the DS forbids

Cases are YAML with a Pydantic schema: required components, forbidden patterns, assertions, rubric, viewport.

id: component-select-001
category: component
prompt: |
  Create a country selector containing 10 countries.
  It should have a label, helper text and error state.
requirements:
  required_components: [Select]
  forbidden: [native-select]
  accessibility:
    keyboard_navigation: true
    label_required: true

Graders

The overall score is not a single LLM opinion. Deterministic graders and LLM judges are tracked separately on the dashboard.

  • Static / AST deterministicDS imports vs raw <button>, hardcoded colors, spacing tokens, duplicate components
  • Component compliance deterministicRequired / forbidden components, variants, props, composition rules
  • Runtime deterministicBuild, Playwright, console errors, dialogs, tabs, form submit
  • Accessibility deterministicaxe-core plus labels, keyboard, focus, semantics, ARIA
  • Visual LLM judgeLayout, hierarchy, spacing, alignment, consistency, polish — each 0–5 with reasoning
  • Product UX LLM judgeTask completion, pattern correctness, error prevention, feedback, IA
  • Code quality LLM judgeStructure, duplication, types, React practices, unnecessary abstraction

Scoring

Weights are configurable. Default mix for the overall score:

AxisWeightSource
DS Compliance25%AST + component grader
Functional20%Playwright interactions
Visual15%Vision LLM, rubric 0–5
Product UX15%UX judge
Accessibility10%axe + assertions
Code Quality10%Static + LLM
Build Reliability5%Install / build / boot

Every generation and judge call stores input/output tokens, estimated cost, and latency. The dashboard shows benchmark cost and cost per eval.

Regression

Compare two runs — model versions or prompt versions — and see what moved.

ds-eval compare runs/sonnet-v1 runs/sonnet-v2
overall +2.7
DS compliance +4.1
Visual +1.2
Accessibility −3.4 ⚠

7 improved · 3 regressions · 90 unchanged

Each flipped case keeps before/after screenshots, grader explanations, and the generated source.

Dashboard

A local research UI over SQLite run history. Overview, models, runs, cases, and a comparison view for any two outputs of the same task.

  • OverviewOverall score, categories, pass rate, cost, tokens, latency
  • CasesFilter by model, category, difficulty, pass/fail, grader
  • Case detailPrompt, rendered UI, screenshot, source, rubric, logs, cost

Custom Design System

Ships with a demo DS (Button, Input, Select, Dialog, Table, tokens, a few UX patterns) so the repo runs from clone. Point it at yours with a manifest:

name: Demo DS
version: 1.0
components:
  path: ./src/components
tokens:
  path: ./src/tokens
documentation:
  path: ./docs
storybook:
  url: http://localhost:6006
ds-eval init ./my-design-system
ds-eval run --system ./my-design-system --suite smoke

Limitations

LLM judges are not ground truth. Visual and UX scores are subjective, even with a rubric and optional multi-judge setup. Deterministic graders (imports, tokens, axe, Playwright) are the reliable core; treat the rest as signal, not a verdict.

A high score means the model followed this design system on this suite. It does not mean the generated UI is production-ready for your product.

Published run

The 100-case table above is the shape of the dataset, not the size of this run. The published smoke lives in ds-eval/runs/fixture-good: model fixture, 10 cases, overall 87.8, cost $0. It is not a Claude or Codex call. The compare block under Regression shows the report shape, not this run.

CaseCategoryOverallDSA11yCode
a11y-labels-001accessibility88.387.5100.0100.0
adversarial-actions-001adversarial88.387.5100.0100.0
adversarial-color-001adversarial88.387.5100.0100.0
component-button-001component88.387.5100.0100.0
component-input-001component88.387.5100.0100.0
component-select-001component86.787.5100.090.0
edge-long-001edge-cases85.282.5100.090.0
page-settings-001pages88.387.5100.0100.0
pattern-confirm-001patterns88.387.5100.0100.0
pattern-empty-001patterns88.387.5100.0100.0

One case in full: component-select-001. The task is a country select with a label, helper, and error. The fixture answer is source/Task.tsx.

import { Select } from "@ds";

const COUNTRIES = ["Austria","Belgium","Canada","Denmark","Estonia","Finland","Germany","Hungary","Ireland","Japan"];

export function Task() {
  return (
    <Select label="Country" helper="Used for billing" error="Select a country to continue" defaultValue="">
      <option value="">Choose one</option>
      {COUNTRIES.map((name) => (
        <option key={name} value={name}>{name}</option>
      ))}
    </Select>
  );
}

Axes for this case. The visual and UX judges are off: a zero in the record means skipped, not a judgment of the screen. Their weight moves onto the other axes, which is why the overall is not zero.

AxisScore
DS87.5
Functional70.0
Visual0.0skipped, weight moved to the other axes
UX0.0skipped, weight moved to the other axes
A11y100.0
Code90.0
Build100.0
Overall86.7

Pipeline

ds-eval scores the UI after the model has a context pack and, if needed, a repair pass:

ds-context prepares the pack
model writes UI
ui-repair patches violations
ds-eval scores the result ← you are here
prompt-regress diffs the next prompt

Source