← Back to work
Design SystemsAI Product DesignInternal tooling

Design System Agent

An agent that keeps every product honest to the design system

Role
Product Designer — agent design & scoring model
Company
Campminder
Team
Design systems team, used by every consuming product team
Timeline
2 weeks to first working version

My role focused on

Scoring model designAgent behavior specDashboard designPlain-language reportingDesign system governance

At a glance

01

The problem

Every product drifted from the design system a little differently, and nobody could see the drift until a designer happened to notice it in review.

02

The insight

Built an agent that treats the design system's own code as ground truth, scores any repo against it, and writes fixes in language a PM could read.

03

The impact

Turned a subjective “this doesn't look right” into a concrete score, a ranked list of fixes, and a trail of how adherence changes over time.

Foundations audited

5

Component adoption, color, typography, spacing, radius

Setup per new repo

0config files needed

Reads the design system's own code, not a manifest

New components proposed

Auto

Flags recurring custom patterns as candidates to add upstream

01, Context

Adherence lived in one reviewer's head

Design-system adherence depended on whichever reviewer happened to notice a raw `<button>` or a hardcoded hex value in a PR. There was no repo-wide picture of how far a product had drifted, and no way to catch it before merge instead of months later.

Even the design system's own documentation couldn't always be trusted, a `DESIGN.md` can say one typeface while the real, shipped CSS says another. Adherence needed a source of truth that couldn't quietly go stale.

Why it matters

Manual

design-system review, one PR at a time

Adherence depended on whichever reviewer happened to notice

Drift

nobody could see accumulating

No repo-wide picture of how far a product had drifted

Docs vs. code

quietly disagreed with each other

A design doc could say one thing while the shipped code said another

0

path from finding to fix

Even when drift was found, turning it into a concrete fix was manual

02, How it works

Ground truth is code, not docs

The agent syncs a read-only copy of the design system's real code, then scans the target repo, scores it, and writes both a plain-language report and a local dashboard with the adherence score and its trend over time.

Adherence is defined as one thing, and only one thing: does the product use the design system's component at all? Color, typography, spacing, and radius are scored separately, so a messy-but-adopted component is never confused with one that was never adopted.

.ds-audit/dashboard.html

campminder/registration

Audited against campminder-ui (@campminder/ds), the design system repo.

Last run: Aug 11, 2026 · scope: full · repo @a3f9c21

81%

Adherence

▲ +5 pts vs. previous

At a glance

Files scanned184
Not adopted3
Out of sync5
Custom, no DS match5

Issues by severity

High3
Medium2
Low3

Foundations & component adoption

How consistently this repo uses what campminder-ui defines.

Component adoption81%

3 not adopted, 5 out of sync with the DS.

Color tokens88.4%

2 hardcoded color values found, most often needing --color-accent.

Typography100%

Fully on the DS typeface.

Spacing91.2%

4 off-scale spacing values found (not a multiple of the 4px grid).

Radius83.3%

1 off-scale radius value found.

Component adoption ledger

Every component in the registry, checked against this repo.

Buttonv0.5.0Current
Inputv0.5.0Current
Selectv0.5.0Current
Badgev0.5.0Current
Tablev0.3.1→v0.5.0Outdated
Dialogv0.4.0→v0.5.0Outdated
Tooltipno stampUnstamped
Comboboxv0.5.0 avail.Not adopted
DataTablev0.5.0 avail.Not adopted

Custom components, no DS equivalent

Built locally, candidates to propose upstream.

src/components/CampCard.tsxNot in DS
src/components/PaymentPlanSummary.tsxNot in DS
src/components/RosterFilterBar.tsxNot in DS

184 files scanned · generated by the audit-ds agent — some DS components are persona-specific; “not adopted” doesn’t always mean “should be adopted here.”

Syncs the real design system

Pulls a read-only copy of the design system's own code, the actual source of truth, not whatever a README claims.

Scores what's real, not what's tidy

Adherence is one question: does the product use the design system's component at all? Foundations are scored separately alongside it.

Diff mode for every PR

Defaults to auditing just the diff against main, fast enough to run before a merge, not just as a quarterly audit.

Full-repo mode for real debt

A whole-repo scan surfaces the adoption picture no diff can show, and tracks the score's trend run over run.

Proposes, doesn't just flag

Custom components with no design-system match are surfaced as candidates to propose upstream, not violations to feel bad about.

Plain language first

The report leads with a short recap a PM could read, no jargon, before any of the technical detail.

03, Key decisions

The calls that made the score trustworthy

A scoring tool nobody trusts gets ignored after the first false positive. These decisions were about earning that trust.

Scoring

Adherence = adoption, full stop

Averaging in color, spacing, and typography used to conflate “uses the design system” with “styled cleanly.” Splitting them out means the headline score answers one honest question.

One clear scoreNo conflation
Rigor

Code is ground truth, never a doc

A design system's own docs can drift from its shipped code. The agent always reads the actual CSS and components, never assumes a markdown file is current.

Reads real codeNever assumes
Tone

Custom components are candidates, not violations

A hand-built component might be the design system's next component, not a mistake. The report frames it as something to propose upstream.

Custom code foundPropose, don't shame
Communication

Plain language before any jargon

The person reading the report might be a PM, not an engineer. The recap always comes first, in one paragraph, with no unexplained jargon.

Plain recapJargon after
Scope

Diff by default, full scan on request

Most asks are “does this PR introduce drift,” which needs to be fast. A full scan is reserved for when someone actually wants the whole adoption picture.

Fast by defaultFull scan on ask

04, Results

From a gut feeling to a number you can track

The run below is an illustrative example, not one specific team's real audit, but the shape of it (adoption climbing, drift caught before merge, custom components proposed instead of duplicated) matches what happened in practice.

Adherence score

81%+29 pts since first run

Component adoption, the headline metric

Findings caught pre-merge

7

In a single diff-mode run, before it reached review

Custom components flagged

5

Proposed upstream instead of quietly duplicated

Setup time for a new repo

0 minno config needed

Point it at any consuming repo and run

Reflection

What I took from it

Designing the scoring model taught me more about design-system governance than any audit I'd done by hand. Deciding what adherence should even mean (adoption, not tidiness) was the real design work, the agent just made it consistent.

The plain-language-first rule was the detail I almost skipped and I'm glad I didn't: a report only a design-systems engineer can parse doesn't change anyone's behavior. One a PM can read does.

Next project

Prickly Pear

→