Design System Agent
An agent that keeps every product honest to the design system
- Role
- Product Designer — agent design & scoring model
- Company
- Campminder
- Team
- Design systems team, used by every consuming product team
- Timeline
- 2 weeks to first working version
My role focused on
At a glance
The problem
Every product drifted from the design system a little differently, and nobody could see the drift until a designer happened to notice it in review.
The insight
Built an agent that treats the design system's own code as ground truth, scores any repo against it, and writes fixes in language a PM could read.
The impact
Turned a subjective “this doesn't look right” into a concrete score, a ranked list of fixes, and a trail of how adherence changes over time.
Foundations audited
Component adoption, color, typography, spacing, radius
Setup per new repo
Reads the design system's own code, not a manifest
New components proposed
Flags recurring custom patterns as candidates to add upstream
01, Context
Adherence lived in one reviewer's head
Design-system adherence depended on whichever reviewer happened to notice a raw `<button>` or a hardcoded hex value in a PR. There was no repo-wide picture of how far a product had drifted, and no way to catch it before merge instead of months later.
Even the design system's own documentation couldn't always be trusted, a `DESIGN.md` can say one typeface while the real, shipped CSS says another. Adherence needed a source of truth that couldn't quietly go stale.
Why it matters
Manual
design-system review, one PR at a time
Adherence depended on whichever reviewer happened to notice
Drift
nobody could see accumulating
No repo-wide picture of how far a product had drifted
Docs vs. code
quietly disagreed with each other
A design doc could say one thing while the shipped code said another
0
path from finding to fix
Even when drift was found, turning it into a concrete fix was manual
02, How it works
Ground truth is code, not docs
The agent syncs a read-only copy of the design system's real code, then scans the target repo, scores it, and writes both a plain-language report and a local dashboard with the adherence score and its trend over time.
Adherence is defined as one thing, and only one thing: does the product use the design system's component at all? Color, typography, spacing, and radius are scored separately, so a messy-but-adopted component is never confused with one that was never adopted.
campminder/registration
Audited against campminder-ui (@campminder/ds), the design system repo.
Last run: Aug 11, 2026 · scope: full · repo @a3f9c21
Adherence
▲ +5 pts vs. previous
At a glance
Issues by severity
Foundations & component adoption
How consistently this repo uses what campminder-ui defines.
3 not adopted, 5 out of sync with the DS.
2 hardcoded color values found, most often needing --color-accent.
Fully on the DS typeface.
4 off-scale spacing values found (not a multiple of the 4px grid).
1 off-scale radius value found.
Component adoption ledger
Every component in the registry, checked against this repo.
Custom components, no DS equivalent
Built locally, candidates to propose upstream.
184 files scanned · generated by the audit-ds agent — some DS components are persona-specific; “not adopted” doesn’t always mean “should be adopted here.”
Syncs the real design system
Pulls a read-only copy of the design system's own code, the actual source of truth, not whatever a README claims.
Scores what's real, not what's tidy
Adherence is one question: does the product use the design system's component at all? Foundations are scored separately alongside it.
Diff mode for every PR
Defaults to auditing just the diff against main, fast enough to run before a merge, not just as a quarterly audit.
Full-repo mode for real debt
A whole-repo scan surfaces the adoption picture no diff can show, and tracks the score's trend run over run.
Proposes, doesn't just flag
Custom components with no design-system match are surfaced as candidates to propose upstream, not violations to feel bad about.
Plain language first
The report leads with a short recap a PM could read, no jargon, before any of the technical detail.
03, Key decisions
The calls that made the score trustworthy
A scoring tool nobody trusts gets ignored after the first false positive. These decisions were about earning that trust.
Adherence = adoption, full stop
Averaging in color, spacing, and typography used to conflate “uses the design system” with “styled cleanly.” Splitting them out means the headline score answers one honest question.
Code is ground truth, never a doc
A design system's own docs can drift from its shipped code. The agent always reads the actual CSS and components, never assumes a markdown file is current.
Custom components are candidates, not violations
A hand-built component might be the design system's next component, not a mistake. The report frames it as something to propose upstream.
Plain language before any jargon
The person reading the report might be a PM, not an engineer. The recap always comes first, in one paragraph, with no unexplained jargon.
Diff by default, full scan on request
Most asks are “does this PR introduce drift,” which needs to be fast. A full scan is reserved for when someone actually wants the whole adoption picture.
04, Results
From a gut feeling to a number you can track
The run below is an illustrative example, not one specific team's real audit, but the shape of it (adoption climbing, drift caught before merge, custom components proposed instead of duplicated) matches what happened in practice.
Adherence score
Component adoption, the headline metric
Findings caught pre-merge
In a single diff-mode run, before it reached review
Custom components flagged
Proposed upstream instead of quietly duplicated
Setup time for a new repo
Point it at any consuming repo and run
Reflection
What I took from it
Designing the scoring model taught me more about design-system governance than any audit I'd done by hand. Deciding what adherence should even mean (adoption, not tidiness) was the real design work, the agent just made it consistent.
The plain-language-first rule was the detail I almost skipped and I'm glad I didn't: a report only a design-systems engineer can parse doesn't change anyone's behavior. One a PM can read does.
Next project
Prickly Pear