Showcase · AI & Design Systems · 03 of 03

When the system teaches the machine.

AI agents generate UI faster than any team. The question isn't speed - it's whether the output belongs to your system. A practitioner's perspective on AI in design systems.

At a glance

AI changes what design systems arefor.

Design systems used to serve designers and developers. Now they also serve AI agents. This page explores what that shift means - from how we structure tokens to how we govern what machines produce.

New user class

AI agents are now consumers of your design system — alongside designers and developers.

Context engineering

Structured rules files and MCP servers give AI the knowledge to generate system-compliant UI.

Intent architecture

Components named by purpose, not appearance. Semantic tokens that encode meaning for machines.

Agentic governance

AI lints first, humans review edge cases. The DS lead becomes an architect of constraints.

The core question

AI doesn't replace any part of the workflow. It changes thenature of every part.

1.Human defines

Token scales, component APIs, usage rules, governance tiers, brand decisions, quality bars.

2.AI generates

Variants, platform specs, dark mode, documentation, lint checks, compliance reports.

3.Human validates

Edge cases, taste calls, architectural decisions, ambiguous outputs, strategic direction.

01 / 06 · The Workflow

Who does what — and where theline moves.

Six phases of DS work. The line between human and AI shifts at every phase.

01 Token Definition

Human

Designs the scale, names the tiers, sets brand decisions.

AI Agent

Generates 30 token sets from one source.

5 themes × 2 modes × 3 platforms

02 Figma ComponentsOptional

Human

Sets design direction, interaction patterns, a11y behavior.

AI Agent

Generates the full variant matrix in minutes.

Can be skipped when token docs, component APIs, and usage rules are comprehensive enough for agents to work from code directly.

Weeks → minutes

03 Code Components

Human

Architects the component API — props, events, slots.

AI Agent

Fills 75 variants from rules.

5 types × 3 sizes × 5 states

04 Usage Rules

Human

Writes do/don’t guidelines and governance tiers.

AI Agent

Enforces them at generation time.

Review comment → build constraint

05 Documentation

Human

Structures the IA, writes the “why.”

AI Agent

Generates and syncs the “what.”

Docs never drift from code

06 Quality Assurance

Human

Reviews edge cases and taste calls.

AI Agent

Catches 80% instantly.

First-pass: tokens, a11y, naming

“The pattern is the same at every phase: Humans set intent, constraints, and quality bars. AI handles volume, variants, and enforcement.”

That's the overview. Now zoom into the layers underneath.

02 / 06 · Context Engineering

The new DS skill: writing rulesmachines can follow.

Without context, the agent guesses. Four layers later, it generates system-native code.

Output quality Consistent colors - 30%
Design Tokens

Design tokens are named values — --spacing-4 instead of 16px, --color-primary instead of #0369A1. They’re the atoms of your design language. When an AI agent can read them, it stops inventing its own values.

Primitive and semantic values: colors, spacing, typography, elevation. This is the visual vocabulary the agent draws from.

--color-action-primary: var(--sky-700);
--spacing-4: 1rem;
--radius-md: 0.5rem;
Component APIs

A component API defines what a button CAN be — which variants exist, which sizes are allowed, which events it emits. Without this, the agent writes custom HTML for every button it creates.

Prop contracts, variant maps, slot definitions, composition rules. The structural grammar that tells the agent how to build.

<ds-button
  variant="primary | secondary | ghost"
  size="sm | md | lg"
  (click)="handler">
  {{ label }}
</ds-button>
Usage RulesKey layer

Imagine a linting rule, but for design decisions. “Primary buttons: max one per page.” “Destructive actions: always require confirmation.” These rules existed in your team’s heads. Now they exist in a file an AI agent can enforce.

This is the layer most teams skip — and the one that matters most for AI.

Tokens give the agent your palette. APIs give it your structure. But rules give it judgment. The difference between “consistent output” and “correct output.”

# Button — Usage Rules
- primary: max 1 per viewport
- ghost: tertiary actions only
- all: min touch target 44×44px
- danger: requires confirmation dialog
- WCAG: contrast ≥ 4.5:1
MCP Server

MCP — Model Context Protocol — is a protocol Anthropic created in late 2024. It lets AI tools (Claude, Cursor, VS Code Copilot) connect to external data sources like your Figma files, token database, and component docs. Figma launched their MCP server in 2025. It went GA at Schema 2025 — available on all plans.

The bridge between your system and the model. The agent reads your live system — not cached documentation, not outdated wikis.

// Agent calls MCP:
get_component("Button")
→ returns: variants, props, tokens,
   usage rules, do/don’t examples

// Generates with full context

Theory is useful. But seeing it happen is better.

03 / 06 · Same prompt, three systems

Same prompt. Three times.Only the system changes.

The model doesn't get smarter. The infrastructure behind it does. Watch what happens when the DS team invests in tokens, intent-based naming, and an MCP server - without changing a single word in the prompt.

AI Agent - Design System Chat
HC

“Build a settings page with user preferences”

AI

Here’s a settings page:

No context output.html
<div style="background: #f3f4f6;
  padding: 24px; border-radius: 8px;">
  <h2 style="font-size: 20px;
    font-weight: bold; color: #111827;">
    Settings
  </h2>
  <div style="margin-top: 16px;">
    <label>Name</label>
    <input style="border: 1px solid #d1d5db;
      padding: 8px; width: 100%;" />
  </div>
  <div style="margin-top: 16px;">
    <label>Email</label>
    <input style="border: 1px solid #d1d5db;
      padding: 8px; width: 100%;" />
  </div>
  <button style="background: #3B82F6;
    color: white; padding: 8px 16px;
    margin-top: 16px;">
    Save
  </button>
</div>

The agent guesses. Hardcoded hex values, inline styles, flat HTML. Every team member who runs this prompt gets different output.

Context improves

What the DS team built

  • Defined a 3-tier token architecture (primitive → semantic → component)
  • Named tokens by intent: --color-action-primary, not --blue-500
  • Documented component APIs with props, slots, and states
  • Published everything in a structured rules file the agent can read
HC

“Build a settings page with user preferences”

Same prompt as before.

AI

Here’s a settings page using your tokens:

Tokens + APIs output.html
<div class="bg-surface p-spacing-6
  rounded-radius-lg">
  <h2 class="text-heading-lg
    text-on-surface">Settings</h2>
  <div class="mt-spacing-4">
    <label class="text-label">Name</label>
    <input class="form-input" />
  </div>
  <div class="mt-spacing-4">
    <label class="text-label">Email</label>
    <input class="form-input" />
  </div>
  <button class="btn bg-action-primary
    text-on-action rounded-radius-md">
    Save
  </button>
</div>

Colors are right. Spacing follows the scale. But the structure is still improvised — the agent knows what to paint with, not how to compose. Component names describe appearance (CardBase, TextFieldBox), not purpose.

Context improves

What the DS team added

  • Renamed components by intent: CardBase → FeatureHighlight, TextFieldBox → SecureInput
  • Added usage rules: "email fields are PII → require consent flag, mask display"
  • Connected everything via an MCP server — a live API the agent queries at generation time
HC

“Build a settings page with user preferences”

Same prompt as before.

AI

Thinking: Settings page → SettingsLayout. Email field = PII → SecureInput with consent flag. Save = primary action → max 1 per viewport. Checking usage rules via MCP…

Here’s a settings page using your design system:

Full system (MCP) output.html
<ds-page-layout variant="settings">
  <ds-page-header>
    <ds-heading level="2">
      {{ pageTitle }}
    </ds-heading>
  </ds-page-header>
  <ds-form (submit)="save.emit()">
    <ds-form-field label="Name">
      <ds-input [(value)]="name" />
    </ds-form-field>
    <ds-form-field label="Email">
      <ds-secure-input
        sensitivity="pii"
        mask="partial"
        autocomplete="off"
        data-gdpr="consent-required" />
    </ds-form-field>
    <ds-button variant="primary"
      type="submit">Save</ds-button>
  </ds-form>
</ds-page-layout>

Same model. Same prompt. Three times. The difference is the system behind it.

What changed

Colors
Hardcoded hex
Token-based
Token-based
Structure
Flat HTML
Flat HTML
Component tree
Data
Static strings
Static strings
Bound props
Rules
None
None
Enforced at generation

The agent knows the system. Now — who watches the agent?

04 / 06 · Agentic Governance

The DS lead's job isn'tbuilding components anymore.

Four review gates: three automated, one human. Each catches what the others would miss.

Component submitted
AI Lint Check Automated

Gate 1 — Automated

Linting for design consistency — the same way ESLint catches code style issues, this catches design system drift. Hardcoded colors, wrong spacing values, components that aren’t in the library.

Scans generated code against the design system rules file. Checks token usage, naming conventions, component API compliance.

⚠️  Line 12: color #3B82F6
→ Use token: var(--action-primary)

✔  Line 24: <ds-button variant="primary">
→ Correct component usage
Accessibility Audit Automated

Gate 2 — Automated

Accessibility isn’t optional — it’s structural. This gate checks the mechanical parts: contrast ratios above 4.5:1, ARIA labels on interactive elements, touch targets above 44×44px.

Automated validation of contrast, ARIA attributes, touch targets, keyboard navigation paths.

⚠️  Button missing aria-label
→ Icon-only buttons require labels

✔  Contrast ratio 6.2:1
→ Meets WCAG AA (min 4.5:1)
Naming Convention Automated

Gate 3 — Automated

Names are the first thing an AI reads. RedAlertBox tells the next agent nothing. CriticalNotification tells it everything — severity, category, when to use it.

Validates component names, CSS class names, and file structure against system naming patterns.

⚠️  Component: "RedAlertBox"
→ Use intent-based name:
   "CriticalNotification"

✔  Token: --color-border-critical
→ Matches naming convention
Human Review Human judgment

Gate 4 — Human judgment

This is the gate that can never be automated. A component can pass every lint rule and still be wrong for its context. “Is this the right component for this situation?” That’s judgment — and it’s where the DS lead’s value lives.

The DS lead reviews edge cases, architectural decisions, and anything the automated pipeline flagged as ambiguous. This is where taste lives.

Flagged for human review:

• New variant "warning" requested
  → Requires DS team approval
  → Check brand guidelines

• Modal inside Modal detected
  → Architectural concern
Production

Structure is leverage. Now there are numbers to prove it.

05 / 06 · The Business Case

For the first time, this work hasnumbers.

Design system investment used to live in the realm of “trust us, it pays off.” Now teams like Storybook are measuring agent performance against system quality - speed, token cost, code conformity. The results are unambiguous.

Generation speed

Fewer tokens, shorter prompts when the DS is the agent’s source of truth. Storybook benchmarks confirm the direction.

Token cost

Illustrative range. Design systems with automated consistency checks reduce visual QA rework significantly — exact gains depend on system maturity and team size.

Code conformity

First-pass accuracy jumps from guesswork to system-native when the agent has tokens, APIs, and rules. Based on Atlassian and comparable DS team reports.

“A well-structured system produces faster, cheaper, more accurate output. A poorly structured one produces guesswork that costs more to fix than it saved to generate.”

The creative paradox

Agents assemble.They don't create.

More control, not less

Designers define the vocabulary - component names, states, descriptions. Agents compose from that vocabulary. The richer the system, the more expressive the output. This isn't automation replacing creativity. It's creativity becoming infrastructure.

The gravity toward generic

Efficiency pulls toward assembly-optimized, lowest-common-denominator systems. Without deliberate creative investment upstream, you get fast systems producing forgettable output - which costs more in the long run. The designer's job is to ensure there's something worth assembling.

Every powerful tool has a shadow side.

06 / 06 · The Hard Questions

What wedon't talk about enough.

Optimism is easy. What earns trust is honesty about the risks.

  1. The de-skilling risk

    When AI handles component creation, junior designers shift from active creators to passive monitors. The foundational skills — understanding spacing, hierarchy, interaction patterns — don’t develop through reviewing agent output. This is the classic automation irony — the systems we build because they outperform humans at routine tasks still require human oversight, by humans whose skills are atrophying because they no longer practice.

    We automate because AI outperforms juniors at routine tasks — then wonder why juniors never develop into seniors. Without intentional apprenticeship models alongside AI tooling, we optimize the present at the cost of the next generation.

  2. Aesthetic homogenization

    Recent research is already framing AI-generated UI homogenization as a structural risk. The pattern echoes a broader finding from generative-AI creativity studies: individual outputs score higher on creativity, but the collective set shows measurably less diversity. Design systems already push toward consistency by definition. Adding AI doubles down on convergence.

    There’s a difference between consistent and generic. Design systems enforce the former. Without careful governance, AI accelerates the latter. The systems that win will be those that encode brand distinctiveness, not just compliance.

  3. Plausible hallucinations

    The hardest bugs to catch are the ones that look right. An AI agent can generate a component that uses your token names, follows your naming conventions, passes linting — but references a token that doesn’t exist or misapplies a variant for the context. Atlassian’s design system team reported this as their #1 challenge: plausible-looking but wrong outputs that broke things — wrong token names, on-brand-looking icons that didn’t exist.

    The response is not to distrust AI. It’s to design your system so hallucinations are structurally impossible — pre-coded templates, JSON schemas for deterministic elements, and mandatory human review gates for anything compositional.

  4. The monitoring paradox

    Automation changes the volume that needs review. A human can carefully review 5 components in a sprint; an AI generates 50 in an hour. “Human in the loop” starts as quality assurance and ends as rubber-stamping the moment generation outpaces attention. This isn’t a sign that the team is sloppy — it’s structural: the same speed that justifies adopting AI also makes meaningful human review uneconomic.

    Designing the oversight layer becomes its own design problem. Which 10% of outputs get human review? Where do you put the mandatory gates the system can’t bypass? The DS lead’s role shifts from reviewing components to designing the review pipeline — and the pipeline matters more than the generation it monitors.

The best design systems don't just serve humans. They teach machines - and the people who build them know when to let the machine work, and when to pull the brake.

Checklist

Is your systemagent-ready?

Seven structural prerequisites. If an agent can't read it, it can't use it.

  1. Semantic token layer.

    Primitives → semantic → component tokens. The semantic layer is non-negotiable — it’s what agents actually read when generating UI.

  2. Parity between Figma and code props.

    Same property names, same values on both sides. Without parity the agent guesses the mapping — and gets it wrong.

  3. Complete interactive states.

    Every state (hover, focus, disabled, error, loading) modeled in Figma variants, Storybook stories, and the a11y spec. Missing states become missing outputs.

  4. Explicit composition slots.

    Compound components declare their named slots (Card.Header, Dialog.Footer, etc.) in code and docs. Without them, agents reinvent wrappers instead of composing what already exists.

  5. Structured, machine-readable docs.

    Usage rules, anatomy, prop descriptions, and examples live in parseable formats — MDX frontmatter, component manifests, llms.txt — not buried in wiki pages an agent will never crawl.

  6. Accessibility metadata on the contract.

    ARIA roles, keyboard patterns, and required labels declared at the component API level. Agents skip a11y unless it’s encoded — make it structural, not documentary.

  7. Types as the contract.

    Discriminated unions over booleans, JSDoc on props, strict TypeScript. Claude, Cursor, and Copilot read types before docs — precise types prevent invalid variant combinations at generation time.

§ Continue

Build for systems, not screens.

That's the AI argument. Two more showcases follow — the system itself, and the UX process behind it.