Skip to content

<design-systems> By Cris Morales

What Does AI-Readiness Mean in Design Systems?

We measured how coding agents can consume the systems we build using an external tool. AI-readiness comes down to great design and engineering.

A few weeks back, I was talking to Christoph Hellmuth about the full range of being a Design Engineer. His background is in engineering, and my background is actually from advertising and brand. The full spectrum right there. We talk about coding capabilities but that’s a totally different article. TJ already wrote his thoughts on this, he called it The Rise of the Alicorn.

Chris and I agree on what the mindset of a design engineer in 2026 is. Explore, experiment and combine different surfaces and digital material, keyword “experiment”.

Experiment. verb [ I ], /ɪkˈsper·əˌment/:
Actively try out new methods, ideas, or techniques to observe their outcomes.

You will need a question, a hypothesis, variables, methodology, and analysis of the output. Pretty much what he is building on his Open Design System Bench project. It’s a plug-in benchmark that runs coding agents against your component library and grades the output on different axes. The high profile test involves putting agents through 3 different scenarios x 3 times each x 10 design tasks.

Table titled System A, one production React design system. A headless coding agent got ten intent-level tasks in a throwaway app, three context levels, three repetitions each: 90 graded generations scored on six axes. 69 components, 268 public exports and 601 documented props, 0 errored generations. Cost and wall clock went from $158.90 and 7h 35m to $94.30 and 5h 57m after the harness refactor.

The results have been surprisingly interesting, not only surfacing components or token issues, but also validating our hypothesis at Southleft, that great engineering and design practices beat complex harnesses.

AI-readiness is about discoverability and readability.

We can measure this on two dimensions, repository readiness (inside the system) and consumption readiness (outside):

Inside means how the agent can navigate your design system repo and complete design tasks. This is particularly important for your design system team. The outside dimension is about the agent using your design system properly when using tools like MCP, and/or just reading its documentation.

Grace Han posted her AI-Readiness scale earlier this year. “Traditional maturity models focused on adoption: reuse, governance, ownership, and contribution. That still matters. But AI adds another layer: 
Can your system be safely understood and operated by machines across functions?”. Same question.

Great engineering practices beat agent harnesses

Let me come back to the idea that great engineering beats harnesses. Last week we ran tests using the Open Design System Bench on a few projects. AI-readiness scored over 90 on the internal dimension and around 40–50 on the external, this one is our estimation, because our test ran directly on the codebase, not through CLI, MCP, other tooling, or just from documentation. That’s expected, because that’s on each client team’s side.

We also ran the same 90 cells on Astryx, Meta’s design system, eight years inside Meta and open-sourced in June 2026. The result was pretty much a tie: 95.68 against 96.19, and the same 77 mechanically perfect cells. A system in its first phase landing in the same band as one with eight years behind it.

Each cell measures imports, API fidelity, token discipline, a11y, compile, and judgment. All of those are measured mechanically and are therefore heavily influenced by agent discoverability, with judgment being the exception.

How well can the agent navigate, search, find elements in your repo, and understand how everything works?

Table titled Five axes ask whether the agent understood the system, one asks whether it understood the design. Weight and mean per axis: imports 10%, 100.0; apiFidelity 25%, 97.6; tokenDiscipline 15%, 100.0; a11yStatic 10%, 100.0; compile 10%, 78.9; judgment 30%, 89.8. One task scored 100 on all five mechanical axes and 30 on judgment, three separate times.

Clean code is machine-readable code.

This depends a lot on two principles: separation of concerns and structured discoverability.

Separation of concerns means splitting a computer program into distinct sections, where each section addresses a separate, single piece of responsibility (a “concern”). This enables structured discoverability, a concept similar to Progressive Disclosure in UX: sequencing information and actions across several layers to avoid overwhelming the user, but for agents.

Simply put: one truth in one place. Everything else redirects to it, working as an index.

A very interesting finding is that prose files like AGENTS.md, llms.txt, or whatever.md file, including agent skills, can produce poor-quality output by providing stale instructions. The problem is stale, duplicated, or authoritative-looking information can conflict with the source of truth. That makes them an additional source of hallucination. AI instructions are only valuable when they remain trustworthy.

Table titled One wrong word in the steering document cost ten generations. AGENTS.md listed error as a status tone, but the library's canon is danger. 19 of 90 generations failed hard, 10 traced to that one wrong value, 0 of those in the bare context, 9 were ordinary TypeScript errors. 147 rejected names were curated across the component specs, 0 reached the machine-readable contract.

We saw bare runs (without checking harness files or skills) scoring higher on the mechanical dimensions, in both systems we tested. This happens mostly when those files hardcode information that should be discovered dynamically, like a component or token list. The bench tasks require the agents to build specific UIs using your component library; the agent looks at those files and generates a component that already exists. Sometimes these files even hold ghost tokens, because the agent that generated the file hallucinated them.

Table titled More guidance didn't mean better output, comparing bare, AGENTS.md and skill contexts. Mean score 95.5, 93.0, 94.2. Compiles against the real types 90.0, 70.0, 76.7. Design judgment 89.3, 89.6, 90.4. API fidelity 98.7, 96.5, 97.5. Hard failures 3, 9 and 7 of 30. Clean passes 24, 19 and 19 of 30.

Time to put on my hard hat and refactor the whole agent harness following these principles. Added a few new docs in strategic places, and removed 90% of the AGENTS.md content. Tested again in high profile (90 cells) and, to my surprise, we scored lower (from 90 to 72). I totally missed that we had a prop list in the file, but I knew we had the right auto-generated contract in place; the file just needed to point to it. Added 107 lines and improved the validation gate. Now we sit at the same 90 points baseline with a document that is 90% smaller. This also reduced the test token cost by around 41%, and time to complete around 22%.

Table titled The fix was to stop hand-writing the vocabulary and derive it from the prop catalog, before and after. agents-md compile 70.0 to 90.0, skill compile 76.7 to 80.0, vocabulary compile failures 6 to 0, other compile failures 10 to 3, 107 lines added to AGENTS.md, rejected names in the machine-readable contracts 0 to 147.

Judgment measures something different

Mechanical correctness ≠ design quality. Whether the agent made the right design decisions is a completely different beast. Was it the correct component for the intent, the correct emphasis hierarchy, the correct interaction pattern? This is different from whether the code compiles or the API is used correctly. It’s totally possible to score 100 on the mechanical dimensions and score 30 here. We know that, because it happened to us.

The lowest task in both systems was a password field with a show/hide button. The text input had no slot for that button, so agents built their own on top of it. Every mechanical check passed, judgment failed. And adding more documentation didn’t move it at all in our system. A missing API is not a documentation gap, the fix lives in the component.

Mechanical dimensions tell us whether the agent understood the system. Judgment tells us whether it understood the design.

This is where agent instructions and skills shine, and they should focus on design language. Don’t give agents a list of answers. Give them the language and principles they can use to arrive at the right answer. A good example of this is Vercel’s DESIGN.md (vercel.com/design.md), which reads more like a brand guideline than other examples of this file found on the web. To build a file like this you should put on your scientist coat and run evals (lots of them).

If judgment is the difficult part, you can’t improve the guidance without evaluating the outcomes. A good starting point is reading Teresa Torres’s article on AI Evals: A Hands-On Guide for Product Teams. Vercel published an article on the methodology they used for building their DESIGN.md file.

Diving into this means scoring higher on the external dimension of the test. There are different strategies on how to expose data from your design system through open knowledge-bases, MCPs and custom internal tooling.

Want to go deeper into what judgment is? Check Pegah Ahmadi’s take on governance and AI-readiness:

“The question is whether an organization can introduce AI into design work without weakening the integrity of its decisions.

That means preserving human judgment, trust, responsibility, and accountability as production becomes faster and more automated. It means knowing not only whether an agent can do something, but whether it should do it and whether it has the authority to decide at all.

This is where AI readiness stops being a tooling problem and becomes an organizational design problem.”

Do AI-Ready, AI-Native, Agentic or similar prefixes hold in the future?

Probably not. Just like we stopped saying “Mobile-Friendly” because everything is just expected to work on mobile now. TJ revisited his view on Context-Based Design Systems and Use AI to Need Less AI.

Optimize for being a great design system. Then AI-readiness emerges as a property of that system. It’s a new way of exposing whether the quality bar you’ve already set is actually being met.

Building them with great engineering, design and governance practices will get you there. That’s what we do at Southleft: we keep moving forward and experimenting in the space, pushing the relationship between AI and the humans in the loop.

If you’re working through this on your team and want to compare notes, we’d love to talk. The conversation is more interesting than the discourse.

<cta>

Is your design system AI-ready?

A 30-minute discovery call is the fastest way to find out where your leverage is.

<theme.console>

// every theme is solved to WCAG AA before it's allowed out. color, type, shape, texture, elevation, motion — 44 tokens, the whole site, no rebuild.