
Button 2026 recap: Why AI content standards break, and what a few content design experts did about it
From drifting style guides to untested prompts, Button 2026 speakers showed where AI content systems break and how they fixed them.
The Ditto team had the opportunity to join Button Conference this year as a sponsor, and a common thread across the conference talks was how fragile AI systems can be. So content designers are getting more creative in how they enforce consistency and quality in rapidly changing product workflows.
Three specific ways these homegrown systems have been breaking down in practice: the rules live in too many places, a machine reads them differently than a person does, and they aren't tested against real cases. Here's what each looked like in practice, and what speakers did about it.
The rule lives in too many places
A single content standard does many jobs. It guides the draft, gates what ships, and grades the result. Each job tends to live in a different place, and the copies of the rule drift.
Emma Pindera, Senior Content Strategist at PointClickCare, described what it takes to keep one style guide change from drifting. When there’s a change, she runs a four-step process to keep it consistent across the AI skill, the design system, and the style guide.
- Refresh the source. Use scheduled tools to update the skill, the design system, or the guide itself, and flag gaps.
- Review the skill package. Claude Cowork updates it automatically, and a person modifies or approves the changes.
- Apply the same change summary to the design system and Figma plugin through MCP.
- Verify. Spot-check the update, then run a sample review to confirm the change shows up.
Skip any step, and the pieces drift out of sync again.
Avalara’s Head of Documentation, Kristina Maultsby, showed what drift looks like once it ships. One of her team's rules required preserving UI labels verbatim. Another converted all labels to sentence case. Both shipped. Her takeaway was that guardrails need to be tested against each other, not just on their own.
Mike Jang, Principal Technical Writer, F5, took a single-source approach. He turned each page of his 30-page style guide into its own skill, written as Markdown files in a git repository, plus an instructions file telling the agent how to use them.
The result is a UX message writing buddy that offers three options and lets a person pick. Jang tests the agent for a range of use cases, like tooltips and empty states, with screenshots of messages that break the guidelines. Then he automates the setup with CI/CD, the kind of pipeline engineers use to test and ship code, so the guide stays the agent's single source of truth.
Writing the standard is just the start. The harder part is keeping every place it lives consistent, which starts with one source and one owner.
A machine reads the rule differently than you do
Himali Kelvekar (Microsoft) shared an exercise that showed the gap between an AI judge and a human evaluator. Teams often use one AI model (an LLM judge) to score another's answers, guided by a written instruction. Picture an AI travel assistant, and a user who lands at a London airport at 8 pm with their family and asks for the "best way" to reach a hotel in central London.
The instruction tells the judge to check that the answer mentions the mode of transport, distance, travel time, and price.
A person reads that scenario and infers a lot. "Best" involves trade-offs, like fastest, cheapest, or simplest, and an 8 pm arrival with a family suggests tired kids and plenty of luggage.
An LLM judge applies the instruction literally. A response can include all four details and still not be useful if it doesn’t offer options or explain the trade-offs.
That matters because manual review isn't possible across thousands of outputs. The judge's reading becomes the standard, and a misreading repeats at scale.
Kristina Maultsby experienced a version of this, too. Her team at Avalara uses AI to draft documentation for its knowledge center. Every page is a defined type with an expected structure, like a short description and keywords. Those fields are how pages are found, reused, and connected to related pages.
The team's prompt asked for the description and keywords "softly," with no step to check, and the how-to pages came out missing them. A person would likely read a politely worded request as part of the job, but the model skipped it. Her pipeline now checks every generated page against the structure for its type, and anything missing triggers a repair.
In both cases, the instruction was written the way you'd write it for a person. Instructions for machines are copy too, and they need the same specificity, editing, and testing as any customer-facing words.
Nobody tested the rule against real cases
Amit Singh's team at Intuit built an AI assistant that answers customers' payroll and tax questions. Every reply followed the same template, acknowledging the person, answering directly, explaining why, and suggesting a next step.
It worked, until real conversations showed it was too rigid. People ask quick follow-ups, and they don't expect a full introduction and explanation every time. His team moved to several response shapes, including a short one for follow-up turns.
The takeaway was that a structure that looks right on paper still has to be tested against real conversations. Singh's rule is "You cannot ship a stated behavior that is not tested." A stated behavior is any claim about what the AI will do, like "it never invents a tax rate."
If you've claimed it, test it before it ships. He does that through a five-step loop:
- State a behavior. Write it into the skill's short description so there's something specific to test.
- Add golden use cases. These are real cases with the answers attached, and every claim is proven with tests.
- Judge at scale. An LLM scores a rubric across hundreds of thousands of conversations.
- Calibrate with humans. A machine creating a machine is only as honest as its rubric and its interpretation of that rubric.
- Ship or fix. Restate the claim, then go around again.
Kelvekar showed what a test set looks like for an AI travel assistant. It starts with the real questions people are likely to ask, like how to get to a hotel after a late landing or whether a flight is still on time, and marks the ones that matter most as high priority.
Then the set deliberately adds the cases that go wrong. One is a known gap, a question the assistant is known to struggle with, like what a traveler is owed after a cancelled flight. Another is an unsupported request, like asking it to book a hotel and charge a card, so the test shows how the assistant responds instead of leaving that to chance.
Kelvekar calls this set the backbone of everything that follows, and says coverage matters more than volume. In other words, start with a small test set that spans the situations people really face, then scale up with automated evaluations.
Across these talks, the pattern looks like a continuous loop. Collect what broke, write the rule, test it against real cases, enforce it, and refresh it as the product changes.
Four steps to make your AI content standards harder to break
Pulling these together, here's a simple place to start.
- List every place your standards live. Style guides, prompts or skills, design tool plugins, and eval rubrics.
- Pick one source and one owner. Decide where a change gets made first and who makes sure every other location follows.
- Reread your rules the way a machine would. Look for anything you left unstated or that could be misinterpreted.
- Build a small test set from real failures, and rerun it whenever a rule changes.
We've designed Ditto to centralize your content system, so the same standards apply everywhere your team works. That includes your design tools, your codebase, and your AI agents.
Thanks to Button
We had a great time at Button this year! Thank you to the Button team and to every speaker. There were many more insightful speakers and topics that we haven’t covered here. You can buy access to the 2026 recordings to watch on demand.
Latest


