Tremoloo - logo
English
English
A group of people sitting at a table with computers

How Do You Test an Interface That's Never the Same Twice?

Generative UI

testing generative UI, usability testing AI interfaces, QA adaptive UI, evaluating generated interfaces, UX research AI products, generative UI quality

Summary

Generative UI breaks the core assumption of usability testing: that there is an interface. When the AI composes a different arrangement per user and context, five moderated sessions on "the checkout flow" no longer sample the same thing.

Testing doesn't disappear β€” it moves to three new objects.

  1. First, components: every component must be independently robust, because it will appear in combinations nobody planned.

  2. Second, the composition rules: you test the system's decisions β€” did it choose the right components, in a sensible arrangement, for this scenario β€” across dozens of scenarios, not five.

  3. Third, distributions: single-session depth gives way to statistical breadth over many generated variants, because quality is now a property of a range of outputs, not a screen.

The research question itself changes: not "is this interface usable?" but "does this system make good interface decisions, and how often does it fail?"

A client asked us recently how we'll usability-test their AI feature when "there's no screen to put in front of users." It's the right question, and it deserves a real answer, because a lot of teams are quietly shipping generated interfaces with no evaluation strategy at all β€” the testing equivalent of hoping.

Why does traditional testing break?

The standard playbook β€” recruit five users, give them identical tasks on identical screens, watch where they struggle β€” works because the artifact is fixed. Findings replicate. "Users miss the button in the top right" is actionable because the button is in the top right for everyone.

Generative UI removes the fixed artifact. User one's dashboard isn't user two's. A finding from Tuesday's session may describe an arrangement that never occurs again. Test the thing the old way and you'll collect anecdotes about a product that no longer exists by the time you write the report.

What are you actually testing now?

1. The components. Since components appear in unplanned contexts, each one must survive alone: does the date picker work regardless of what surrounds it? This resembles design-system QA more than usability testing β€” states, edge cases, and behavior under odd data. Weak components are no longer a local problem; they're a systemic one.

2. The composition decisions. The model is now a junior designer making layout calls at scale. So evaluate it like one: give the system a battery of scenarios β€” the sparse case, the overloaded case, the edge-case data, the angry user mid-task β€” and score whether its choices match what a competent designer would do. You're testing judgment, not pixels. Expect to build a rubric: relevance of selected components, sensible hierarchy, no contradictory elements, states handled.

3. The distribution. Quality becomes statistical. One great generated dashboard proves nothing; you need the failure rate across a hundred generations. What percentage of outputs are unusable, off-brand, broken in Arabic, missing a state? This looks less like a lab study and more like model evaluation β€” because that's what it is.

Does moderated research still matter?

More than ever, but aimed differently. Watching a real user interact with a generated interface tells you what generation quality metrics can't: whether the arrangement made sense to a human, whether the adaptation helped or confused, whether users even noticed the interface adapting β€” and how they reacted when they did. "Wait, this wasn't here before" is a trust event you will never catch in automated evaluation.

The practical program: automated scenario batteries run continuously, moderated sessions run per release on representative scenarios, and a standing log of generation failures fed back into the component library and rules. The teams that treat generated UI as untestable aren't moving fast β€” they're accumulating invisible quality debt.

Who should be doing this work?

Here's the honest shift: evaluating generative UI sits at the intersection of UX research and ML evaluation, and neither discipline owns it yet. Researchers have the human judgment methods; ML teams have the statistical tooling. The evaluation program needs both β€” which is why this is a research-ops decision, not something to hand to whichever team has spare capacity.


We run usability programs for products where the interface itself is becoming dynamic. If you're shipping AI features without an evaluation strategy, let's build one β†’

Ready to Launch
Your Next Project?

If you’re ready to stop iterating in circles, we partner with focused teams to research,
design, iterate that are clear in purpose and ready to perform.