‹  Research

Case study · Frontend standard

frontend-quality

AI-built interfaces get called finished without being measured. This plugin turns “looks right” into numbers: a real browser, a fixed rubric, a frame budget at 20× CPU throttle, and an acceptance agent that never saw the work being done.

Status

Open source · v1.0

Period

September 2026

Stack

JavaScript · Playwright · axe · Claude Code

A rule that nothing checks does not work.

Three levels before any score: what the project is, what this screen must do, and what a real browser measures. Only then the rubric, and acceptance by an agent that did not build it.

1Problem

Interfaces built by an agent fail in three predictable ways. Without project context the result is the statistical mean of its training data — competent and anonymous. “Looks right” gets said without anything being measured. And performance is judged on the developer’s desktop at full speed, where almost everything is fast.

2Task

Give the agent a process, mechanical checks, a scoring rubric and a performance method, so that “done” means “measured” — and make the checks trustworthy enough that people keep reading their reports.

3Approach

  1. Context, contract, checks

    Project context lives in one file; each screen gets a contract of what it must do; mechanical checks run against both.

  2. Checks in a real browser

    Eight scripts cover layout, states, tap targets, visual regression, motion, performance, vocabulary and forbidden patterns. Playwright and axe live inside the skill, so nothing is installed into the project.

  3. Scenarios, not addresses

    A scenario drives the app into the state where defects actually live — logged in, with data, mid-flow — instead of stopping at the first screen.

  4. Performance attributed to functions

    The long-animation-frame API breaks every slow frame down by function, file and caller; input delay is split into waiting, handler and paint. All numbers are normalised by the throttle factor.

  5. A gate at the end of the turn

    A stop hook refuses to end a turn while a check is red — only in projects that opted in.

  6. An independent verifier

    A read-only acceptance agent scores the result, may not fix anything, and must list what it did not check.

4What it measures

50 pointsFive dimensions of ten. Ship at 42 or more, targeted fixes between 30 and 41, rebuild below 30.
14 failure modesCollected from 190 fix commits across two real frontends — each with the check that now catches it.
0.83 msWhat is left of a 16.67 ms frame at 20× CPU throttle. At that speed a frame has to happen without the main thread at all.
8 checksLayout, states, tap targets, regression, motion, performance, vocabulary and forbidden patterns.

A worked example from the performance skill: a login screen at 20× throttle. The median frame was 17.2 ms (58 fps), but 7.9% of frames took longer than 50 ms, the worst one 210 ms, with 3,567 ms of blocking time. The slowest interaction took 216 ms — 20 waiting, 38 in the handler and 159 painting. The fix was in painting, not in JavaScript, which is exactly what a desktop run would never show.

A scenario run on real data found 206 ms of forced layout and a 76 ms delay before the input handler. The same page checked by address alone found almost nothing.

At ×20 a frame must happen without the main thread. Not “fast JS per frame” — no JS per frame at all.skills/frontend-performance/SKILL.md

5Limitations

  • A headless browser is not a device: no thermal throttling, a different GPU and network. The numbers are a lower bound on the trouble.
  • 60 fps at 20× is reachable only where a frame does not need the main thread; elsewhere the honest target is a steady 30.
  • Some failures cannot be automated: focus order, stacking order, a cache that survives switching users.
  • The scripts cannot see an empty cell or a monotonous page — that is left to the rubric and the verifier.

Sources