Status
Open source · v1.0
Period
September 2026
Stack
JavaScript · Playwright · axe · Claude Code
A rule that nothing checks does not work.
Three levels before any score: what the project is, what this screen must do, and what a real browser measures. Only then the rubric, and acceptance by an agent that did not build it.
1Problem
Interfaces built by an agent fail in three predictable ways. Without project context the result is the statistical mean of its training data — competent and anonymous. “Looks right” gets said without anything being measured. And performance is judged on the developer’s desktop at full speed, where almost everything is fast.
2Task
Give the agent a process, mechanical checks, a scoring rubric and a performance method, so that “done” means “measured” — and make the checks trustworthy enough that people keep reading their reports.
3Approach
Context, contract, checks
Project context lives in one file; each screen gets a contract of what it must do; mechanical checks run against both.
Checks in a real browser
Eight scripts cover layout, states, tap targets, visual regression, motion, performance, vocabulary and forbidden patterns. Playwright and axe live inside the skill, so nothing is installed into the project.
Scenarios, not addresses
A scenario drives the app into the state where defects actually live — logged in, with data, mid-flow — instead of stopping at the first screen.
Performance attributed to functions
The long-animation-frame API breaks every slow frame down by function, file and caller; input delay is split into waiting, handler and paint. All numbers are normalised by the throttle factor.
A gate at the end of the turn
A stop hook refuses to end a turn while a check is red — only in projects that opted in.
An independent verifier
A read-only acceptance agent scores the result, may not fix anything, and must list what it did not check.
4What it measures
A worked example from the performance skill: a login screen at 20× throttle. The median frame was 17.2 ms (58 fps), but 7.9% of frames took longer than 50 ms, the worst one 210 ms, with 3,567 ms of blocking time. The slowest interaction took 216 ms — 20 waiting, 38 in the handler and 159 painting. The fix was in painting, not in JavaScript, which is exactly what a desktop run would never show.
A scenario run on real data found 206 ms of forced layout and a 76 ms delay before the input handler. The same page checked by address alone found almost nothing.
At ×20 a frame must happen without the main thread. Not “fast JS per frame” — no JS per frame at all.skills/frontend-performance/SKILL.md
5Limitations
- A headless browser is not a device: no thermal throttling, a different GPU and network. The numbers are a lower bound on the trouble.
- 60 fps at 20× is reachable only where a frame does not need the main thread; elsewhere the honest target is a steady 30.
- Some failures cannot be automated: focus order, stacking order, a cache that survives switching users.
- The scripts cannot see an empty cell or a monotonous page — that is left to the rubric and the verifier.