Skip to main content

Learning Track

Filter content by your role to see the most relevant sections

Accessibility Tooling

Test, validate, and enforce accessibility at every stage: jsx-a11y errors at lint speed, axe-core auditing rendered components in vitest, token-level pair validation at build, keyboard-driven e2e walks — plus the manual checks automation cannot replace.

Darian Rosebrook
Staff Design Technologist, Design Systems Architect
Expertise:Design Systems, Accessibility, Testing

Why This Matters

Accessibility has more good tooling than almost any other quality property—and more ways to use it badly. The failure mode is always the same: running one tool, late, as a gate, and calling the absence of its findings "accessible." The tools are excellent at their slices and silent about everything else; using them well means knowing both the slice and the silence.

This system runs four automated layers, each positioned where its cost is lowest: static markup rules at lint speed, rendered-tree audits inside component tests, numeric claims at build time, and real keyboard operation in the e2e suite. This page covers what each layer sees, what none of them see, and how the manual pass completes the picture.

Core Concepts: Four Layers and Their Blind Spots

Layer 1: jsx-a11y at Lint Speed

Static rules over markup, in seconds, on every commit. This repo registers the plugin through next/core-web-vitals and promotes key rules to hard errors:

// eslint.config.mjs (excerpt)
'jsx-a11y/alt-text': 'error',
'jsx-a11y/anchor-has-content': 'error',
'jsx-a11y/anchor-is-valid': 'error',

Sees: missing alt, empty anchors, invalid hrefs, obviously wrong ARIA attributes—the mechanical sins. Blind to: everything runtime—rendered contrast, focus behavior, announcement quality. A file can lint clean and be unusable.

Layer 2: axe-core in the Component Tests

The rendered tree, audited where components are already tested. The vitest setup wires jest-axe with a deliberate tag selection:

// test/setup.ts (behavior)
expect.extend(toHaveNoViolations);
// axe configured with tags: ['wcag2a', 'wcag2aa', 'best-practice']
// (color-contrast promoted to error only at the AAA level,
//  because the pair ledger already enforces AA at the token layer)

Sees: name/role/value problems in real DOM, landmark structure, contrast of rendered combinations. Blind to: keyboard interaction sequences (it inspects state, not journeys), screen reader behavior beyond the tree, and anything behind interactions the test does not drive.

Layer 3: The Token Validator at Build

The numeric claims—contrast pairs, per mode—checked at the source rather than the screen: WCAG_LEVELS + validateColorPair walk the declared pairs. This is the cheapest layer of all (the data is tiny) and the only one that catches a failing pair before any component renders it.

Sees: declared numeric claims. Blind to: usage—every combination that was never declared (which is why the pair ledger discipline from the tokens page is this layer's precondition, not its product).

Layer 4: Keyboard-Driven E2E

The Playwright suite operates the product the way a keyboard user does: tabbing through navigation, clicking prerequisite links, asserting landmarks and focus visibility, walking the mobile viewport. It is the only layer that catches focus-trap failures, tab-order absurdities, and reflow breakage in context.

Sees: real interaction behavior, real viewports. Blind to: what it is not written to walk—coverage is per-pattern and grows only as patterns gain walks.

The Honest Boundary

All four layers share one silence: none of them experiences the product. Comprehension, announcement quality, task flow, and real screen reader behavior under real browsing habits remain human checks with tools—VoiceOver or NVDA, run per pattern on a schedule (the assistive-tech page's protocol). The maturity marker is not more automation; it is knowing exactly where the automation stops and the pass begins.

System Roles: Who Runs What

Engineering Impact

Engineers own the layers' wiring and health: rules promoted to errors stay errors, axe tags stay deliberate, e2e walks grow with patterns. A disabled rule is a policy decision, made in config with a comment—not in code, silently.

Accessibility Impact

The a11y function owns the boundary map: what each layer covers, where the manual schedule picks up, and the periodic audit that the layers still say what they claim to.

Design Impact

Design-side checkers (contrast plugins) calibrate to the same numbers as Layer 3—one answer across the design/code line, so findings are routing decisions rather than disputes.

Design & Code Interplay

The design side sees the tooling as shared instruments: the same WCAG numbers in the plugin and the validator, the same floors in the size scales and the e2e asserts. When both sides trust one number, a finding is never an argument about the number—it is a routing question about the layer that missed it.

Applied Example: Routing a Finding to Its Layer

A user reports: "the filter panel is unusable with a screen reader." Route it:

  1. Reproduce with the tree: VoiceOver walk transcribes the announcements—headings missing, controls announced as "group, group".
  2. Ask which layer should have caught each line: unnamed buttons → Layer 2 (axe names rule) or Layer 1 (jsx-a11y); missing landmarks → Layer 2; tab trap in the panel → a Layer 4 walk that does not exist yet.
  3. Fix at the component, encode at the layer: the filter gains names and landmarks; the axe assertion joins its test; the e2e gains the panel walk—so this specific finding becomes unrepeatable, not just fixed.
  4. Log the gap honestly: the findings that no layer could catch (announcement phrasing) stay on the manual schedule—tracked, not pretended away.

Constraints & Trade-offs

  • Layer breadth vs depth: four cheap layers beat one deep one—the failure classes differ, and the cheap layers run constantly.
  • Error vs warning promotion: every rule promoted to error removes discretion and adds friction; demoted rules reacquire silence. The ledger of placements is policy worth reviewing.
  • Automation trust vs manual reality: green boards breed the belief that accessibility is solved; the boundary map exists to keep the conversation honest inside the confidence.

Common Pitfalls & Failure Modes

1. The single-tool audit

Running one scanner and shipping its absence of findings. Each layer's blind spot list is the argument for the other three.

2. Silent rule demotion

// Bad: Per-file suppression, forever
/* eslint-disable jsx-a11y/... */
// Good: Config-level policy with a comment and an expiry review

3. Testing the happy render only

axe on the default state misses the opened dialog, the errored form, the loading skeleton—audit the states users actually suffer in.

4. Scores as goals

Lighthouse scores are directional, gameable, and viewport-specific. Targets belong to criteria and patterns; scores are diagnostics.

Verification Checklist

Additional Resources

  • Assistive Technology Support — the manual protocol this layering serves (/blueprints/foundations/accessibility/assistive-tech)
  • Automation & CI/CD — where each layer runs (/blueprints/foundations/tooling/automation)
  • axe-core — the rule engine and its documentation of what it cannot check (https://github.com/dequelabs/axe-core)
  • The sources — test/setup.ts, eslint.config.mjs, utils/accessibility/tokenValidator.ts, test/e2e/foundationNavigation.spec.ts

Related Concepts

Reflection Questions