Skip to main content

Learning Track

Filter content by your role to see the most relevant sections

Atomic vs Semantic Tokens

Raw value tokens versus purpose-driven roles, decided by provenance: constraint-derived answers (accessibility, perception, composition) live one layer, brand choices live another — and granularity is floored by what anyone could ever perceive.

Darian Rosebrook
Staff Design Technologist, Design Systems Architect
Expertise:Design Systems, Token Architecture, Accessibility

Why This Matters

"When do I make an atomic token, and when do I make a semantic one?" is the question that decides whether a token system compounds or collapses. Teams that answer it by feel accumulate both failure modes: vocabularies too raw to theme (everything references blue.600, dark mode becomes a rewrite) and vocabularies too baroque to choose from (a takeout menu of sixty near-identical roles, where every screen is a fresh negotiation).

The working answer is a decision procedure, not a vibe: every token answers a question ("what is our decision for X?"), and the provenance of the answer decides the layer. Some answers are determined by constraints—accessibility floors, perceptual thresholds, composition rules—and those are derived, documented, and wrong-able. Others are brand choices—defensible, themeable, and meaningless to argue about on the merits. Confusing the two is how systems get both brittle and bloated: derived values locked in as if they were taste, and taste defended as if it were physics.

This page turns that procedure into checks you can run on any token, using this repository's scales as the worked evidence.

Core Concepts: Provenance, Granularity, Availability

Provenance: Derived or Chosen?

Interrogate any token with one question: could a sufficiently careful analysis have computed this value from requirements? If yes, it is derived and belongs with the machinery that derives it. If no—if the value is one defensible pick among many—it is chosen and belongs where choices are centralized and swapped.

The derived set in this repository is larger than most teams expect, and that is the system's quiet strength:

  • Accessibility floors: tapTargetMin 44px, actionMinHeight 36px, the 4.5:1/3:1/7:1 thresholds, the 320px reflow point, reduced-motion handling. These are not opinions; a violation is a defect.
  • Perceptual derivations: the color ramps (contrast-keyed at 1.15:1, luminance-uniform, shared per-level targets), the frame-quantized durations, the one-frame stagger floor. Wrong values are perceptually indistinguishable noise.
  • Composition rules: control heights from spacing steps, icon sizes aligned to the same scale, nesting radius step-downs, depth numbers that order with shadow levels.

And the chosen set is equally specific: the accent family per brand, the radius personality, the 1.5px stroke weight, the easing characters, the breakpoint stops. None of these can be argued from requirements—but all of them are localized: one field in a brand file, one semantic mapping, one slot in the generator's anchors. The layer boundary exists precisely so that changing a chosen value never requires touching a derived one.

The Test That Separates Them

// Ask of any token:
// 1. What question does it answer? ("What is our decision for ___?")
// 2. Who could re-derive the value, and from what?
//    - From WCAG / perception / composition  → derived (core or a rule)
//    - From brand identity / a decision      → chosen (brand/theme data)
// 3. What breaks if the value changes?
//    - A validator, an audit, a physical law → it was derived
//    - Nothing; it just looks different      → it was chosen

// Worked: feedback.border.warning = orange.700 (light) / orange.300 (dark)
//   The hue family is chosen (warning ≈ orange, by convention).
//   The 700/300 split is derived (contrast parity per mode).
//   Both facts live in one token — that is normal, and the layer
//   boundary is at the reference, not inside the value.

Granularity: The Perceptual Floor

How fine should a scale be? Fine enough that adjacent steps are distinguishable—no finer. A just-noticeable difference for quantities like duration and size is roughly 15%: adjacent scale steps closer than that are two names for one answer, and two names for one answer is where drift breeds (teams pick by name aesthetics; reviewers cannot arbitrate because nothing looks different).

Audit this system's scales against that floor and the verdicts are instructive. Spacing passes cleanly—every adjacent step is 33–100% apart. Icons (20–33%) pass. The composite durations consumers actually see (83/100/167/250/333ms) pass at 20–67%. The raw duration scale does not: 150/167ms, 300/333ms, 600/667ms sit at ~11%, and medium/medium1 are both 250ms—name-only distinctions, kept tolerable only because the consumption tier never offers them.

Availability: The Tier That Choosers See

The takeout-menu problem is solved at the consumption layer, not the storage layer. Raw scales can afford breadth (sixteen durations, nine color families) precisely because choosers are offered the semantic tier: a handful of roles and composites, each with a sentence-test name and a decided meaning. The rule for sizing that tier:

  • Enough decided answers that new work is mostly looking up what the system already decided—the arbitrary decisions become outliers.
  • Few enough that choosing is not itself a project—if two roles theme the same, fail the same, and compose the same, they are one role.

In practice this system's working menu is small: seven motion composites, a dozen color role groups, three border roles, six radius roles. Everything else is either storage (core) or wrapping (component tokens)—and the layering makes each tier's size a separate, deliberate decision.

Atomic Is a Storage Word, Semantic Is a Job Word

The naming asymmetry is the tell: atomic tokens are named for what they are (a family, a step, a frame count); semantic tokens are named for what they do (the hover background of the primary action). When you catch a "semantic" token whose name is secretly atomic—blueForButtons—the layer boundary has leaked, and the refactor is usually to promote the intent into a proper role rather than to rename the color.

System Roles: Who Guards the Boundary

Design Impact

Designers own the question list: which decisions the system has decided, phrased as roles. A designer's comp that uses a color no role answers for is either a new role (does it theme or fail differently?) or off-system (absorb it). The provenance questions keep the role list honest.

Engineering Impact

Engineers own the reference directions and the tiering: semantic references core, components reference semantic, and nothing offers raw scale values to product code. The reference linter is the mechanical guard; the code review question "which composite/role is this?" is the human one.

Accessibility Impact

Accessibility owns the derived floor's authority: when a derived value conflicts with a chosen one, derived wins where users are affected (contrast over brand hex, targets over sleekness) and the choice re-routes (another family, another step) rather than the constraint bending.

Design & Code Interplay

The design tool mirrors the tiers as variable structure: collections hold the atomic scales (storage), modes and styles hold the semantic roles (the menu designers actually pick from). A designer choosing from raw palettes inside a comp is the design-side equivalent of a component consuming --core-*—technically possible, structurally wrong, and the review catch is identical: "which role is this?"

Applied Example: Auditing a Token With the Procedure

Someone proposes spacing.size.18—an 18px step "between 16 and 24." Run the full procedure:

  1. Question: "what is our decision for medium-plus gaps?"—vague already. Good sign the proposal is symptom, not need.
  2. Provenance: is 18px derivable? No constraint yields it—not a target multiple, not a perceptual anchor. It is chosen; and chosen values that are just other steps belong to density, which already exists as an axis with slots for exactly this.
  3. Granularity: 16→18 is +12.5%—below the perceptual floor. Two names, one answer: the classic drift seed.
  4. Availability: the consumption tier offers semantic spacing and density modes; an 18px literal widens nothing that a density switch doesn't already decide.
  5. Verdict: reject the step; route the need —if the real complaint is "default feels cramped in tables," that is a density or role conversation with a name that will survive the next refactor.

Five checks, two minutes, and the vocabulary stayed tight while the actual need got named. That is the procedure earning its keep.

Constraints & Trade-offs

  • Derived-heavy vs chosen-heavy systems: deriving everything you can (contrast-keyed ramps, frame quantization) front-loads rigor and pays off in audits; it also makes "just change the value" harder, which is the point and occasionally the friction.
  • Raw breadth vs consumption tightness: broad storage with a narrow menu gives coverage and guardrails, but only if the tiering is enforced—breadth that leaks to choosers is a menu with no floor.
  • Perceptual floors vs engineering quantization: frame-clean values and perceptually-distinguishable steps collide at ~11%; perception must win at the vocabulary layer because nobody sees a frame boundary.
  • Splitting roles vs merging them: every split multiplies theme and audit surface; the merge rule (same theming, same failure, same composition ⇒ one role) is deliberately conservative.

Common Pitfalls & Failure Modes

1. Semantic tokens wearing atomic names

// Bad: A job word costume over a value
--semantic-color-blue-for-links: #0a65fe;
// Good: The role owns the job; the value resolves upward
--semantic-color-foreground-link: var(--core-color-palette-brand-primary-600);

2. Derived values defended as taste

"We chose 3:1 for borders" is not a choice; it is WCAG 1.4.11. Treating derived floors as negotiable style produces audit findings at the worst possible time. The fix is moving the number next to its citation.

3. Chosen values hardened as physics

The inverse: defending a brand hex or a 1.5px stroke as if it were accessibility. The cost is rebrands that touch derived layers—and the fix is routing every choice through theme data so changing it is a mapping edit.

4. Sub-perceptual vocabulary

// Bad: +12.5% apart — two names, one answer
size.06: 16px   size.065: 18px   size.07: 24px

5. Menu leakage

Product code consuming --core-* directly re-opens every decision the semantic tier closed. The reference linter exists for exactly this; a codebase that disables it has chosen drift.

Verification Checklist

Additional Resources

  • Core vs Semantic deep-dive — the layer mechanics worked end to end (/blueprints/foundations/tokens/core-vs-semantic)
  • Token Naming & Hierarchy — how the layers stay legible (/blueprints/foundations/meta/token-naming)
  • Color Foundations — the contrast-keyed ramp derivation in full (/blueprints/foundations/color)
  • Motion & Duration — the JND audit and tiering that keeps the duration menu clean (/blueprints/foundations/motion)

Related Concepts

Reflection Questions