What is the design Token tier system, and how does it differ from flat token naming?
Flat naming was the earliest, most intuitive approach: name a color or value directly and descriptively, like blue-500 for a specific shade of blue. This works fine when a Design System has few tokens, but as the system grows, a problem surfaces — if a button's background color uses blue-500, and later you want to change "all primary button backgrounds" to a different color, you discover blue-500 is also used in a dozen places that have nothing to do with buttons (link text, icon colors, borders), so changing one value drags along a wide swath of places that shouldn't have been affected.
The tier system splits naming into three independent layers to solve this. The first layer, "Global/Primitive," is only responsible for defining the raw value itself, like blue-500 = #2D5BD6, with no reference to usage at all. The second layer, "Semantic/Alias," gives that primitive value a purpose-oriented name — color-action-primary points to blue-500 — answering "what context is this color used in." The third layer, "Component-specific," has a particular component reference the semantic layer's token — button-primary-background points to color-action-primary. To change a color, you only need to adjust the mapping at the second layer, without touching the underlying primitive values or going through the component layer one by one.
Why did the tier system need to develop, and what underlying problem does it solve?
This architecture directly addresses the "Token chaos" problem that inevitably emerges once a Design System scales. When a system has only a few dozen tokens, flat naming works just fine, since people can still remember where every token is used. But once token counts grow into the hundreds or thousands, spanning multiple platforms (web, iOS, Android) and multiple themes (light, dark, brand variants), flat naming turns "can this token be safely changed" into a question that requires a full codebase search to answer, dramatically slowing down design and development iteration.
The core idea behind the tier system is splitting "what the raw value is" from "what purpose this value represents" into two concerns that can change independently. The primitive value can stay stable (a specific blue's hex code doesn't change often), but the semantic layer's mapping can shift as the brand evolves or dark mode gets toggled, while the component layer barely needs to change at all, since it just stably points at the semantic layer. This layering turns common large-scale changes — rebranding a color, switching dark mode, adjusting brand identity — into something that only requires adjusting that middle layer, rather than auditing the entire codebase one reference at a time.
How does the tier system actually work, and what's the reference relationship between layers?
The reference chain is one-directional: component-layer tokens reference semantic-layer tokens, semantic-layer tokens reference global-layer tokens, and global-layer tokens are themselves the endpoint, referencing nothing further. A concrete example: the global layer defines gray-900 = #1A1A1A (a pure color value); the semantic layer defines color-text-primary = {gray-900} (giving this color the purpose of "primary text color"); the component layer defines heading-text-color = {color-text-primary} (a heading component referencing this semantic purpose). If dark mode needs adjusting later, you only need to have the semantic layer point to a different global value under dark mode (say, color-text-primary pointing to gray-100 instead of gray-900 in dark mode) — the component layer doesn't need to change at all, since it references the stable semantic-layer name rather than a specific color value.
In practice, teams commonly use tooling (Token transformation tools like Style Dictionary) to write these three layers of definitions into structured files (usually JSON or YAML), then automatically generate the formats each platform can read (CSS variables, iOS .swift constants, Android XML resources). This means a designer adjusting the semantic layer's mapping in Figma can directly sync code changes across all platforms, without an engineer manually editing each one by hand.
What's the real-world impact for readers or teams currently building a Design System?
If your team is building design tokens from scratch, the tier system is worth adopting from day one, rather than waiting until the Token count has ballooned into something unmaintainable before refactoring — refactoring a flat naming system that's already directly referenced by a large number of components costs far more than designing it in layers from the start. A concrete way to get started: first list out the "semantic purposes" your team actually needs (primary action color, secondary text color, danger/warning color, etc.), then work backward to define the global primitive values that support those purposes, rather than defining a pile of color values first and figuring out how to categorize them afterward.
If your team already has a flat-named token system and is weighing whether to refactor into a tiered one, the practical cost-benefit question is: does the team expect, over the next six months to a year, to need changes that span a large number of components — a large-scale color rebrand, adding dark mode, supporting multiple brands? If none of that is on the horizon, staying as-is may be more cost-effective than the engineering cost of refactoring; if that kind of need is a near-certainty, the earlier you refactor, the lower the cost of every subsequent large-scale change.
Material Design 3 (Google's design system) uses a token architecture that is a real-world instance of this exact three-way split — Reference Tokens (global), System Tokens (semantic), and Component Tokens. The system-semantic layer defines purpose-oriented roles like primary and on-primary; switching dark mode only requires adjusting which reference values the system layer points to, and the thousands of components using those roles require zero individual modification.
The advantage of the tier system is that large-scale changes (rebranding colors, switching dark mode, supporting multiple brands) only require adjusting the middle semantic layer, dramatically cutting maintenance and tracking cost; the drawback is that initial setup requires extra upfront thinking and planning time (semantic-purpose categories need to be clearly defined first), and for a project with few tokens and a small team, the architecture's upfront complexity can exceed the scale of the problem it actually solves.