Design · October 5, 2026
Design System Observability & Platform Ops: Why Dashboards Lie and Observability Tells the Truth
By Morgan Ross · Senior Technical Lead
Photo by Karsten Winegeart on Unsplash
The Hidden Cost of Design System Governance
Most design system teams already know they should version their components, communicate deprecation in advance, and measure system health. They’ve read the blog posts, attended the talks, and nodded along at conferences. Per Trueman 2026, most teams still aren’t doing these things with real rigour.
Most of them still aren’t doing these things with any real rigour, and I don’t think that’s a knowledge problem. It’s an infrastructure problem. The tooling, processes, and organisational habits that make these practices work in other disciplines haven’t been built for design systems yet, and most teams are trying to improvise their way through operational challenges that platform engineering, API management, and DevOps observability solved years ago.
The Monitoring-Observability Gap
Most design system teams can tell you component adoption numbers, if they’ve set up analytics at all. Some track which components are used and how often. Very few can tell you how components are being used, where props are being overridden, which tokens are getting bypassed in favour of raw values, or which teams have built parallel solutions because the system’s components didn’t fit their needs. Per Daley 2026, this gap is exactly what observability through OTel pipelines is designed to close.
That last point matters most. In DevOps, observability’s value comes from letting teams discover unknown problems. If you only look at what you already expect to find, you’ll catch known issues and miss everything else. Design system teams operate almost entirely in this mode. You check adoption dashboards and see that Button is used 4,000 times. What you don’t see is that 600 of those instances have overridden the padding token, 200 have wrapped it in a custom div to add behaviour the component doesn’t support, and three teams have built their own button entirely because yours doesn’t handle their loading state correctly.
Spotify’s Encore is one of the more sophisticated examples of component-level tracking in the space. They query repositories daily to understand where and how often components are used, and their slot pattern analytics can show which props are overridden most frequently and what configurations teams choose. That starts to get into observability territory, asking “how are people using this?” rather than just “is this being used?” But most teams aren’t Spotify, and even Encore’s approach is narrower than what DevOps observability provides. They’re still primarily answering known questions about known components. The DevOps model lets teams discover problems they weren’t looking for.
The Three Pillars of Design System Observability
The conventional framing of observability in DevOps rests on three pillars. Metrics track aggregated measurements over time, things like component adoption rates, token usage patterns, and detachment rates. Logs capture event-level detail — each component instantiation, each prop override, each token bypass. And traces follow a component’s journey across the codebase, from Figma Dev Mode through development into production, so you can see where things slow down or break.
Together, they let a design system team diagnose problems they’ve never encountered before without having to add new instrumentation first.
Most design system teams have almost none of this. The tooling, processes, and organisational habits that make observability work in other disciplines haven’t been built for design systems yet.
Agent Traces as Free Telemetry
The fact that data is being auto-deleted (30-day timer) creates urgency without being alarmist. Four agent-trace signals: hallucinated props, bypassed tokens, novel components, corrective turns. Every hallucinated prop is a documentation gap, not a prompt problem.
The Seven Moves for 2026
Move 1: Instrument CI/CD with Automated Health Checks. Per Arant 2026, automated checks replace the need for manual governance.
Every pull request can trigger automated checks that verify token usage, component implementation, and accessibility compliance. These checks fail the build when violations occur, providing immediate feedback rather than discovering problems weeks later. This isn’t gatekeeping. It’s the same pattern your codebase uses for linting, testing, and security scanning.
Move 2: Adopt SLOs for System Health
Define what “healthy” means in measurable, specific terms. For a design system, that might be “fewer than 5% of component instances in production use overridden tokens” or “new components reach 60% adoption within their target product teams within 90 days.” Start with three or four SLIs (Service Level Indicators) you currently can’t measure, and figure out how to instrument for them.
Move 3: Implement Golden Paths for Common Workflows
A golden path is an opinionated, well-maintained route through a common workflow that covers roughly 80% of use cases. Instead of publishing components and hoping teams assemble them correctly, provide assembled starting points and make the places where teams need to diverge visible and trackable.
The Infrastructure Mindset
Treating design systems as infrastructure changes how you approach governance entirely. You think in terms of health and sustainability rather than rules and reviews. Feedback loops replace enforcement. Production systems have SLAs, uptime monitoring, and automated alerts. Design systems can have component integrity scores, token consistency tracking, and automated checks that run on every pull request.
From Compliance to Health
The most effective design systems aren’t the ones with the most comprehensive documentation or the strictest governance. They’re the ones that continuously adapt based on real usage patterns. They’re the ones that measure health rather than enforce compliance.
This requires changing how you think about your role. If you’re a design systems lead, you’re not building a comprehensive library that teams should use. You’re providing infrastructure that teams rely on. That infrastructure needs monitoring, maintenance, and continuous improvement based on production data.
Practical Code Example: CI Health Check
// Example: CI check that validates token usage and component adherence
const { scanComponentUsage } = require('react-scanner');
const { initOTelSDK, trace } = require('@opentelemetry/sdk-node');
async function ciHealthCheck() {
const sdk = new NodeSDK({
spanProcessors: [
new BatchSpanProcessor(
new OTLPTraceExporter({
url: process.env.OTEL_ENDPOINT,
}),
),
],
serviceName: 'design-system-health',
});
sdk.start();
try {
const usage = await scanComponentUsage('./src');
// Emit spans with key attributes
for (const [componentName, data] of Object.entries(usage)) {
const span = tracer.startSpan('component instance');
span.setAttribute('component.name', componentName);
span.setAttribute('component.isDeprecated', data.isDeprecated);
span.setAttribute('component.props', JSON.stringify(data.props));
span.end();
}
// Query recent trends
const trends = await queryObservabilityTrends();
console.log('Design system health trends:', trends);
} finally {
sdk.shutdown();
}
}
Summary Table: Monitoring vs Observability
| Aspect | Monitoring (Adoption Dashboards) | Observability (Platform Ops) |
|---|---|---|
| Focus | What is used? | What breaks? how? why? |
| Data type | Lagging indicators (adoption rates) | Leading indicators (drift, overrides, gaps) |
| Feedback speed | Quarterly audits | Continuous (CI/CD, traces) |
| Feedback nature | Enforcement (compliance) | Improvement (data-driven) |
| Scale | Manual, periodic | Automated, continuous |
| Power dynamics | Teams policed | Teams served (infrastructure mindset) |
Conclusion
Design systems need observability. Not because you need more ways to catch violations, but because you need better ways to maintain infrastructure that teams depend on. The question isn’t whether teams are following your guidelines. The question is whether your system is healthy enough to serve their needs. Monitoring is how you answer that question continuously, rather than discovering the truth too late to prevent entropy from taking hold.
References
- “We know how to build design systems, but we don’t know how to operate them,” Murphy Trueman (May 2026) — The canonical piece on DS vs ops gap. Covers observability (three pillars), SLOs, deprecation lifecycle, golden paths.
- “Stop policing your design system,” Murphy Trueman (Jan 2026) — Health metrics vs compliance metrics. Component integrity scores, token consistency tracking, CI/CD integration. The “measure health, not compliance” framing.
- “Using Observability Tools to Track Frontend Migrations,” Adrianne Daley (Jan 2026) — Concrete implementation: react-scanner → OTel → Honeycomb. Code configs, gotchas (maxQueueSize), dashboards.
- “How to govern a design system across 10+ consumer projects,” Adam Arant (Apr 2026) — Seven-step practical framework. Contribution models, frozen API, SemVer, codemods, manifest-per-consumer, deprecation timelines, fixed cadence.
- “Agents are the most honest reviewers your design system will ever have,” Murphy Trueman — Agent traces as free telemetry. Four signal types (hallucinated props, bypassed tokens, novel components, corrective turns).
- “Design System Tooling 2026,” DesignSystems.one — The Figma → MCP → Agent → Repo pipeline. Token tooling landscape. Wire-it-up commands for Claude Code, Cursor, Codex.
- “We know how to build design systems, but we don’t know how to operate them,” Design Systems Collective (May 2026) — The infrastructure problem. Tooling, processes, and organisational habits.
- “Stop policing your design system,” Murphy Trueman (Jan 2026) — Health metrics vs compliance metrics.
- “Using Observability Tools to Track Frontend Migrations,” Adrianne Daley (Jan 2026) — Static code analysis → OTel → Honeycomb for migration tracking.
- “How to govern a design system across 10+ consumer projects,” Adam Arant (Apr 2026) — Seven-step governance model. Contribution models, frozen API, SemVer, codemods, manifest-per-consumer.
Primary sources:
- https://www.designsystemscollective.com/we-know-how-to-build-design-systems-but-we-dont-know-how-to-operate-them-0f06e79de7fb
- https://blog.murphytrueman.com/stop-policing-your-design-system/
- https://adriannedaley.com/blog/2026/01/23/observability-for-migrations.html
- https://adamarant.com/en/blog/how-to-govern-a-design-system-across-10-consumer-projects
- https://www.designsystems.one/designops/tooling
- https://blog.murphytrueman.com/agents-are-the-most-honest-reviewers-your-design-system-will-ever-have/
Primary Sources
The following primary sources underpin this analysis. Each link provides deeper detail on the concepts discussed:
-
We know how to build design systems, but we don’t know how to operate them (Murphy Trueman, May 2026): https://www.designsystemscollective.com/we-know-how-to-build-design-systems-but-we-dont-know-how-to-operate-them-0f06e79de7fb — The canonical piece on the DS vs ops gap. Covers observability (three pillars), SLOs, deprecation lifecycle, golden paths.
-
Stop policing your design system (Murphy Trueman, Jan 2026): https://blog.murphytrueman.com/stop-policing-your-design-system/ — Health metrics vs compliance metrics. Component integrity scores, token consistency tracking, CI/CD integration. The “measure health, not compliance” framing.
-
Using Observability Tools to Track Frontend Migrations (Adrianne Daley, Jan 2026): https://adriannedaley.com/blog/2026/01/23/observability-for-migrations.html — Concrete implementation: react-scanner → OTel → Honeycomb. Code configs, gotchas (maxQueueSize), dashboards.
-
How to govern a design system across 10+ consumer projects (Adam Arant, Apr 2026): https://adamarant.com/en/blog/how-to-govern-a-design-system-across-10-consumer-projects — Seven-step practical framework. Contribution models, frozen API, SemVer, codemods, manifest-per-consumer, deprecation timelines, fixed cadence.
-
Agents are the most honest reviewers your design system will ever have (Murphy Trueman): https://blog.murphytrueman.com/agents-are-the-most-honest-reviewers-your-design-system-will-ever-have/ — Agent traces as free telemetry. Four signal types (hallucinated props, bypassed tokens, novel components, corrective turns).
-
Design System Tooling 2026 (DesignSystems.one): https://www.designsystems.one/designops/tooling — The Figma → MCP → Agent → Repo pipeline. Token tooling landscape. Wire-it-up commands for Claude Code, Cursor, Codex.