Why one number on a heatmap causes so much anxiety
Three weeks before a board meeting, a People Analytics Lead opens a spreadsheet she built at a private equity owner's request. Every role in the company sits in a row. Every row has a color: green, yellow, orange, red. Her VP of Customer Support scrolls to his department and sees a red cell next to "Support Representative — Tier 1." He asks the only question that matters to him: "Are we cutting this team?"
That question is the wrong one to ask of the number, and it is also the most natural question in the world to ask. An AI exposure score was never built to answer it. This article explains what the score is actually measuring, what the Low/Moderate/High/Critical bands represent, and how to read a result without turning a diagnostic input into a verdict nobody asked the model to deliver.
What the score is actually measuring
An AI exposure score is a measure of task composition, not a prediction about a person's employment. It describes how much of a role's day-to-day work — as defined by its O*NET occupational profile — consists of tasks that current AI systems are structurally well-suited to assist with or partially perform, versus tasks that depend on physical presence, situational judgment, or original creative work.
ONET, maintained by the U.S. Department of Labor's Employment and Training Administration, is the right foundation for this because it is granular and standardized. The database currently covers 1,016 occupational titles (923 with full data, spanning more than 55,000 job titles in practice) and describes each one with roughly 277 descriptors, updated on a regular cycle with a primary annual refresh each third quarter. That granularity is what allows an exposure score to be built from actual task statements rather than a job title alone. Two roles both called "Analyst" can carry very different scores once you look at what their underlying ONET tasks actually involve.
This site incorporates information from ONET. Used under the CC BY 4.0 license. ONET is a trademark of USDOL/ETA.
The four dimensions behind every exposure score
Every task in a role's O*NET profile is scored against four dimensions, each on a 0–100 scale:
- Cognitive routine — how standardized, rules-based, and pattern-repeatable the cognitive work is.
- Physical routine — how much of the task depends on repeatable physical action versus adaptive physical presence.
- Social/judgment — how much the task depends on negotiation, persuasion, mentorship, or contextual human judgment.
- Creative — how much the task depends on original idea generation, design, or synthesis without a clear precedent.
Each task also carries an O*NET-defined importance weight — how central that task is to the occupation overall. The final role-level score is a weighted composite: dimension scores multiplied by task importance, summed across all of a role's tasks, and normalized. That weighting matters. A task that occupies 5% of a role's time and scores high on cognitive routine should not move the needle as much as a task that occupies 40% of it.
If you want the occupation-level detail behind this — how O*NET task and skill data map onto exposure specifically — the companion piece on how AI exposure is scored by occupation walks through that mapping in more depth. For the index construction itself, see how the AI occupational exposure index is built.
How Low, Moderate, High, and Critical map to action, not automation
The score bands exist to route a role toward a workflow, not to hand down a sentence. Read them this way:
- Low / Monitor — the role's task composition leans on physical presence, situational judgment, or creative synthesis in ways current tools do not structurally replicate. Action: revisit on the normal planning cycle.
- Moderate / Monitor–Review — a meaningful share of tasks show AI-assist potential, but the composite is not concentrated. Action: watch for tooling changes; no immediate reallocation needed.
- High / Review — a large share of the role's weighted task time sits in cognitive-routine or pattern-repeatable territory. Action: this role deserves a structured conversation about task redesign, tooling investment, or skill development — not a headcount decision made from the score alone.
- Critical / Redeployment Candidate — the composite is heavily concentrated in the most automatable dimensions across the role's highest-importance tasks. Action: this is the strongest signal in the model, and it still means "flag this role for a redeployment planning conversation," not "eliminate this role." A Redeployment Candidate label points toward internal mobility pathways — skill-adjacent occupations the same O*NET data can surface — as the next analytical step, not toward a layoff list.
Notice what none of the four labels say: "automated," "safe," or "replaced." That is deliberate. The score describes task composition today, scored against a rubric. It does not model your company's technology roadmap, your customers' tolerance for AI-assisted service, your union agreements, your budget cycle, or your leadership's actual intentions. Those inputs belong to the humans in the room, not the spreadsheet.
A worked example: scoring a single role
Here is a simplified, illustrative calculation — round numbers, for teaching the method, not a real published score.
Take a Payroll Specialist role with three representative O*NET-style tasks:
| Task | Importance (1–5) | Cognitive routine | Social/judgment | Creative |
|---|---|---|---|---|
| Process biweekly payroll runs in established software | 4.5 | 85 | 20 | 5 |
| Reconcile payroll discrepancies with HR and Finance | 4.0 | 40 | 65 | 10 |
| Advise employees on benefits elections | 3.5 | 15 | 80 | 15 |
To build a composite cognitive-routine score, multiply each task's dimension score by its importance weight, sum, and divide by total importance:
(85 × 4.5) + (40 × 4.0) + (15 × 3.5) = 382.5 + 160 + 52.5 = 595 Total importance = 4.5 + 4.0 + 3.5 = 12 Weighted cognitive-routine score = 595 ÷ 12 ≈ 49.6
Run the same weighting across the social/judgment column and the creative column, and the role ends up with a composite profile that leans moderate-to-high on cognitive routine but is meaningfully offset by social/judgment work in reconciliation and benefits conversations. That combination is exactly why the score is reported across all four dimensions rather than collapsed into a single automation percentage — a role can carry real exposure in one dimension and real durability in another, and a manager needs to see both to make a sensible call.
What the score doesn't know — and why judgment still matters
The rubric is transparent by design, which means its limits are visible too. It does not know that your company just signed a three-year contract requiring human sign-off on every reconciliation. It does not know that your Tier-1 support team is also your best source of product feedback. It does not know your local labor market, your union contract, or your customers' preferences.
This is consistent with what the broader research says about AI and work: the signal is real, but it is a starting point for planning, not a finished plan. McKinsey's 2023 research estimated that 60–70% of work hours across the economy could theoretically be automated with current and emerging technology, up from roughly 50% in earlier estimates — a technical ceiling, not a company-specific forecast. The World Economic Forum's Future of Jobs Report 2025 estimated 39% of core skills will be transformed or outdated by 2030, and separately found that of every 100 workers needing new training, 29 are expected to be upskilled in their current role and 19 reskilled or redeployed internally — figures that describe an economy-wide skills transition, not a verdict on any specific team.
Gartner's 2024 research adds a useful caution from the planning side: 86% of HR leaders report they have not implemented strategic workforce planning, and 66% say their planning work is still limited to headcount forecasting rather than skills or task-level analysis. An exposure score is only useful inside a planning process that can actually receive it — feeding a High or Critical result into a redeployment or upskilling conversation, not a spreadsheet that ends at the color-coding.
What macro research tells us — and what it can't tell you about your company
Anthropic's Economic Index found that the share of occupations using its models for at least a quarter of their tasks grew from 36% in early 2025 to 49% on a pooled basis later in the year — evidence that usage is broadening across occupation types, not evidence about your organization's specific roles. Goldman Sachs' 2023 analysis estimated a global equivalent of 300 million full-time jobs exposed to automation in some form. Brookings' 2024 research found the highest-exposure sectors cluster in STEM, business and finance, engineering, and law, with roughly 12.9 million workers — about a third of those occupations — classified as highly exposed.
Every one of these figures describes a labor market. None of them describes your payroll team, your support desk, or your compliance function. That gap between macro research and company-specific reality is exactly what a role-level exposure score, built on your actual O*NET-mapped role list, is designed to close — carefully, and still only as an input.
Turning a score into a next step
A score without a next action is just a color on a spreadsheet. The useful sequence looks like this: map your roles to O*NET occupations, generate the four-dimension composite for each, group results into department-level heatmaps, and route High and Critical roles into a structured conversation — about task redesign, tooling investment, or internal redeployment pathways — rather than a binary stay-or-go decision.
If you want a starting structure for that conversation, the AI Exposure Scorecard template lays out the four dimensions, the band definitions, and a worksheet for translating a score into a Monitor, Review, or Redeployment Candidate action plan your leadership team can actually use. You can also see how the underlying pricing tiers scale with organization size on the pricing page, browse the full template store, or read more on exposure methodology on the blog.
The number on the heatmap is real. What it means is narrower — and more useful — than most readers assume the first time they see a red cell next to a name they recognize.
