A board question with no clean answer
Three weeks before a board meeting, a People Analytics Lead gets a one-line request from the CFO: "Which of our roles are exposed to AI, and what are we doing about it?" It sounds like a simple ask. It isn't. The question quietly assumes that "exposed to AI" and "at risk of automation" mean the same thing — and that assumption is where most workforce AI analyses go wrong before they even start.
Exposure and automation risk are related, but they are not interchangeable, and the difference is not academic. Score a role's exposure and stop there, and you get a task-level map of where AI-capable tools intersect with the work. Treat that same score as a risk verdict, and you get something else entirely: a claim about outcomes that no task-level analysis — including this one — is actually built to make.
This article draws the line clearly, because the teams that skip this distinction tend to make one of two costly mistakes: over-reacting to roles that are highly exposed but poorly suited to full automation, or under-reacting to roles where moderate exposure masks a real redeployment need. Getting the framing right is the difference between a defensible board answer and a number nobody can stand behind under follow-up questions.
What "AI exposure" actually measures
Exposure, in the sense used across the O*NET-grounded analysis in this piece, is a measure of task-level intersection — how much of a role's day-to-day work touches activities that current AI tools are capable of performing or substantially assisting with. It is descriptive, not predictive. It answers "where does the work overlap with what AI can do," not "what will happen to this job."
That distinction matters because ONET itself was never designed to produce automation forecasts. It's an occupational classification system — currently covering 1,016 occupational titles and 923 data-level occupations across roughly 55,000 job titles in the U.S. economy, per the ONET Resource Center — built around detailed task, skill, and knowledge descriptors (roughly 277 of them, updated on a regular cycle by the U.S. Department of Labor). Exposure scoring layers an AI-capability lens on top of that existing occupational data. It inherits O*NET's descriptive intent, not a forecasting mandate it was never given.
A defensible exposure methodology scores each task against four separate dimensions, each rated 0–100:
- Cognitive routine — how standardized and rule-based the reasoning or information-processing component of the task is.
- Physical routine — how standardized and repeatable any physical component is.
- Social/judgment — how much the task depends on interpersonal negotiation, trust, or context-specific judgment calls.
- Creative — how much the task depends on originality, novel synthesis, or generative problem-solving rather than pattern-matching.
Each task in a role is weighted by its O*NET-defined importance before the four scores roll up into a single role-level exposure figure. A role is not "high exposure" or "low exposure" in the abstract — it is high or low on a specific combination of these four dimensions, and that composition is exactly what a raw single-number automation-risk label discards.
What "automation risk" implies — and why that's a different claim
"Automation risk" carries a payload that "exposure" does not: an implicit claim about outcome. It suggests a probability that a role will actually be automated — displaced, restructured, or eliminated — as opposed to a description of where the work currently overlaps with AI capability.
That implicit claim is not something a task-level exposure model, or any tool built on O*NET data, can responsibly make. Whether a highly exposed task is actually automated depends on variables outside the occupational data entirely: a company's technology budget, its risk tolerance, labor market conditions, regulatory constraints, union agreements, customer expectations, and dozens of judgment calls that sit with human leaders, not with a scoring rubric.
This is not a hedge — it's a structural fact about what the underlying research actually supports. McKinsey's 2023 analysis of generative AI's economic potential found that 60–70% of work hours could theoretically be automated with current and emerging technology, up from roughly 50% in prior estimates — a statement about technical feasibility, not a prediction of what will occur inside any specific company. Anthropic's 2026 research on labor market impacts found that only about 68% of actual AI usage maps to fully-feasible task completion, while roughly 3% falls into tasks the tools cannot meaningfully perform at all — meaning even where AI is being used heavily, the picture is one of partial task assistance far more often than full task replacement. Exposure describes potential overlap; what an organization does with that overlap is a separate decision, made by people, not by the score.
Conflating the two collapses a nuanced, four-dimensional, task-level picture into a binary that the underlying data was never built to support.
Why the conflation leads to bad workforce decisions
Treating exposure as a verdict tends to produce two failure modes, and both are expensive.
The first is premature action on the wrong roles. A role can score high on cognitive-routine exposure — meaning a large share of its reasoning tasks overlap with what AI tools handle well — while scoring low on social/judgment and creative dimensions. That combination often describes work that is a strong candidate for augmentation (tools handling first-draft analysis, research synthesis, or routine documentation) rather than full replacement. Read the composite number as "at risk," and a leadership team may restructure or reduce a function that is actually becoming more valuable once routine sub-tasks are offloaded — just organized differently.
The second is under-reaction where it's needed most. A role with a moderate composite exposure score can still contain a specific cluster of high-exposure tasks worth redesigning around now — before attrition or a hiring freeze forces the decision under worse conditions. A single "risk: moderate" label hides that texture. This is precisely why the World Economic Forum's Future of Jobs Report 2025 frames the transition in terms of skills, not job elimination: it projects 39% of workers' current skill sets will be transformed or become outdated by 2030, and separately estimates that of every 100 workers needing training, 59 need it — with 29 upskilled in their current role and 19 reskilled and redeployed internally, versus only 11 unlikely to receive that training at all. That is a picture of skill transformation and internal redeployment, not blanket job loss — and it only becomes visible when exposure data is read at the task and skill level rather than compressed into a single risk score.
Both failure modes trace back to the same root error: reading a task-level, multi-dimensional exposure figure as if it were a single-dimension outcome prediction.
A worked example: same score, two different roles
Consider two roles that land at a similar composite exposure figure — for illustration, call it 62 out of 100 on each — arrived at through very different underlying compositions.
Role A (illustrative): cognitive routine 85, physical routine 20, social/judgment 40, creative 25. The high cognitive-routine score suggests a large share of the role's reasoning tasks — data reconciliation, standard reporting, first-pass classification — overlap heavily with current AI capability. The low creative and moderate social/judgment scores suggest the remaining work leans toward exception-handling and stakeholder communication.
Role B (illustrative): cognitive routine 45, physical routine 70, social/judgment 55, creative 30. Here the exposure is driven by standardized physical-task components rather than cognitive ones — a very different intervention path, since the tools relevant to physical-task exposure (robotics, process automation) are not the same tools relevant to cognitive-task exposure (language and reasoning models).
Same composite number. Two entirely different conversations about what to monitor, what to redesign, and what — if anything — to reskill toward. This is why a defensible exposure output surfaces the four underlying dimensions and the task list behind them, not just the rolled-up figure.
Reading the score correctly: exposure as input, not instruction
The practical output of a well-built exposure analysis is not a red/yellow/green risk label. It's a classification that keeps the human decision explicit: Monitor, for roles where exposure is present but the current task mix doesn't warrant near-term action; Review, for roles where a leadership team should examine the specific task composition and make a deliberate call; and Redeployment Candidate, for roles where the task-level pattern suggests a structured internal transition is worth actively planning — not roles being flagged for elimination.
None of those three labels predicts what will happen. All three describe what deserves a closer look, and by whom. The analysis behind them draws on ONET data used under license — this site incorporates information from ONET, used under the CC BY 4.0 license; O*NET is a trademark of USDOL/ETA — and applies the four-dimension rubric described above at the individual task level before rolling anything up to a role or department view.
If you want the deeper mechanics of how a raw score turns into one of these three labels, see what an AI exposure score actually measures and the companion piece on how augmentation differs from automation of jobs. For a department-by-department view of how exposure varies across common mid-market org structures, AI exposure by occupation, explained walks through the pattern in more detail, and which jobs AI is more likely to augment versus replace extends the Role A / Role B logic above across a broader set of occupations.
Bringing this back to the board conversation
Back to the CFO's question. The honest, defensible answer isn't a list of "safe" and "unsafe" roles — it's a task-level map of where AI capability and current work genuinely overlap, broken out by the four dimensions that actually drive different responses, with a clear label for which roles warrant monitoring, review, or a deliberate redeployment conversation. That is a harder answer to deliver in one sentence. It is also the only one that survives a follow-up question about how the number was built.
Teams that want a structured way to walk their own leadership through this distinction — before the board meeting, not during it — can start with the Workforce AI Readiness Assessment Guide, which lays out the reading and decision framework in more detail. To see how the underlying scoring and tiers are packaged for a full org chart, the pricing page breaks down what's included at each tier.
If this framing is useful, subscribe below for future pieces on reading exposure data responsibly — one distinction at a time, starting with this one.
