The board doesn't ask about jobs — it asks about work
Three weeks before the board meeting, a People Analytics Lead gets a one-line request from the company's PE owner: "Which roles are exposed to AI, and what's the plan?" The instinct is to open the org chart and start labeling job titles — this one's fine, that one's at risk. It takes about ten minutes of trying before the exercise falls apart. A "Financial Analyst" who spends most of the week reconciling spreadsheets is a very different exposure story than a "Financial Analyst" who spends most of the week in stakeholder negotiations and judgment calls about deal structure. Same title. Same box on the chart. Completely different work.
This is the point at which a credible AI exposure exercise stops asking "is this job exposed" and starts asking "which tasks inside this job are exposed, and how much does each one matter to the job as a whole." That shift — from job-level guessing to task-level, importance-weighted scoring — is the entire methodology behind a defensible task based AI exposure methodology, and it's the difference between an analysis a CFO will accept and one they'll pick apart in the first five minutes. This article walks through why tasks are the right unit of analysis, how the underlying O*NET data makes that possible, and how importance-weighted scoring turns a list of tasks into one role-level number a board can act on.
Why "is this job exposed to AI" is the wrong question
Job titles are administrative conveniences. They group together a bundle of tasks for payroll, reporting, and org-design purposes, but they say almost nothing about the actual composition of a person's week. Two people with the same title in two different departments — or even the same department, six months apart — can have meaningfully different task mixes. A job-level exposure label averages all of that away before the analysis even starts, which means it's already wrong for a meaningful share of the people it's applied to.
Macro research on AI's labor impact makes the same point at scale. McKinsey's 2023 analysis of generative AI's economic potential found that 60–70% of employees' work hours could theoretically be automated with current and anticipated technology, up from roughly 50% in pre-generative-AI estimates — a jump driven entirely by which tasks generative tools newly reach, not by whole occupations disappearing. Anthropic's Economic Index reported that 36% of occupations in early 2025 (49% on a pooled basis) used Claude for at least a quarter of their associated tasks — again, a task-level measurement, not a job-level verdict. Every credible exposure framework we're aware of operates at the task level for the same reason: tasks are where the actual work — and the actual automation potential — lives.
This is also why a company-specific exposure exercise can't stop at reading macro research and applying it to a job title. Sector-level findings tell you where to look; they don't tell you what a specific role at a specific company actually does. For a deeper walkthrough of applying macro research to your own org chart, see our guide to running a company-specific AI workforce exposure assessment.
Inside O*NET: how the government already broke jobs into tasks
The good news is that nobody has to invent a task taxonomy from scratch. The U.S. Department of Labor's Occupational Information Network (ONET) has already done it, at a level of detail few HR teams realize exists. ONET currently documents 1,016 occupational titles, of which 923 have detailed data, covering an estimated 55,000+ distinct jobs in the U.S. economy (O*NET Resource Center, USDOL/ETA, 2019). Each occupation is described by roughly 277 standardized descriptors — including a list of specific tasks — that the Department of Labor updates on a regular cycle, with the primary annual refresh landing in the third quarter (U.S. DOL, 2025).
Two things make this data usable for exposure scoring rather than just descriptive. First, every task attached to an occupation carries an importance rating — a measure of how central that task is to the job, independent of how often it's performed. Second, occupations map cleanly onto the Bureau of Labor Statistics' Standard Occupational Classification system, which organizes 867 detailed occupations into 459 broad occupations, 98 minor groups, and 23 major groups under the 2018 SOC structure (U.S. BLS). That crosswalk matters practically: it's what lets a company's messy, inconsistent internal job titles get mapped onto a standardized occupational baseline before any scoring happens. We cover that mapping step, including how occupational titles are matched and where judgment calls are still required, in our explainer on AI exposure by occupation.
The task-importance rating deserves its own explanation, because it's the single most load-bearing number in the entire methodology — it's what allows a handful of task scores to roll up into one role score without simply averaging everything equally. We go deeper on how those ratings are constructed and what they do and don't measure in our piece on O*NET task importance ratings.
This site incorporates information from ONET. Used under the CC BY 4.0 license. ONET is a trademark of USDOL/ETA.
The four-dimension task exposure framework, explained
Once a role's tasks are pulled from O*NET, each individual task gets scored against four dimensions, each on a 0–100 scale:
- Cognitive routine — how standardized and rule-based the thinking involved in the task is. High scores describe tasks like data entry reconciliation or standard report generation; low scores describe tasks that require synthesizing ambiguous or novel information.
- Physical routine — how repeatable and specified the physical component of the task is, where applicable. Most white-collar tasks score low here by default; this dimension matters more for roles with a hands-on component.
- Social / judgment — how much the task depends on reading people, navigating organizational politics, negotiating, or making a call under genuine uncertainty with limited precedent. High social/judgment content pulls a task's exposure down, not up.
- Creative — how much the task depends on generating genuinely novel output rather than recombining known patterns. Like the social/judgment dimension, high creative content is a downward pressure on exposure, not an upward one.
Each task in a role's O*NET task list gets a score on all four dimensions. That's a deliberate design choice: a single "exposure score" per task would collapse information that's actually useful separately. A task can be high on cognitive routine and still carry meaningful social/judgment weight — think of a routine-sounding task like "prepare compliance documentation" that nonetheless requires judgment calls about materiality. Scoring the four dimensions independently, rather than blending them into one number too early, is what makes the framework auditable — a reviewer can ask "why is this task scored this way on this dimension" and get a specific answer, rather than being handed an opaque composite. We walk through each dimension in more depth, including illustrative task examples for each, in our full four-dimension task exposure framework guide.
A worked example: importance-weighted task exposure scoring
Here's how the pieces fit together, using round, illustrative numbers to show the arithmetic rather than any specific company's real result.
Say a role has been mapped to an ONET occupation with six core tasks. Each task carries an ONET importance rating (on O*NET's underlying 1–5 scale, normalized here to a 0–100 weight for clarity) and gets scored on the four exposure dimensions. A simplified version might look like this:
| Task | Importance weight | Cognitive routine | Physical routine | Social/judgment | Creative |
|---|---|---|---|---|---|
| Reconcile monthly account entries | 90 | 85 | 5 | 10 | 5 |
| Prepare standard variance reports | 75 | 80 | 5 | 15 | 10 |
| Review flagged discrepancies with department heads | 70 | 30 | 5 | 75 | 15 |
| Recommend process changes to close cycle | 55 | 25 | 5 | 60 | 45 |
| Train junior staff on close procedures | 40 | 20 | 10 | 70 | 20 |
| Respond to ad hoc leadership requests | 60 | 15 | 5 | 65 | 40 |
To get the role's score on any single dimension, each task's score on that dimension is weighted by its importance rather than simply averaged. For cognitive routine, that means multiplying each task's cognitive-routine score by its importance weight, summing those products, and dividing by the sum of the importance weights. Using the numbers above, the importance-weighted cognitive routine score comes out in the high 40s — noticeably lower than the simple average of the six raw scores (about 43 either way in this particular example, but in most real task lists the gap between weighted and unweighted results is larger, because the highest-importance tasks and the highest-cognitive-routine tasks don't always overlap).
That's the mechanical core of importance-weighted task exposure scoring: high-importance, high-cognitive-routine tasks pull the role's score up; high-importance tasks with heavy social/judgment or creative content pull it down; low-importance tasks — whatever their raw scores — have proportionally less influence either way. The same weighting procedure runs independently for all four dimensions, producing four dimension-level scores per role before anything gets blended into a single figure. This is deliberately more work than eyeballing a job description, and that's the point — it's what lets someone reproduce the number from the same inputs and get the same answer. For the full weighting formula, including how ties and missing data are handled, see our detailed guide to importance-weighted task exposure scoring.
From task scores to a role-level result: Monitor, Review, Redeployment Candidate
Four dimension scores per role are informative, but a board wants one legible output per department, not four numbers per person. The four dimensions blend into a single composite exposure score, which then places the role into one of three bands:
- Monitor — current task composition shows low-to-moderate exposure; no near-term action indicated beyond normal periodic review.
- Review — exposure is elevated enough, or task composition is shifting quickly enough, that the role warrants a closer look at workflow, tooling, or training investment.
- Redeployment Candidate — task composition suggests significant capacity may be freed up as AI tools mature, and the role is a candidate for a structured internal redeployment conversation grounded in adjacent-occupation skill overlap.
Notice what none of those bands say. None of them say a role is "safe." None of them say a role "will be automated" or "will be replaced." That's not a branding choice — it's a boundary the methodology enforces on purpose. An exposure score describes the task composition of a role today, against a rubric that measures automatability, not a prediction about a specific person's employment outcome. Employment decisions depend on factors an occupational task list cannot see: budget, strategy, performance, tenure, internal politics, timing. The score is an input to a human decision, not a substitute for one. If you want a broader walkthrough of how to translate a set of role scores into an actual departmental plan — including how "Redeployment Candidate" roles get matched to adjacent occupations by skill overlap — our guide on how to measure AI's impact on jobs covers that translation step in detail.
It's also worth sitting with how large the underlying shift is likely to be over the next several years, so the three-band framing doesn't read as either alarmist or dismissive. The World Economic Forum's Future of Jobs Report 2025 projects 170 million jobs created and 92 million displaced globally by 2030 — a churn rate of 22% of today's total employment, netting out to roughly 78 million new jobs — alongside a finding that 39% of workers' current skill sets will be transformed or become outdated by 2030. Handled at the task level, with importance weighting and an explicit refusal to predict individual outcomes, that scale of change is something a workforce plan can actually respond to. Handled at the job-title level, with a red/yellow/green label slapped on an org chart, it's just noise dressed up as analysis.
What task-based scoring can't tell you — and why that's by design
A task-level exposure score cannot tell you whether a specific person will keep their job. It cannot tell you whether your company will actually adopt the AI tools capable of reaching a given task — adoption is a business decision, not a property of the task itself. It cannot account for factors like a role's strategic importance to a specific initiative, a manager's stated performance concerns, or a union contract's protections. And it's only as good as the task-to-occupation mapping underneath it: a role mapped to the wrong O*NET occupation, or one whose day-to-day work has drifted well outside its official task list, will produce a misleading score no matter how careful the weighting math is.
None of that is a flaw in the framework so much as an honest description of its scope. A tool that claimed to predict individual job outcomes from a standardized government task taxonomy would be overstating what any dataset like O*NET can support. The value of task-based scoring is narrower and more defensible: it gives an HR strategy team a consistent, reproducible way to describe where the automatable work concentrates across a department, so that the human judgment calls that follow — retraining budgets, redeployment conversations, hiring freezes, restructuring — are made against a shared, auditable picture rather than a collection of individual hunches about who seems replaceable.
Building this yourself, or not
Everything described above — O*NET task pulls, four-dimension scoring, importance weighting, banding — can be built manually in a spreadsheet by a careful analyst willing to spend the time. Independent consultants who do this kind of mapping work typically charge in the $100–$350/hr range, with a median around $150–$200/hr (ConsultFees, 2026), and a manual build brings the same tradeoff every custom analysis does: it can be tailored precisely to one company, but it isn't repeatable without redoing much of the work the next time the org chart changes.
WorkforceAnalysis runs this exact methodology — O*NET task pulls, four-dimension scoring on cognitive routine, physical routine, social/judgment, and creative, importance weighting, and Monitor / Review / Redeployment Candidate banding — as a self-serve workflow, so a People Analytics Lead can regenerate the analysis after every reorg rather than commissioning it once and letting it go stale.
If you'd rather run the task-mapping step by hand first and see the mechanics up close, the O*NET Task-Mapping Methodology Handbook walks through the full weighting worksheet, including the importance-normalization step used in the worked example above, so you can build a first-pass exposure map for one department before deciding how much of the process to automate. Download the handbook or see current WorkforceAnalysis pricing if you'd rather run the full workflow directly.
