Why a simple average of task scores misleads a board
A People Analytics Lead building a role-level exposure picture for the first time usually starts the same way: list the role's tasks, score each one for AI exposure, average the scores, and call that the role's number. It feels rigorous. It is also quietly wrong, and it is the kind of wrong that surfaces at the worst possible moment — when a CFO or board member asks why two roles with the same average score are treated so differently in the redeployment plan.
The problem is that a simple average treats every task as equally central to the job. In practice, it never is. A financial analyst role might include a low-importance task like formatting slide decks alongside a high-importance task like building forecasting models. Both can carry a similar raw exposure score, but they do not carry the same weight in what the job actually is. An exposure score that does not account for that difference is not measuring the role — it is measuring an arbitrary list of activities someone happened to write down.
This is where importance weighting — the practice of scaling each task's exposure score by how central that task is to the occupation, using O*NET's importance ratings — turns a rough estimate into a number a board can actually rely on. It is also one of the core mechanics behind our own four-dimension task exposure framework, and worth understanding on its own terms before you apply it to a live role list.
What O*NET importance ratings actually measure
ONET — the U.S. Department of Labor's Occupational Information Network — maintains a structured database covering 1,016 occupational titles and 923 data-level occupations representing more than 55,000 jobs across the U.S. economy, built from roughly 277 descriptors and updated on a regular cycle, with a primary annual update typically landing in the third quarter of each year. Among those descriptors are task importance ratings: for every task associated with an occupation, ONET records not just whether the task is performed, but how important that task is to successful performance in the role, based on structured occupational analysis.
That distinction matters because a task list on its own is flat. Importance ratings give it shape. They tell you which tasks are load-bearing for the occupation and which are peripheral — information a simple task inventory cannot provide on its own. We go deeper on how these ratings are constructed and what they cover in O*NET task importance ratings explained; the short version for scoring purposes is that importance data lets you move from "here is a list of things this role does" to "here is what actually defines this role."
This site incorporates information from ONET. Used under the CC BY 4.0 license. ONET is a trademark of USDOL/ETA.
The four-dimension framework, applied with weights
Our transparent AI exposure rubric scores every task a role performs across four dimensions, each on a 0–100 scale:
- Cognitive routine — how standardized and rule-based the mental work is
- Physical routine — how standardized and repeatable the physical work is
- Social/judgment — how much the task depends on relationship management, negotiation, or contextual judgment calls
- Creative — how much the task depends on original synthesis, design, or novel problem-solving
Scoring a task on these four dimensions produces a task-level exposure profile. The question importance weighting answers is: once you have that profile for every task in a role, how do you combine them into a single role-level score without letting a marginal task distort the picture?
The mechanics are described in full in our task-based AI exposure methodology writeup, but the underlying logic is straightforward: each task's four-dimension score is multiplied by its O*NET importance weight before it is rolled into the role total, rather than simply averaged with every other task at equal standing. High-importance, high-exposure tasks push the role score up more than low-importance tasks with the same exposure profile. That is the entire point.
A worked example: two roles, identical task lists, different weights
Consider a simplified example — round numbers, illustrative only, not drawn from any live client role.
Say a role has three tasks, each scored on cognitive-routine exposure (the dimension most relevant here) and each carrying an O*NET importance weight on a 1–5 scale:
| Task | Cognitive-routine score | Importance weight |
|---|---|---|
| Task A | 80 | 5 |
| Task B | 60 | 3 |
| Task C | 20 | 1 |
A simple average gives you (80 + 60 + 20) ÷ 3 = 53.3.
An importance-weighted average gives you: (80×5 + 60×3 + 20×1) ÷ (5+3+1) = (400 + 180 + 20) ÷ 9 = 66.7
The gap between 53.3 and 66.7 is not noise. It is the difference between a role where the dominant, most-central task is highly routine — which the simple average understates because it lets a marginal task (Task C) drag the number down — and the weighted score, which reflects the fact that Task A is what this job is actually for. Flip the weights so the low-exposure task carries the most importance, and the two methods diverge in the opposite direction. Either way, the simple average is systematically less accurate at the tails, which is exactly where a board is going to ask the sharpest questions.
Where weighting changes the Monitor / Review / Redeployment Candidate call
This is not an academic distinction. Our scoring output classifies roles into three tiers — Monitor, Review, or Redeployment Candidate — as an input to human judgment, never as a verdict on any individual's job. Whether a role lands in one tier or the next is often decided at the margin, and the margin is exactly where a 13-point gap between a simple average and a weighted average matters most.
A role that scores 53 under a simple average might sit comfortably in Monitor. The same role at 67 under a weighted average might land in Review — a materially different conversation with the department head. Getting the weighting right is not a cosmetic improvement to the math. It is the difference between a defensible recommendation and one that falls apart under the first follow-up question.
Importance weighting does not change what a role's tasks are. It changes whether your score reflects what the role actually does most, or an arbitrary average of everything it does at all.
None of this tells you what to do about a given score — that remains a judgment call for the people leading the redeployment conversation, informed by context the model does not have.
Building this into a repeatable process
The value of importance weighting compounds when you are scoring dozens or hundreds of roles rather than one. Applied by hand across a full department, task-by-task importance weighting in a spreadsheet is exactly the kind of exacting, repeatable work that consumes analyst-hours and invites transcription error at scale. That is the structural gap self-serve, task-level tooling is built to close — applying the same weighted methodology consistently across every role, rather than re-deriving it manually each time.
If you are formalizing this process for your own workforce assessment, our O*NET Task-Mapping Methodology Handbook walks through the full weighting workflow with downloadable templates, and our pricing page outlines the tiers available if you want the scoring automated end to end rather than rebuilt from scratch.
