The title list is never the problem you think it is
You export the HRIS role list three weeks before the board meeting and open it expecting a clean column of job titles. What you get instead is 340 rows of "Sr. Financial Analyst," "Financial Analyst II," "Financial Analyst — FP&A," and one lone "Financial Analyst (Contract)" that someone typed by hand in 2019. None of these strings exist in a standardized occupational taxonomy. All of them need to.
This is not a data-entry failure. It's what internal job title lists always look like, because they were built for payroll and org charts, not for occupational classification. Standardized systems like ONET describe 1,016 occupational titles across 923 data-level occupations, covering more than 55,000 individual job titles in practice, according to the ONET Resource Center. Your HRIS export was never going to line up with that vocabulary on its own — and it shouldn't have to. That's the job fuzzy matching is built to do.
This article covers how fuzzy matching works, what a confidence score actually measures, and why — no matter how good the match — a person still has to confirm it before it feeds into any exposure analysis.
What fuzzy matching actually does
Fuzzy matching compares a messy input string against a reference list of standardized labels and returns the closest candidates, ranked by similarity, rather than requiring an exact string match. Instead of looking for "Financial Analyst" and failing on "Sr. Financial Analyst," the algorithm measures how close the two strings are — accounting for word order, abbreviations, pluralization, and extra tokens like seniority markers or department tags — and surfaces the O*NET occupation it most likely corresponds to.
This matters because O*NET occupations are drawn from the 2018 Standard Occupational Classification system, which organizes 867 detailed occupations into 459 broad occupations, 98 minor groups, and 23 major groups, per the U.S. Bureau of Labor Statistics. That's a deliberately compact, standardized structure — it was never going to contain every title variant a mid-sized company invents over a decade of job postings. Fuzzy matching is the bridge between the two vocabularies: the messy, company-specific one you have, and the standardized, comparable one you need for any credible O*NET SOC code lookup by job title.
What fuzzy matching does not do is guarantee correctness. A high similarity score tells you two strings look alike. It does not tell you the role actually performs the tasks associated with that occupation. That distinction is the entire reason a confidence score exists — and the entire reason it isn't the last step.
Inside the match: how a confidence score gets built
A confidence score is a way of quantifying how sure the system is that a proposed O*NET match is correct, so a person reviewing the output knows where to spend their attention. Here's a simplified, worked illustration of the kind of factors that go into one — using round, illustrative numbers to show the mechanics, not a claim about any specific dataset.
Say the input title is "Sr. Financial Analyst — FP&A." The candidate match is O*NET-SOC occupation 13-2051.00, Financial Analysts. A matching process might score this on a few weighted factors:
- Core term overlap (does "Financial Analyst" appear intact?) — strong signal, weighted heavily.
- Noise-token handling (does the system correctly discount "Sr.," "FP&A," and other modifiers rather than treating them as mismatches?) — moderate signal.
- Alternate title match (O*NET occupations list known alternate titles; if "Financial Analyst II" or "Senior Financial Analyst" already appears as a recognized alternate title for 13-2051.00, that's a strong confirming signal.)
- Ambiguity check (is there a second O*NET occupation — say, a budget analyst or a management analyst code — that scores nearly as well? A close second candidate should lower confidence, not just get discarded.)
Combine those into an illustrative composite: strong core-term overlap plus a recognized alternate title, minus a small deduction for the unrecognized "FP&A" token, might land around a 90 out of 100. That's a high-confidence match. Now take a genuinely ambiguous title like "Program Manager" — a string that could reasonably map to a project management occupation, a general/operations management occupation, or a nonprofit-program-specific one — and the same kind of scoring might land closer to 55 or 60, with two or three candidates clustered near each other. That's not a system failure. That's the system correctly telling you the title itself is ambiguous and a person needs to look at it.
Why every match still needs a human to confirm it
This is the point that separates a useful crosswalk from a risky one: a confidence score is a triage signal, not a verdict. High-confidence matches still deserve a spot check, because a title can score well on string similarity while describing a role that, at your company, does meaningfully different work than the O*NET occupation's task list assumes. Low- and mid-confidence matches need a named person to make the call, informed by what the role actually does day to day — not just what it's called.
This is the same discipline that governs everything downstream. An O*NET occupation match feeds directly into task-level exposure scoring — the four-dimension rubric (cognitive routine, physical routine, social/judgment, and creative, each scored 0–100) that produces a role's exposure profile and its classification as Monitor, Review, or Redeployment Candidate. A wrong occupation match at the front end doesn't just mislabel a title. It quietly corrupts every task-importance weight and every downstream classification built on top of it. Get the crosswalk wrong, and you're not producing an input to human judgment — you're producing noise dressed up as analysis.
That's why a defensible workflow treats fuzzy matching as the first pass, not the finish line: auto-accept the clear high-confidence matches, route mid-confidence matches to a reviewer with the two or three closest alternates shown side by side, and require an explicit manual override — with a note on why — for anything ambiguous or unmatched. If you haven't yet walked through that full workflow end to end, the companion piece on how to map job titles to O*NET occupations covers it step by step, and the SOC code crosswalk tool guide goes deeper on building a reusable mapping table rather than a one-off cleanup.
Building a crosswalk you can reuse, not just a list you clean once
The title list you clean today will look different next quarter — new postings, new internal title conventions, another acquisition with its own HRIS export. A fuzzy-matched crosswalk is only worth building if it's structured to be reused: every title mapped once, every override documented with a reason, every occupation code tied back to its ONET source data (the ONET database is refreshed on a regular update cycle, with a primary annual update, per the U.S. Department of Labor) so the mapping doesn't silently go stale. If you're new to what O*NET actually is and why it's the reference standard here, what is O*NET and how it works is the right starting point.
Once titles are matched and confirmed, occupation-level data — task lists, importance ratings, and the descriptors that feed exposure scoring — can be pulled consistently across the whole organization, department by department, instead of being reconstructed by hand every time someone asks "which roles look exposed?" That consistency is what turns a messy title export into something a board can actually rely on. You can see how that full pipeline, from matched titles to department heatmaps, is priced and packaged on the pricing page, and the O*NET Task-Mapping Methodology Handbook walks through the crosswalk-building process in template form if you'd rather build the mapping table yourself before running it through a scoring engine.
This site incorporates information from ONET. Used under the CC BY 4.0 license. ONET is a trademark of USDOL/ETA.
