Quick answer: Best in class UNSPSC classification is not one accuracy number applied everywhere. It means every "high confidence"-labeled result is genuinely trustworthy, every lower confidence line is honestly flagged as such, and the categories that matter to the business, the ones bought regularly and negotiated on, get the attention. A one off purchase that lands in a technically correct but unhelpful category is a smaller problem than a regular category that is confidently wrong.
On this page: Why one accuracy number is not the point · An example from a construction supplier · What high confidence should actually mean · Hard FM, cleaning, and construction differ · FAQ
Why one accuracy number is not the point
Many procurement teams I speak to ask the same question first: what accuracy does your classification hit. It is a reasonable question, and it is also the wrong one to lead with. If we invest in a stock, we do not ask only how the company performed last quarter. We ask how confident the analysis behind that number actually is, because a confident number built on thin evidence is worth less than an honest, uncertain one. If we plan a route to an important meeting, we build in the delays we expect, we do not just trust a single estimated time and hope.
Classification accuracy works the same way. A headline percentage tells you almost nothing about whether you can trust the specific lines your reports are built on. What you actually need is to know, line by line, whether a classification is high confidence, medium confidence, or low confidence, and to be able to trust that the high confidence bucket really is high confidence. That is a different, harder requirement than a single accuracy number, and it is the one that determines whether the data underneath your reporting is something you can act on.
An example from a construction supplier
A construction client of ours had a personnel event, and the team bought ribbons, the kind handed out like medals. The classification system placed them under jewelry. Technically, that's not incorrect. Ribbons handed out at an event are, by category definition, closer to jewelry than to anything else in a standard taxonomy. It did not tell the client much of anything useful.
The reason is simple: those ribbons are a one off purchase. Nobody is negotiating a supplier contract for event ribbons, and nobody needs to track that category over time. What procurement actually cares about with a one off purchase is that it lands somewhere reasonable and does not clutter the categories that do matter, not that its classification is philosophically precise. Compare that to a regular purchase, steel fixings ordered every month from the same supplier, and the standard flips. There, classification accuracy is not a nice to have. It is the thing that determines whether the company can negotiate the next contract with real numbers behind it.
What high confidence should actually mean
This is where the review, not the raw model, does the real work. Any AI classification system will produce results with a range of certainty attached, whether or not that certainty is shown to the person using the results. The first thing to check is not the accuracy of the whole dataset. It is whether the system actually tells you which lines are high confidence, medium, and low, and whether you can trust the high confidence bucket specifically, since that is the bucket you will not want to look at again.
Worst case, without that distinction, every team member ends up doing their own research on ambiguous lines, further slowing the process and creating inconsistencies across classification, because nobody has a shared, trustworthy signal for which lines are actually settled. The alternative is not classifying everything perfectly. It is being honest about what is settled and what is not, and putting the limited review time where it changes a decision, on the regular, negotiated categories, rather than on a one off box of ribbons.
Hard FM, cleaning, and construction differ
Best in class looks different in each, because the categories are used differently. Where the same spend recurs every month, the bar is consistency at commodity (item-specific) level. In hard services this shows up like this: a facilities team buying HVAC filters every month needs those lines classified consistently at commodity level, because filter spend feeds a supplier negotiation.
Where the risk is invisible creep, the bar is keeping categories separated over time. An example from cleaning: consumables purchased in bulk, gloves, cloths, chemicals, need to be separated from one off equipment purchases, because consumables creep is a trend that only shows up once the categories are held apart and tracked over time.
Where nothing repeats exactly, the bar is comparability across projects rather than a single correct label. An example from construction: material costs vary project to project, so the standard for best in class there is less about a single perfect category and more about whether the same material, steel fixings, electrical conduit, is classified the same way across every project it appears in, so costs can actually be benchmarked.
What counts as best in class shifts with what the category is used for, not with the industry label on the file. A one off purchase in any of these sectors can tolerate a technically correct but unhelpful classification. A regular, negotiated category cannot. Getting this distinction right, rather than chasing one accuracy figure across the whole dataset, is what separates classification that supports a decision from classification that just looks complete on a dashboard.
Frequently asked questions
What does best in class UNSPSC classification actually mean?
It means every high confidence result is genuinely reliable, every medium or low confidence line is honestly flagged rather than hidden inside an aggregate accuracy figure, and the categories that matter for negotiation and budgeting get the review attention. A single accuracy percentage applied across an entire dataset does not tell you this, because it treats a one off purchase the same as a regularly negotiated category.
Why does classification sometimes look wrong even when it's technically correct?
A classification can follow category definitions exactly and still be unhelpful. Event ribbons handed out at a personnel event, for instance, are closer to jewelry than to any standard category by strict definition, so a classifier calling it jewelry isn't wrong. It's just irrelevant, because the purchase was a one off with no supplier relationship or budget line attached, which is exactly the kind of case where technical accuracy doesn't matter to any decision.
Should a one off purchase get the same classification scrutiny as a regular purchase?
No. A one off purchase only needs to land somewhere reasonable and avoid cluttering the categories procurement actually manages. A regular, negotiated purchase needs consistent, trustworthy classification because it feeds contract negotiation and category level reporting. Treating both the same way wastes review time on cases that do not change any decision.
How does UNSPSC best practice differ between Hard FM and Soft FM, cleaning, and construction?
Hard FM depends on consistent commodity level classification for regularly negotiated categories such as HVAC parts. Cleaning depends on separating recurring consumables from one off equipment purchases so spend trends are visible over time. Construction depends on classifying the same material the same way across different projects, since project to project variation is what makes benchmarking difficult otherwise.
How does Pearstop make classification confidence trustworthy rather than just reporting an accuracy number?
Pearstop scores every line for confidence rather than publishing a single blended accuracy figure, and routes lines below a set confidence threshold to a person for review instead of leaving that judgement inside an opaque model. This lets a team trust the high confidence bucket specifically, and focus limited review time on the regular, negotiated categories where a wrong classification actually changes a decision.
Where can I compare UNSPSC classification against other approaches to see how accuracy is typically measured?
A comparison of the main classification tools and approaches, including how each measures and reports accuracy, is covered in UNSPSC classification tools compared. That comparison focuses on how different vendors define and report an accuracy figure, which is a useful companion to the confidence-level approach described here, since the two questions, which tool to use and how to trust its output, are related but not the same one.
Veelgestelde vragen
What does best in class UNSPSC classification actually mean?
It means every high confidence result is genuinely reliable, every medium or low confidence line is honestly flagged rather than hidden inside an aggregate accuracy figure, and the categories that matter for negotiation and budgeting get the review attention. A single accuracy percentage applied across an entire dataset does not tell you this, because it treats a one off purchase the same as a regularly negotiated category.
Why does classification sometimes look wrong even when it's technically correct?
A classification can follow category definitions exactly and still be unhelpful. Event ribbons handed out at a personnel event, for instance, are closer to jewelry than to any standard category by strict definition, so a classifier calling it jewelry isn't wrong. It's just irrelevant, because the purchase was a one off with no supplier relationship or budget line attached, which is exactly the kind of case where technical accuracy doesn't matter to any decision.
Should a one off purchase get the same classification scrutiny as a regular purchase?
No. A one off purchase only needs to land somewhere reasonable and avoid cluttering the categories procurement actually manages. A regular, negotiated purchase needs consistent, trustworthy classification because it feeds contract negotiation and category level reporting. Treating both the same way wastes review time on cases that do not change any decision.
How does UNSPSC best practice differ between Hard FM and Soft FM, cleaning, and construction?
Hard FM depends on consistent commodity level classification for regularly negotiated categories such as HVAC parts. Cleaning depends on separating recurring consumables from one off equipment purchases so spend trends are visible over time. Construction depends on classifying the same material the same way across different projects, since project to project variation is what makes benchmarking difficult otherwise.
How does Pearstop make classification confidence trustworthy rather than just reporting an accuracy number?
Pearstop scores every line for confidence rather than publishing a single blended accuracy figure, and routes lines below a set confidence threshold to a person for review instead of leaving that judgement inside an opaque model. This lets a team trust the high confidence bucket specifically, and focus limited review time on the regular, negotiated categories where a wrong classification actually changes a decision.
Where can I compare UNSPSC classification against other approaches to see how accuracy is typically measured?
A comparison of the main classification tools and approaches, including how each measures and reports accuracy, is covered in UNSPSC classification tools compared. That comparison focuses on how different vendors define and report an accuracy figure, which is a useful companion to the confidence-level approach described here, since the two questions, which tool to use and how to trust its output, are related but not the same one.
Not sure which UNSPSC code to use?
Paste any product or service description and get the correct 8-digit code instantly — or explore the full UNSPSC tree to understand the hierarchy.

Stephanie Wiechers
CEO & Co-founder, Pearstop
Stephanie leads Pearstop's go-to-market and strategic direction. She works directly with procurement and FM leaders across Europe to understand how data quality affects margins, contracts, and AI readiness.
LinkedIn →Further reading
UNSPSC Classification Accuracy: What 90–95% Actually Means
What does 90–95% automated UNSPSC classification accuracy mean in practice? How is it measured, what does the remaining 5–10% look like, and how does accuracy improve over time?
Read more →ProcurementUNSPSC for MRO: Classifying Maintenance, Repair, and Operations Spend
MRO procurement is the hardest category to classify consistently. Here is why UNSPSC works for maintenance spend, and what it takes to get clean MRO data from SAP or Maximo.
Read more →

