Quick answer: A data cleaning AI project fails most often not because the model is weak, but because nobody manages it the way they would manage a team doing the same work. That means starting small enough to build trust before scaling, naming a specific person accountable for the outcome, and expecting the same reporting, timelines, and contact point from the software that you would expect from a team of people doing this work for you.
On this page: Start small before you scale · Treat the software like a team · What accountability requires from software · Why pilots and validation mattered · FAQ
Start small before you scale
Recent industry research puts a hard number on a pattern we see directly in client engagements: over 80 percent of AI projects fail to deliver, and the average organisation scrapped nearly half of its AI proofs of concept in 2025 before they reached production. The recurring reason is not that the underlying models are not capable enough. It is unclear definitions of success, weak data foundations, and no governance framework attached to the rollout.
One FM client we worked with recently ran their data cleaning project through pilot phases rather than a single big bang rollout across every site and every category at once. The project ran narrower in its first phase than the eventual scope, tested against a validated supplier list before expanding, and only then rolled out across the rest of the business. At the project review afterward, the pilot phases and the validated supplier list, not the underlying AI itself, were what the client identified as the reason the project held together. A model that works well in a controlled first phase and is then handed the whole company's data at once has no chance to surface its edge cases before they are expensive.
Treat the software like a team
If we hire a team to do a piece of work, we do not just hand over the brief and disappear until the deadline. We expect to know who is accountable, how the work is progressing, and where to raise a concern. Somewhere in the shift to AI tools, that expectation quietly drops away, and a data cleaning project gets treated as a system you configure once rather than a working relationship you manage continuously.
The fix is not complicated, and it does not require new AI capability. It requires applying the same expectations. Name the software, in the practical sense of giving the project a clear owner and a clear point of contact, the way you would name the lead on a team doing this work. Define, explicitly, who is responsible for what the system produces, because "the AI did it" is not an answer anyone accepts from a person, and it should not be accepted from a system either. These are not new disciplines invented for AI. They are the same disciplines any outsourced or in-house team is already held to, applied to a different kind of worker.
What accountability requires from software
Some expectations transfer directly. A contact point, a named person who can answer a question about the work, is needed whether the work is done by a team or a system, and a data cleaning AI project without one is a project nobody can escalate a problem to. Other expectations surface differently but still apply. If a team member did this work for you, you would expect them to update you on progress without having to ask each time. The equivalent for software is not a person checking in verbally, it is the system surfacing what has happened and what needs attention on a schedule you do not have to chase.
That difference matters because the failure mode without it is the same either way: nobody knows where to look until something is already wrong. Rather than a client repeatedly checking in without knowing when to look, it is far easier to manage a project that tells you, on its own schedule, what has changed and where your attention is needed first. A system that produces results with no visibility into confidence, no flag for what needs review, and no record of what changed is not more capable for lacking those things. It is simply unmanaged, and unmanaged work fails for the same reasons whether a person or a model produced it.
Why pilots and validation mattered
The specific lesson from that project, generalised beyond the one client, is that the two things that actually made the difference were structural, not technical. Pilot phases meant edge cases surfaced against a small, contained dataset where a wrong result was cheap to catch and correct, rather than against the full company dataset where the same wrong result would have propagated into a board report before anyone noticed. A validated supplier list meant the classification work started from a foundation that was already checked, rather than compounding uncertainty in the supplier data with uncertainty in the classification on top of it.
Neither of those things required a better model. Both required someone managing the project with the same discipline they would apply to a team of people: sequencing the work, checking the foundation before building on it, and keeping a named person accountable for what came out the other end. That is the actual argument for treating an AI data project as a managed engagement rather than a tool deployment. The project succeeds or fails on the same structural decisions either way. The AI just changes who, or what, is doing the labour inside that structure.
Frequently asked questions
Why do most AI data cleaning projects fail?
The recurring causes are structural rather than technical: unclear definitions of success, a weak data foundation to start from, and no governance framework around the rollout. Over 80 percent of AI projects fail to deliver on this pattern, and the failure typically traces back to how the project was managed, not to a limitation in the underlying model.
How can I run a successful AI data cleaning project?
Running a successful data cleaning project with AI ironically hinges on the people involved. Before starting to work with an AI data cleaning company, always ask them how they structure their projects and whether you will have a dedicated success manager. A human point of contact means there is someone overseeing the quality of the work, which is needed to ensure the automations keep delivering in the long run.
What does it mean to treat AI software like a team you manage?
It means applying the same basic expectations you would apply to a team of people doing the same work: a named person accountable for the outcome, a clear point of contact, regular visibility into progress, and a defined process for escalating a problem. These are project management disciplines, not AI capabilities, and skipping them is what turns a capable model into an unmanaged project.
Why did pilot phases matter more than the AI model itself in this project?
Pilot phases let edge cases surface against a small, contained dataset where a wrong result is cheap to catch, rather than against the full company dataset where the same error would have reached a report before anyone noticed. That sequencing decision, not a difference in model capability, was what the client identified as the reason the project held together.
What is a validated supplier list, and why does it matter for a data cleaning project?
A validated supplier list is a checked, deduplicated reference of supplier records confirmed accurate before classification work builds on top of them. Starting classification from an unvalidated supplier list compounds two sources of uncertainty, supplier identity and category assignment, instead of resolving one before tackling the other.
How does Pearstop apply project management discipline to AI data projects?
Pearstop runs data cleaning and classification engagements with a dedicated contact person responsible for success. Projects are run through pilot phases against a validated supplier list before scaling to full volume, and assigns a named point of contact accountable for the outcome, with progress and confidence flagged on a set schedule rather than left for the client to chase. The step-by-step approach allows to check properly while volume is still small and tweak before scaling.
Is a data cleaning AI project ever safe to roll out across a whole company at once?
It is rarely the right sequencing choice, even when the underlying technology can technically handle the full volume. A phased rollout against a smaller, validated dataset surfaces edge cases while they are still cheap to fix, and the cost of that phasing is small compared to the cost of a wrong classification reaching a board report from an unvalidated, company-wide first pass.
Frequently asked questions
Why do most AI data cleaning projects fail?
The recurring causes are structural rather than technical: unclear definitions of success, a weak data foundation to start from, and no governance framework around the rollout. Over 80 percent of AI projects fail to deliver on this pattern, and the failure typically traces back to how the project was managed, not to a limitation in the underlying model.
How can I run a successful AI data cleaning project?
Running a successful data cleaning project with AI ironically hinges on the people involved. Before starting to work with an AI data cleaning company, always ask them how they structure their projects and whether you will have a dedicated success manager. A human point of contact means there is someone overseeing the quality of the work, which is needed to ensure the automations keep delivering in the long run.
What does it mean to treat AI software like a team you manage?
It means applying the same basic expectations you would apply to a team of people doing the same work: a named person accountable for the outcome, a clear point of contact, regular visibility into progress, and a defined process for escalating a problem. These are project management disciplines, not AI capabilities, and skipping them is what turns a capable model into an unmanaged project.
Why did pilot phases matter more than the AI model itself in this project?
Pilot phases let edge cases surface against a small, contained dataset where a wrong result is cheap to catch, rather than against the full company dataset where the same error would have reached a report before anyone noticed. That sequencing decision, not a difference in model capability, was what the client identified as the reason the project held together.
What is a validated supplier list, and why does it matter for a data cleaning project?
A validated supplier list is a checked, deduplicated reference of supplier records confirmed accurate before classification work builds on top of them. Starting classification from an unvalidated supplier list compounds two sources of uncertainty, supplier identity and category assignment, instead of resolving one before tackling the other.
How does Pearstop apply project management discipline to AI data projects?
Pearstop runs data cleaning and classification engagements with a dedicated contact person responsible for success. Projects are run through pilot phases against a validated supplier list before scaling to full volume, and assigns a named point of contact accountable for the outcome, with progress and confidence flagged on a set schedule rather than left for the client to chase. The step-by-step approach allows to check properly while volume is still small and tweak before scaling.
Is a data cleaning AI project ever safe to roll out across a whole company at once?
It is rarely the right sequencing choice, even when the underlying technology can technically handle the full volume. A phased rollout against a smaller, validated dataset surfaces edge cases while they are still cheap to fix, and the cost of that phasing is small compared to the cost of a wrong classification reaching a board report from an unvalidated, company-wide first pass.

Stephanie Wiechers
CEO & Co-founder, Pearstop
Stephanie leads Pearstop's go-to-market and strategic direction. She works directly with procurement and FM leaders across Europe to understand how data quality affects margins, contracts, and AI readiness.
LinkedIn →Further reading
When manufacturer and type fields swap in your spend data
A silent field swap between manufacturer and type corrupts classification without a single error message. Here is how it happens and how it gets caught.
Read more →ProcurementUNSPSC vs eCl@ss: when your ERP and reporting disagree
UNSPSC and eCl@ss answer different questions. Here is how to pick when your ERP was built for one and head office reporting expects the other.
Read more →

