"We cannot start using AI until we get our data sorted."
I keep hearing this.
It sounds responsible. It sounds like good governance. In some companies it has become the respectable reason to delay every useful AI project for another year.
There is only one problem.
Your data will never be perfect.
You may improve it. You should improve it. You may move it into a better platform, agree common definitions, remove duplicates, add ownership and repair years of neglected integrations.
But it will never be finished.
The company will change. Customers will change. Systems will change. Somebody will buy another piece of software. A field that was optional will become important. A merger, product launch or regulation will alter what the data is meant to describe.
If perfect data is the starting line for AI, you may never start.
The question is not whether the data is perfect. The question is whether it is good enough, understood enough and controlled enough for this particular task.
We have already had decades to sort the data
Many organisations have been trying to fix their data for years.
They have bought databases, data warehouses, master-data tools, customer platforms, data lakes and new reporting systems. The latest programme may involve Snowflake, Databricks or another very capable platform.
That work can be valuable. Keep doing it.
But do not pretend that one more transformation programme will create a permanent moment in which every record is complete, every definition is agreed and every system tells the same story.
Data quality is relative to a purpose. The UK Government's Data Quality Framework describes quality as fitness for purpose and recommends targeting improvements where they add most value. It also treats quality as something managed throughout the data lifecycle, not a cleaning exercise that ends.
That gives us a much more useful question:
What does this data need to be good enough to do?
A chat box is useful. It is not the whole operating model
While the data programme runs, companies often give people access to a basic online chat tool and call that their AI strategy.
Those tools can be useful for drafting, summarising public information and thinking through a problem. But they do not, by themselves, teach the organisation how to perform real work safely.
The more interesting capability is in the harness around the model: the workspace, files, tools, permissions, instructions, tests, logs and repeatable workflows. A harness such as Codex or Claude's work environment can help a person inspect information, produce an artefact, run checks and leave evidence of what happened.
That does not mean giving it access to everything.
It means giving it the smallest useful set of data and authority for a clearly defined job.
Run two programmes at the same time
I would not cancel the long-term data programme. I would run two lanes.
Lane one improves the estate.
It establishes ownership, shared definitions, lineage, platform architecture, security, retention, quality monitoring and reusable data products. It may take one or two years. In reality, it will probably become an ongoing capability rather than a project with a neat finish line.
Lane two delivers bounded value now.
A small team takes one business outcome, finds the minimum data needed, checks it, records what is unreliable and builds a safe route through it. It uses copied or read-only information first. A person reviews the result. The workflow cannot silently write back into the system of record.
The second lane does not compete with the first. It tells the first lane which problems are actually worth solving.
Build a small data delivery team
Call it a tiger team, a data delivery team or a SWAT team if that language works inside your company.
Its job is not to clean the whole organisation. Its job is to make one useful outcome possible without hiding the weaknesses in the source information.
- Name the business outcome. What decision, document, service or piece of work are we trying to improve?
- Select the minimum useful data. Do not begin by connecting every system. Start with the few sources the task genuinely needs.
- Name the owner and the meaning. Who understands the fields, the exceptions, the dates and the way this information is actually created?
- Profile the weaknesses. Record missing values, stale records, conflicting definitions, duplicates and known blind spots.
- Set task-specific quality rules. Decide what must be complete, current or accurate for this use case, and what uncertainty can remain visible.
- Use a bounded copy or read-only connection. Keep the first route away from live write authority, payments, publication and irreversible actions.
- Test the result against reality. Use known examples, awkward edge cases and competent human review. Measure corrections and failures, not just attractive demos.
- Feed the defects back. Every recurring data problem becomes evidence for the long-term data programme.
That is not doing it half-heartedly.
It is doing the smallest complete piece of work that can teach the company something real.
Imperfect data is not permission to guess
I want to be careful here.
"Start with imperfect data" must not become "dump everything into a model and hope".
The more consequential the outcome, the stronger the evidence and controls need to be. A rough internal classification task is not the same as setting a price, deciding somebody's eligibility, giving clinical advice or moving money.
For each use case, I would want a short receipt:
- the intended outcome;
- the source and effective date of the data;
- what is known to be incomplete or unreliable;
- what the agent can read, create or change;
- how the result is checked;
- who is accountable for accepting it;
- how an error is stopped and corrected;
- and what evidence is kept.
NIST's voluntary AI Risk Management Framework is useful here because it treats governance, context, measurement and ongoing management as connected work. You do not solve risk once and then forget it. You keep measuring the system in the context in which it is being used.
Use AI to discover the data work
There is another reason not to wait.
Until people try to improve a real task, they often do not know which data matters.
A central programme may spend months standardising fields that nobody uses while the operations team still cannot tell whether a customer has received the thing they paid for.
A live, bounded use case exposes that quickly.
It finds the missing identifier, the date that means three different things, the spreadsheet maintained by one person, the customer name that does not match finance and the status field that everybody ignores.
The AI project is not sitting on top of the data programme. It is producing evidence for it.
Do not wait to start learning
I worry that some organisations are using data quality as a postponement strategy.
They will spend another year buying platforms, drawing architecture and promising that everything will be ready after the migration. Meanwhile, their people will not learn how to describe work to an agent, inspect its reasoning, test its output, control its permissions or decide where it should stop.
Those are organisational capabilities. They take practice.
So yes, improve your data.
Fund the long-term programme. Establish ownership. Repair the systems. Remove the duplication. Build the platform that will serve the company for years.
And at the same time, choose one useful problem. Find the minimum data. Make its weaknesses visible. Put a safe harness around the work. Review the outcome. Learn from what breaks.
Do not wait for perfect data.
Build a safe path through the data you have.
Related reading
- If It Made You Different, Start AI At The Edge
- How Do We Get Our Teams To Use AI?
- When A Proof Of Concept Is Really Research
- Start With A Safe Folder, Not A Live System
- The AI Harness Is Becoming The Operating System
Sources and notes
- UK Government Data Quality Framework
- UK Government Data Quality Framework: guidance
- NIST Artificial Intelligence Risk Management Framework
This is a practical operating argument, not legal, security or data-governance advice. Controls should reflect the data, the people affected and the consequences of getting the task wrong.
