Data governance
What to put in order in your data before bringing AI into the organization
When an AI project slips, the first assumption is that the problem is technical. In most cases we have seen, the problem was organizational: nobody knew who owned the data, which system held the correct version, or who was even permitted to see it.
Here is what we ask to have in order before building starts. This is not a year-long program - in a mid-size organization it is usually a matter of weeks.
1. One source of truth per entity
Every core entity - customer, order, product, employee - needs one system defined as the source of truth. If three systems hold a "customer address" and all three differ, an AI agent will answer with one of them and you will not know which. A better model does not solve that.
2. Permissions at the data layer, not the interface
The basic rule: an agent should see exactly what the user invoking it is permitted to see. If permissions today are implemented in the screen rather than in the data layer, an AI system will walk around them without anyone intending it. This is the most important item on the list and usually the most painful.
3. Classify data into three levels
You do not need an elaborate classification model. Three categories are enough: public, internal, sensitive. Each level gets one clear rule about what may be sent to an external model and what stays inside the organization. A policy you can explain in two minutes is a policy that gets followed.
4. Logging and traceability
Every question asked, every answer given, every action an agent took - recorded and reproducible. It is required for audit, but mostly it is what lets you understand why the system got something wrong and what it costs.
5. Retention and deletion policy
How long conversations with an AI system are kept, who can delete them, and what happens when a customer requests deletion. If you have obligations under privacy law or GDPR, they apply to this data too.
6. Defined ownership
Every data domain needs an owner by name and role - a person who decides who gets access and is accountable for quality. Without that, every access request becomes a two-week discussion at a management meeting.
7. Contract terms with vendors
Three clauses to check in any AI vendor contract: whether your data is used for training, where it is stored geographically, and what happens to it when the engagement ends. A serious vendor answers all three without hesitating.
And what you do not need
You do not need a two-year data warehouse program before starting with AI. That advice sounds responsible and in practice freezes organizations. The practical path is to put the data for the first use case in order only - source of truth, permissions, classification - and expand case by case.
The organizations that succeed with AI are not the ones whose data is perfect. They are the ones that know exactly which data is in order and which is not, and build accordingly.
Want to talk about what you just read?
If the article touched something you are dealing with, we are happy to have a short, no-obligation conversation.
Book a meeting