Data governance: the invisible prerequisite before automating anything
What the RAND Corporation study reveals about the causes of AI project failure, and how to assess your data governance before automating.
Upleo
More than 80% of enterprise AI projects fail to deliver the expected value — a rate twice as high as for traditional IT projects, according to a RAND Corporation study conducted among 65 data scientists and engineers. Among the five root causes identified, only one is purely technical. Insufficient, inaccessible or poor-quality data is a direct, documented cause — and it precedes, in the order of operations, any choice of model or tool.
Why data, not the model, is the leading cause of failure
What the RAND study shows (2024)
The study The Root Causes of Failure for Artificial Intelligence Projects, published by the RAND Corporation, synthesized the experiences of 65 data scientists and engineers to identify five root causes of AI project failure: a poorly understood definition of the problem to solve, insufficient or inaccessible training data, a technology approach chosen before the business problem, inadequate deployment infrastructure, and a problem too complex for the current state of the art. Of these five causes, only one — the last one — is genuinely a technological limitation. The other four, including insufficient data, are organizational and procedural.
The direct cost of poor data quality
Gartner estimates the average cost of poor data quality at around $12.9 million per year per organization. This figure isn't specific to AI projects — it affects every decision made on degraded data — but it directly illustrates why an automation project built on that same data structurally inherits the problem before its first deployment.
The link to the already-identified organizational factor
A concrete component of the "learning gap"
We already detailed, in our analysis of generative AI's real value in the enterprise, the concept of the "learning gap" identified by MIT NANDA — an organization's inability to integrate an AI tool into existing workflows — as well as the weight of organizational readiness measured by McKinsey. Data governance is one of the most concrete and measurable components of that organizational readiness: it isn't a separate topic, but one of the precise mechanisms through which that organizational factor translates — or doesn't — into results.
What the Informatica survey confirms
A survey conducted by Informatica among chief data officers (CDO Insights, 2025) found that 43% of organizations surveyed cite data quality and readiness as the top obstacle to the success of their AI projects — ahead of budget constraints or a lack of technical skills.
The 4 dimensions of data governance ready for automation
Ownership: who is responsible for which data
Without a clearly identified owner for each data category, no one is genuinely in charge of its quality, its updates, or its correction when an error is detected. This is one of the most frequently cited gaps in the literature on AI project failure.
Quality: accuracy, completeness, freshness
Incomplete, duplicated or outdated data produces a model that faithfully reproduces those flaws, at scale. Quality isn't a one-time state achieved once, but a property that must be maintained over time as the data evolves.
Accessibility: silos vs unified architecture
Good-quality data scattered across siloed systems that don't talk to each other remains, in practice, just as unusable as poor-quality data. Technical accessibility — not just the data's existence — directly determines whether it can be mobilized in an automation project.
Compliance: retention, privacy, traceability
Regulatory compliance (retention periods, consent management, traceability of processing) isn't a peripheral constraint: an automation project that ignores these rules upfront risks having to be rebuilt once the problem is discovered, usually after deployment rather than before.
Worked example: assessing your readiness before launching a project
A simple self-assessment grid, scored out of 100 points and evenly split across the four dimensions above, makes it possible to objectively locate where an organization stands before launching a project:
- Ownership (25 points): before any governance work, a typical organization where data responsibility has never been formally assigned scores around 10/25.
- Quality (25 points): with unaddressed duplicates and incomplete fields, the initial score generally sits around 12/25.
- Accessibility (25 points): data scattered across several systems with no unified architecture typically scores 8/25.
- Compliance (25 points): basic regulatory compliance without documented traceability of processing reaches around 15/25.
That's an initial total of 45 points out of 100. After structured work to assign owners, clean up duplicates, set up centralized access and document processing, these same dimensions typically reach 22, 20, 20 and 23 points respectively, for a total of 85 points out of 100 — a level of readiness that fundamentally changes the probability of success for the planned automation project, consistent with the failure causes identified by the RAND Corporation.
What this concretely means before launching an AI project
The correct order: governance first, tooling second
The same trap already identified in our article on B2B marketing automation — configuring the tool before mapping the actual need — shows up here in a different form: launching an AI project before assessing the readiness of the underlying data. In both cases, the mistake is the same: treating tooling as the starting point rather than the final step of a broader preparation.
Who should be involved: not just IT
Leaving data governance solely to the IT department is a common mistake identified in the literature: IT generally owns infrastructure and technical accessibility, but the definition of what counts as quality data in a given business context, along with the responsibility for keeping it up to date, needs to be owned by the business teams themselves.
Sources
- RAND Corporation (2024). The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed. rand.org
- Gartner (2024-2025). Estimate of the average cost of poor data quality, cited across several industry analyses.
- Informatica (2025). CDO Insights Survey.
Is your automation project resting on data that no one formally owns? Let's talk about it, starting from an honest diagnostic of your actual readiness.
Frequently asked questions
Related articles

Generative AI in the enterprise: where the value is real, and where it's still a gadget
What randomized studies (MIT, NBER, METR) reveal about where generative AI creates real enterprise value, and where it remains a gadget.

Lead scoring and qualification: what automation actually changes in commercial prioritization
How to build a reliable B2B account score: a 4-pillar method, a step-by-step worked example, and calibration pitfalls to avoid.
A question about this article?
Let's talk about your context and how we can help.