TechIntelligenceJuly 20, 2026 · 5 min read

Data governance: the invisible prerequisite before automating anything

What the RAND Corporation study reveals about the causes of AI project failure, and how to assess your data governance before automating.

Upleo

Upleo

More than 80% of enterprise AI projects fail to deliver the expected value — a rate twice as high as for traditional IT projects, according to a RAND Corporation study conducted among 65 data scientists and engineers. Among the five root causes identified, only one is purely technical. Insufficient, inaccessible or poor-quality data is a direct, documented cause — and it precedes, in the order of operations, any choice of model or tool.

Why data, not the model, is the leading cause of failure

What the RAND study shows (2024)

The study The Root Causes of Failure for Artificial Intelligence Projects, published by the RAND Corporation, synthesized the experiences of 65 data scientists and engineers to identify five root causes of AI project failure: a poorly understood definition of the problem to solve, insufficient or inaccessible training data, a technology approach chosen before the business problem, inadequate deployment infrastructure, and a problem too complex for the current state of the art. Of these five causes, only one — the last one — is genuinely a technological limitation. The other four, including insufficient data, are organizational and procedural.

The direct cost of poor data quality

Gartner estimates the average cost of poor data quality at around $12.9 million per year per organization. This figure isn't specific to AI projects — it affects every decision made on degraded data — but it directly illustrates why an automation project built on that same data structurally inherits the problem before its first deployment.

The 5 root causes of AI project failure according to RANDA diagram listing the five root causes of AI project failure identified by the RAND Corporation: poorly defined problem, insufficient data, technology chosen before the problem, inadequate infrastructure, problem too complex. Only one of these five causes is purely technical.Poorly defined problemInsufficient or inaccessible dataTechnology chosen before the needInadequate deployment infrastructureProblem too complexOnly causethat is truly technical4 organizational causes out of 5 — including data

The link to the already-identified organizational factor

Data governance as a component of organizational readinessA diagram showing that organizational readiness, already identified as a key success factor for enterprise AI, translates concretely into data governance among other components.Organizational readiness(MIT NANDA, McKinsey)DatagovernanceInternal cultureand skillsProcessredesign

A concrete component of the "learning gap"

We already detailed, in our analysis of generative AI's real value in the enterprise, the concept of the "learning gap" identified by MIT NANDA — an organization's inability to integrate an AI tool into existing workflows — as well as the weight of organizational readiness measured by McKinsey. Data governance is one of the most concrete and measurable components of that organizational readiness: it isn't a separate topic, but one of the precise mechanisms through which that organizational factor translates — or doesn't — into results.

What the Informatica survey confirms

A survey conducted by Informatica among chief data officers (CDO Insights, 2025) found that 43% of organizations surveyed cite data quality and readiness as the top obstacle to the success of their AI projects — ahead of budget constraints or a lack of technical skills.

The 4 dimensions of data governance ready for automation

The 4 dimensions of data governance ready for automationA structural diagram showing the four dimensions of data governance ready for automation: ownership, quality, accessibility, and compliance.OwnershipWho isresponsibleQualityAccuracy,freshnessAccessibilitySilos vsunified architectureComplianceRetention,traceabilityAll four dimensions must be addressed together, not sequentially

Ownership: who is responsible for which data

Without a clearly identified owner for each data category, no one is genuinely in charge of its quality, its updates, or its correction when an error is detected. This is one of the most frequently cited gaps in the literature on AI project failure.

Quality: accuracy, completeness, freshness

Incomplete, duplicated or outdated data produces a model that faithfully reproduces those flaws, at scale. Quality isn't a one-time state achieved once, but a property that must be maintained over time as the data evolves.

Accessibility: silos vs unified architecture

Good-quality data scattered across siloed systems that don't talk to each other remains, in practice, just as unusable as poor-quality data. Technical accessibility — not just the data's existence — directly determines whether it can be mobilized in an automation project.

Compliance: retention, privacy, traceability

Regulatory compliance (retention periods, consent management, traceability of processing) isn't a peripheral constraint: an automation project that ignores these rules upfront risks having to be rebuilt once the problem is discovered, usually after deployment rather than before.

Worked example: assessing your readiness before launching a project

Data governance readiness audit gridA diagram comparing an initial readiness score of 45 out of 100, split across ownership, quality, accessibility and compliance, to a score of 85 out of 100 after structured governance work.Before governance work45 / 100Breakdown: Ownership 10, Quality 12,Accessibility 8, Compliance 15After governance work85 / 100Breakdown: Ownership 22, Quality 20,Accessibility 20, Compliance 23Each dimension scored out of 25 points,total out of 100

A simple self-assessment grid, scored out of 100 points and evenly split across the four dimensions above, makes it possible to objectively locate where an organization stands before launching a project:

  • Ownership (25 points): before any governance work, a typical organization where data responsibility has never been formally assigned scores around 10/25.
  • Quality (25 points): with unaddressed duplicates and incomplete fields, the initial score generally sits around 12/25.
  • Accessibility (25 points): data scattered across several systems with no unified architecture typically scores 8/25.
  • Compliance (25 points): basic regulatory compliance without documented traceability of processing reaches around 15/25.

That's an initial total of 45 points out of 100. After structured work to assign owners, clean up duplicates, set up centralized access and document processing, these same dimensions typically reach 22, 20, 20 and 23 points respectively, for a total of 85 points out of 100 — a level of readiness that fundamentally changes the probability of success for the planned automation project, consistent with the failure causes identified by the RAND Corporation.

What this concretely means before launching an AI project

The correct order: governance first, tooling secondA diagram showing the correct order for an automation project: first assess and govern the data, then only choose and configure the tool, rather than the reverse.1. Assess and governthe available data2. Choose and configurethe automation toolThe most common mistake: reversing this order

The correct order: governance first, tooling second

The same trap already identified in our article on B2B marketing automation — configuring the tool before mapping the actual need — shows up here in a different form: launching an AI project before assessing the readiness of the underlying data. In both cases, the mistake is the same: treating tooling as the starting point rather than the final step of a broader preparation.

Who should be involved: not just IT

Leaving data governance solely to the IT department is a common mistake identified in the literature: IT generally owns infrastructure and technical accessibility, but the definition of what counts as quality data in a given business context, along with the responsibility for keeping it up to date, needs to be owned by the business teams themselves.

Sources

  • RAND Corporation (2024). The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed. rand.org
  • Gartner (2024-2025). Estimate of the average cost of poor data quality, cited across several industry analyses.
  • Informatica (2025). CDO Insights Survey.

Is your automation project resting on data that no one formally owns? Let's talk about it, starting from an honest diagnostic of your actual readiness.

Frequently asked questions

A question about this article?

Let's talk about your context and how we can help.