Ohalo Logo

79% Of Enterprises Say Their Unstructured Data Is Ready For AI. Only 29% Know Where It Is

Companies are racing toward agentic AI. However, the unstructured data for AI sitting underneath those systems may be nowhere near ready.

New research from BARC, co-sponsored by Ohalo, exposes a striking gap between confidence and capability. While 79% of enterprises believe they can extract value from files, emails and documents for AI without compromising governance, only 29% fully know where their relevant data resides.

That gap matters. As businesses move AI from experiments into production, autonomous agents increasingly need access to corporate data to make decisions and take actions. If organisations cannot find, classify and govern that data, giving AI more autonomy could simply automate the underlying risk.

Put another way: the most expensive time to discover a data problem is after building an AI system on top of it.

BARC Research Exposes The Unstructured Data For AI Gap

The report, Harnessing Unstructured Data for AI Innovation, surveyed 225 enterprises across North America and Europe.

It is the first report in BARC’s four-report AI series to focus specifically on the unstructured data that powers AI and whether enterprises can govern it.

Co-sponsored by Ohalo, the enterprise file governance company, the research finds that confidence is moving faster than capability.

Enterprises believe they are ready for AI. Yet their answers about where data sits, what it contains and how they govern it tell a different story.

Consequently, organisations risk discovering serious data problems only after AI projects have started. At that point, exposure can grow while projects stall, budgets expand and expected returns disappear.

Why Unstructured Data For AI Creates A Security Problem

Industry estimates suggest roughly 80% of enterprise data exists in unstructured form. This includes files, emails, images and documents.

That information can contain years of institutional knowledge. It can also contain personal information, intellectual property, financial records, contracts, credentials and other sensitive material.

These are exactly the sources that many enterprise AI systems need for context.

However, the BARC research suggests that many organisations cannot adequately account for that information.

“Roughly two-thirds of AI adopters cannot effectively discover unstructured data, enforce governance policies on that data, or trace how they consume it,” said Kevin Petrie, VP of Research, BARC US. “This hurts their ability to feed agentic AI the deep context it needs to take safe actions and generate business value.”

For agentic AI, that problem becomes more serious.

Traditional AI systems may retrieve or summarise information. Agentic systems can go further. Depending on their permissions, they can interact with applications, retrieve records, trigger workflows and take actions.

That means bad data governance no longer affects only what an AI system knows. It can affect what the system actually does.

AI Projects Can Stall Before They Deliver Value

The problem often starts with good intentions.

A business approves an AI project and connects it to file shares containing years of useful material. That could include research, client information, contracts, product documentation and internal knowledge.

Then the questions begin.

What exactly is inside those files? Which documents contain sensitive information? Who should have access? Which versions are current? Can the information legally or safely enter an AI workflow?

Without reliable answers, the project can quickly hit a wall.

Teams may discover that metadata alone cannot provide enough visibility. As a result, someone must manually inspect and classify the information.

Progress slows. Costs rise. Meanwhile, executives begin asking when the AI investment will produce a return.

According to the report, 70% of enterprises say less than half of their data is discoverable for AI.

That is not a minor housekeeping problem. It is a fundamental AI readiness problem.

Manual Governance Cannot Keep Pace With Agentic AI

The research also highlights how heavily enterprises still depend on people to manage data governance.

“The governance holding all this together is human duct tape: 74% of enterprises still rely on manual effort or individual expertise. And only 29% fully know where their relevant data even resides. That’s where AI initiatives stall, because when it comes to data, you can’t govern what you can’t find,” said Kyle DuPont, CEO and Co-Founder, Ohalo.

The “human duct tape” description cuts to the heart of the issue.

Manual processes may work when employees access a manageable number of documents. They become far harder to defend when AI systems can search, combine and act on information across thousands or millions of files.

Moreover, an AI agent does not magically fix poor data governance. It can simply reach the problem faster.

Unstructured Data For AI Changes The Cybersecurity Equation

From a cybersecurity perspective, this is where the findings become particularly important.

Enterprises have spent years building controls around identities, applications, endpoints and networks. However, AI increasingly works across those traditional boundaries.

An authorised AI agent could access information that individual employees would previously have needed to locate manually. It could also combine information from multiple sources and expose relationships that were previously difficult to see.

Therefore, organisations need to think beyond whether an AI model itself is secure.

They also need to ask what information the AI can reach, who authorised that access and whether the organisation understands the data well enough to enforce those decisions.

Otherwise, a perfectly functioning AI system could become a remarkably efficient way to discover information the organisation did not realise it had.

That is not an AI hallucination. That is a data governance failure.

What The BARC Report Recommends

BARC recommends that organisations audit and catalogue their data before using it in production AI systems.

The report also points to closing governance gaps at the file level rather than relying only on perimeter controls.

That distinction matters because enterprise information rarely sits neatly inside one system. Files can spread across cloud storage, collaboration platforms, email systems, archives, local infrastructure and legacy environments.

Before an organisation gives AI access to that information, it needs to understand what exists, where it resides and what controls should apply.

Ohalo Targets The Unstructured Data Visibility Gap

Ohalo’s Data X-Ray platform addresses that visibility problem directly.

The platform connects to existing enterprise infrastructure and examines files including documents, PDFs, emails, contracts, scanned images and nested archives. It then classifies their contents and feeds those results into existing governance, security and AI tools.

Data X-Ray can run on-premises, in the cloud or in air-gapped environments without moving the underlying data.

The approach targets one of the central problems identified by the BARC research: organisations cannot effectively govern information they cannot first discover and understand.

Cyber News Live Take: AI Readiness Starts Below The Model

The AI industry spends enormous amounts of time discussing models, agents, copilots and the latest capabilities.

The less glamorous question is what those systems are being connected to.

An organisation can buy the best model available. It can deploy sophisticated AI security controls and build impressive agentic workflows. However, none of that fixes years of forgotten files, uncontrolled permissions and sensitive information scattered across enterprise infrastructure.

Agentic AI makes that problem more urgent because agents do not simply wait for employees to find information. They can actively search for it, consume it and potentially act on it.

The BARC findings therefore point to a basic security principle that should survive every AI hype cycle: know what you have before you give something else permission to use it.

If only 29% of enterprises fully know where their relevant data resides, the biggest obstacle to enterprise AI may not be the model at all.

It may be everything the organisation forgot was sitting underneath it.

Download The Full Unstructured Data For AI Report

The full BARC report examines how enterprises are preparing their unstructured information for AI and where governance, discovery and visibility gaps remain.

Download the full report here: https://hs.ohalo.co/barc-report.

Shopping Cart0

Cart

Login