Security

80% of Enterprise Data Is Unstructured, and 81% of Employees Feed It to AI Tools Nobody Approved

Cyberhaven, UpGuard, and a review of the underlying unstructured-data research all describe the same shape: a data estate that is mostly unstructured, growing faster than governance can track, and now feeding directly into AI tools most security teams cannot see.

September 15, 2026 · Security
A darkened office desk with dual monitors displaying a security operations dashboard of file counts, access maps, and data flow charts, no people in frame

Key Takeaways

  • Cyberhaven found 39.7% of AI interactions in 2026 involved sensitive data, with employees pasting it into an AI tool roughly once every three days.
  • UpGuard's 2026 Enterprise AI Security Index found 81% of employees routinely use AI tools their employer never approved, and more than 15% of business-critical files already sit exposed to misconfiguration or oversharing.
  • Unstructured data, the contracts, chat logs, and spreadsheets that AI tools are built to read, already makes up roughly 80% of enterprise data and is growing 55 to 65% a year.
  • Personal accounts carry the deepest exposure: Cyberhaven found 58.2% of Claude use and 60.9% of Perplexity use inside the organisations it studied ran through accounts IT cannot monitor.

Most data security programs were built around a database. They have an owner, a schema, and a fixed set of columns a compliance team can point to and say this is what we are protecting. The actual bulk of enterprise information was never that tidy. It lives in contracts, meeting transcripts, spreadsheets, and chat threads, and a growing body of 2026 research suggests that unstructured mass is exactly what today's AI tools are best at reading, and exactly what nobody has finished governing before employees started feeding it in.

The Blind Spot Was Already 80% of the Estate

Unstructured data already accounts for roughly 80% of enterprise information, according to Lepide research cited in a recent review of the underlying data security posture management research, and that share is expanding 55 to 65% a year as file storage, collaboration tools, and messaging platforms multiply faster than any team can classify them. Legacy discovery tools built for structured databases were never designed for that scale. They lose visibility once an environment crosses roughly 100 terabytes, and by the time an enterprise reaches petabyte-scale storage, most of that content has never been scanned for what it actually contains.

That gap matters more now than it did two years ago, because the tools employees reach for to make sense of a messy file share are AI tools, and AI tools do not care whether a document was ever classified before it landed in a prompt.

Shadow AI Is Not an Edge Case, It Is the Default

UpGuard's 2026 Enterprise AI Security Index found that 81% of employees routinely use AI tools their employer never approved, a figure high enough that shadow AI is no longer the exception a policy needs to catch, it is closer to the baseline behaviour a policy has to assume. The same index found that more than 15% of business-critical enterprise files already face active exposure risk through misconfiguration or oversharing, meaning the content those unapproved tools are most likely to touch was frequently exposed before an AI tool ever entered the picture.

Put the two findings together and the sequence looks less like a single new risk and more like two existing gaps compounding each other. Files were already overshared. Employees are already routing around whatever approval process exists. An AI tool sitting between the two does not create either problem, it just gives both a much faster way to surface something a security team was never told to look for.

Employees Already Know Which Accounts to Use

Cyberhaven's 2026 analysis found that 39.7% of all AI interactions it tracked involved sensitive data, and that employees input sensitive material into an AI tool on average once every three days. The behaviour is not evenly spread across tools. Cyberhaven found 58.2% of Claude usage and 60.9% of Perplexity usage inside the organisations it studied ran through personal accounts rather than enterprise-managed ones, compared with 32.3% for ChatGPT and 24.9% for Gemini, a gap that tracks closely with which tools have the weakest enterprise admin controls rather than which ones employees trust least.

None of this reads as malicious. It reads as employees solving a real problem, finding the answer buried in a contract or a spreadsheet, the fastest way they know how, through whichever account happens to be already logged in. The account choice is rational for the employee and invisible to the security team, which is exactly the combination that turns an individual shortcut into an organisation-wide blind spot.

The number worth sitting with is not 39.7%, the share of AI interactions already touching sensitive data. It is 80, the share of the enterprise that was unstructured, ungoverned, and largely unwatched long before AI gave it anywhere new to go.

Share

More in Security

All Resources →