Cyberhaven, UpGuard, and a review of the underlying unstructured-data research all describe the same shape: a data estate that is mostly unstructured, growing faster than governance can track, and now feeding directly into AI tools most security teams cannot see.
Key Takeaways
Most data security programs were built around a database. They have an owner, a schema, and a fixed set of columns a compliance team can point to and say this is what we are protecting. The actual bulk of enterprise information was never that tidy. It lives in contracts, meeting transcripts, spreadsheets, and chat threads, and a growing body of 2026 research suggests that unstructured mass is exactly what today's AI tools are best at reading, and exactly what nobody has finished governing before employees started feeding it in.
Unstructured data already accounts for roughly 80% of enterprise information, according to Lepide research cited in a recent review of the underlying data security posture management research, and that share is expanding 55 to 65% a year as file storage, collaboration tools, and messaging platforms multiply faster than any team can classify them. Legacy discovery tools built for structured databases were never designed for that scale. They lose visibility once an environment crosses roughly 100 terabytes, and by the time an enterprise reaches petabyte-scale storage, most of that content has never been scanned for what it actually contains.
That gap matters more now than it did two years ago, because the tools employees reach for to make sense of a messy file share are AI tools, and AI tools do not care whether a document was ever classified before it landed in a prompt.
UpGuard's 2026 Enterprise AI Security Index found that 81% of employees routinely use AI tools their employer never approved, a figure high enough that shadow AI is no longer the exception a policy needs to catch, it is closer to the baseline behaviour a policy has to assume. The same index found that more than 15% of business-critical enterprise files already face active exposure risk through misconfiguration or oversharing, meaning the content those unapproved tools are most likely to touch was frequently exposed before an AI tool ever entered the picture.
Put the two findings together and the sequence looks less like a single new risk and more like two existing gaps compounding each other. Files were already overshared. Employees are already routing around whatever approval process exists. An AI tool sitting between the two does not create either problem, it just gives both a much faster way to surface something a security team was never told to look for.
Cyberhaven's 2026 analysis found that 39.7% of all AI interactions it tracked involved sensitive data, and that employees input sensitive material into an AI tool on average once every three days. The behaviour is not evenly spread across tools. Cyberhaven found 58.2% of Claude usage and 60.9% of Perplexity usage inside the organisations it studied ran through personal accounts rather than enterprise-managed ones, compared with 32.3% for ChatGPT and 24.9% for Gemini, a gap that tracks closely with which tools have the weakest enterprise admin controls rather than which ones employees trust least.
None of this reads as malicious. It reads as employees solving a real problem, finding the answer buried in a contract or a spreadsheet, the fastest way they know how, through whichever account happens to be already logged in. The account choice is rational for the employee and invisible to the security team, which is exactly the combination that turns an individual shortcut into an organisation-wide blind spot.
The number worth sitting with is not 39.7%, the share of AI interactions already touching sensitive data. It is 80, the share of the enterprise that was unstructured, ungoverned, and largely unwatched long before AI gave it anywhere new to go.
Guide
With enterprise content already 80% unstructured and growing 55 to 65% a year, this guide shows how intelligent content management surfaces what is actually in that pile before an unapproved AI tool gets there first.
Download
Whitepaper
UpGuard traced over 15% of business-critical files to active misconfiguration or oversharing, and this whitepaper shows how mapping information flow catches that exposure before it reaches a tool nobody approved it for.
Download
Whitepaper
Cyberhaven found 39.7% of AI interactions now involve sensitive data, and this whitepaper covers the training-data and lifecycle controls needed once that data is already inside a model nobody vetted.
Download
Cisco Talos, CrowdStrike, and Verizon all point to the same shift.
Only 9% of organisations can scan unstructured data in real time.
Unit 42 found encryption present in just 78% of 2025 extortion cases.