With the stock of high-quality public text on the internet nearly exhausted, OpenAI, Anthropic and other leading AI developers are quietly purchasing internal corporate communications — employee emails, Slack-style chats, video meeting transcripts and code-review histories — to train their next-generation models, according to a report by The Information published on August 13 citing about 20 industry sources.
Enterprise Data Becomes the New Oil for AI
The market for private business data has exploded in 2026. Bobby Samuels, chief executive of data brokerage firm Protege, said his company’s gross transaction volume jumped from $30 million in 2025 to at least $100 million this year as AI labs compete for datasets that capture real human problem-solving in professional settings.
Buyers are zeroing in on startups heading toward bankruptcy or acquisition. Warmly, an AI agent startup recently bought by HubSpot, turned down as many as four offers of up to $300,000 each for its internal meeting minutes and email archives even after signing its acquisition deal, according to people familiar with the matter. Such records — showing how developers debug code over email, how finance teams discuss quarterly results, or how sales teams demo products on video calls — are viewed as gold for training AI agents that must operate inside real companies.
Hitting the “Data Wall”
The rush into private data comes as researchers warn that AI models are running out of fresh public text. Epoch AI estimates the effective supply of quality human-written tokens on the open web at roughly 300 trillion, a pool that frontier models could fully consume between 2026 and 2032; under more aggressive training assumptions, that window shrinks to as early as 2027. While new text continues to appear online, its growth rate lags far behind the exponential expansion of AI training datasets.
From Books to Boardrooms: A Costly Shift
This pivot follows a controversial chapter in AI data sourcing. Anthropic’s internal “Project Panama” programme involved buying copyrighted books, removing their bindings for high-speed scanning, and then destroying the physical volumes. In July, a federal judge in San Francisco approved a $1.5 billion class-action settlement over the initiative. An unsealed planning document described the effort as an attempt “to destructively scan all the books in the world,” The Washington Post reported.
Now, with corporate communications in focus, anonymisation and privacy compliance have become major bottlenecks. Shub Sinha, chief executive of data processing company Integral, said his firm is running “a rigorous anonymization process to maintain data value while complying with privacy regulations” as it prepares datasets for AI buyers. The result: training costs, once dominated by compute expenses, are seeing data acquisition and cleaning emerge as a significant new line item for AI developers.
Privacy, Consent and the Indian Context
For India’s media and technology sector, the trend carries immediate relevance. Employee chats and emails often contain personally identifiable information, client details and commercially sensitive material. Any sale or reuse of such data must align with India’s Digital Personal Data Protection Act and comparable global regimes, raising complex questions about consent, purpose limitation and cross-border data transfers.
Journalists and editors should also note that as AI models rely more on opaque, privately sourced corporate data, the transparency of their training sources may decline. This could make it harder to audit or challenge AI-generated claims — a critical concern for newsrooms that increasingly use automated tools for summarisation, translation and initial drafting.
What Newsrooms and Creators Should Watch
Style homogenisation: AI assistants trained heavily on corporate communications may produce more uniform, business-like prose, potentially narrowing the range of writing styles available in automated tools.
Access inequality: As high-quality private datasets become expensive and proprietary, advanced AI capabilities may concentrate among well-funded organisations, affecting smaller Indian newsrooms and independent creators.
Policy preparedness: Media companies are advised to document their editorial workflows and style guides as potential assets, while establishing clear internal policies on whether — and under what conditions — their communications can be used for external AI training.
As one data industry executive put it to The Information: “The era of scraping the public web is over. The next battle is for the data inside companies.”









