The AI Data Problem Moved Downstream
Last year I wrote that robotics had a data problem language models were lucky enough to avoid. Text, images, and code had accumulated for decades before anyone decided to train models on them. Robots had no comparable internet of physical experience.
That distinction is starting to look less clean. Recent systems have trained on enormous collections of human video. One company reported using more than a million hours. Another turned roughly 1,900 hours of first-person human footage into more than 18,000 hours of robot-format training data. The interesting part is not that robotics suddenly found its missing dataset. It is that every one of these systems needs substantial machinery to turn human experience into something a robot can actually use.
My last two articles focused on data companies already have. This article looks at a different part of the problem: information can exist without being usable by the system.
If you find this useful, consider becoming a paid supporter 🙏
The Web Is Becoming a Market With Rules
Something related is happening on the digital side, but from the opposite direction. The information is already there, but access to it is becoming more conditional.
An academic publisher recently made selected books available for retrieval by AI research products while explicitly withholding model-training rights. This distinction matters. “Use this for AI” is starting to break into separate permissions for training, retrieval, inference, and perhaps eventually other uses.
The technical infrastructure is moving in the same direction. Website owners can increasingly distinguish between crawlers used for search, training, and agents. One closed-beta system even lets a crawler receive an HTTP 402 response with a price for access. The market is still immature, but the machinery for negotiating access already exists.

Trust is getting harder too. More than a third of webpages published since ChatGPT launched now show significant signs of AI authorship. More troubling are sites producing hundreds of thousands of pages specifically designed to become evidence for AI recommendation systems.
I recently wrote about the risks attached to where a dataset came from before deployment. The newer problem starts after deployment.
When Retrieved Data Can Act
A training dataset can be inventoried, licensed, audited, and removed. Nobody can pre-audit every web page or document an agent will encounter tomorrow. Some of that information will come from sources the company does not control.
That difference becomes important once the system can act. Researchers recently found documentation across more than 100 websites that led coding agents to install software whose ownership could not be established. Another study showed how a malicious tool retrieved from a shared library could be copied by a self-evolving coding agent, stored as a new skill, and propagated further.

This is why I think retrieval is becoming a security boundary rather than merely a search-quality problem. Researchers are already experimenting with source trust scores, cross-document consistency checks, and sanitization at that boundary.
A search system can turn bad information into a bad answer. An agent can turn it into an action. For agents, provenance starts to look less like metadata and more like software supply-chain security. You need to know not only what the system read, but what that information was allowed to cause.
Robotics Found a Pipeline, Not a Corpus
Human video looks like an obvious answer to robotics’ data shortage until you look closely at what the new systems actually do.
One pipeline retargets human actions to robot bodies, inserts robot arms into the video, filters the results, and then combines the synthetic data with real robot demonstrations. Another learns from human and robot video, then grounds what it learned in robot trajectories before adapting it to the particular machine doing the work. Human video is useful raw material. It is not robot experience.
The infrastructure behind the largest effort makes the point even more clearly. Processing throughput had to increase from 14,000 to 440,000 episode-hours per week. Turning the raw data into something the training system could use once took about 48 hours and now takes under a minute. At this scale, the work includes transformation, quality checks, metadata, storage, caching, and continuous preparation for training.

All this changes how I think about the robotics data problem. Robotics did not discover an internet of physical experience. It is building the machinery that can turn human experience into one.
The distinction matters because the machinery may become as economically important as the robots themselves. Collecting, cleaning, translating, grounding, and continually refreshing physical experience is starting to look like a separate layer of the robotics stack.
Useful Data Is the Output
The common thread is not that digital agents and robots use the same kind of data. They obviously do not. It is that useful data increasingly sits at the end of a process rather than at the beginning. The technology is new, but the data problem is familiar.
For digital agents, that process decides what outside information can enter and what the system is allowed to do with it. For robots, it turns human video and demonstrations into data the machine can learn from. The mechanisms are different, but in both cases useful data has to be made from the raw material.

That leads to a management question I suspect many companies cannot answer cleanly: who owns the work of turning raw information and experience into something an AI system can safely use?
The Work in Between
Models are getting easier to buy. The web may be full of information that an agent still cannot safely use. There is plenty of human video. The hard part is turning it into something a robot can use. In both cases, the data problem has moved from finding raw material to doing the work required to make it usable.

Beyond Transformers: A Different Way to Build Intelligence

Six Reasons I Think Open AI Models Will Win

