Data Isn't Scarce, Imagination Is
Key point
The AI data shortage isn't a problem of quantity, but a limitation of imagination that fails to capture new data domains including interaction and causality.
Details
The AI industry currently worries that training data is scarce, but this isn't because the absolute amount of data is insufficient—it's because the scope of data we imagine is too narrow. Existing data such as text, code, and images is only a small fraction of the data we could imagine.
The decisive problem revealed in attempts to automate SRE (Site Reliability Engineering) wasn't model performance, but the absence of end-to-end, long-horizon data. What matters isn't simple fragmentary data, but data that captures the entire cause and effect chain from the occurrence of an incident to its resolution.
Recent research points to the possibility of new kinds of data.
- Thinking Machines Lab: Enables models to acquire proactive capabilities on their own through appropriate interaction data.
- Neural Computers (Meta AI & KAUST): Converts a computer's behavior in CLI and GUI environments into a modelable data stream.
Ultimately, every human-machine interface creates a new class of data. As new interfaces emerge—voice, visual interaction, biosignals, and more—the data market will transform from a fixed pie into a continuously expanding map.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.