Training Data
In five years, 90% of enterprise tokens will come from work nobody asked for
The bet that one or two labs capture 95% of AI value is the one worth fading. Models being smart is not the bottleneck: real workflows need data connections, human-in-the-loop moments, idle agents waiting on delays, change management, and legacy systems nobody modernized, and no research organization wants to attack that list. Token subsidization from the labs is temporary because public markets will eventually apply the same laws of capitalism to everyone, and non-economic actors like Meta, China, and NVIDIA will happily run inference at 10% margin, which pushes value down to the application layer either way. The sharpest prediction: "in five years from now, I would bet, like, 90% of all tokens in the enterprise are things that a user never kicked off, and they just see a result." On why coding diffused instantly and nothing else has: code's entire value is text a person types at a computer, the audience debugs its own MCP errors instead of calling IT, and there is no equivalent of "just connect your GitHub" for knowledge work.
- #agents
- #enterprise
- #open-source
