It started in 2012, with three engineers quitting Oracle to build a cloud database, a Berkeley research project turning into the company Databricks, and Amazon and Google quietly launching cloud data warehouses. It is worth stepping back and asking what it all adds up to, where the data-and-AI world stands now, and what the long arc of this story actually tells us. Viewed from here, the scattered milestones, the funding rounds, the acquisitions, the model releases, the rivalries, resolve into a single coherent transformation, one of the most important in the history of technology.

The Arc in One Sentence
If you had to compress the whole story into one idea, it would be this: over roughly fifteen years, the ability to store, process, and ultimately reason over data has gone from a costly capability reserved for the largest and most technical organizations to an accessible utility available to almost anyone, and in the process data infrastructure has become the foundation of artificial intelligence, the defining technology of the age. The data revolution and the AI revolution are the same revolution, viewed at two different points in time.
Every step builds on the last. The cloud warehouses made analytics affordable and democratized access to serious data work. The separation of storage and compute, and the rise of cloud-native platforms, have made data infrastructure flexible and scalable. The open-source frameworks and the Transformer architecture made advanced machine learning and then generative AI possible. The data platforms evolved from places to run reports into the foundations on which enterprises build AI. And the agentic shift is turning AI from something that answers into something that acts. The destination, AI woven into the fabric of how organizations operate, was implicit in the very first move toward making data cheap and accessible.
Where Things Stand Now
The landscape has a recognizable shape. A small number of dominant data-and-AI platforms anchor the industry: Databricks, valued at 134 billion dollars and still pointedly private, and Snowflake, public and reinvented around AI, as the great independent rivals, alongside the cloud giants Amazon, Microsoft, and Google with their own deeply integrated offerings. Beneath them all sits NVIDIA, supplying the chips and increasingly the models and systems, the most powerful single force in the economy. Around them orbits an ecosystem of specialized tools, many likely to be absorbed by the big platforms, and a rich, globally contested supply of AI models, open and closed, American and Chinese.
The central competitive question of the whole era remains unresolved, which is part of what makes it still so dynamic: who will be the single platform where an enterprise’s data and AI truly live. Databricks and Snowflake continue to converge and compete for that position, the cloud giants press their own claims, and the answer differs from company to company. No one has won definitively, and the agentic AI era has simply raised the stakes of the same long contest.
The Recurring Lessons
A few themes recur so insistently that they amount to lessons. The first is that architecture matters more than incumbency: the companies that won, Databricks and Snowflake, built for where the world was going, the cloud, while those that bet on the prevailing model of their day, the Hadoop vendors, were overtaken despite their early lead. Being early is not the same as being durable.
The second is the power of open source and openness as a strategy: again and again, from Spark to TensorFlow to the open-weights models, freely released technology drove adoption, built ecosystems, and created the foundations on which businesses were built. The third is the relentless logic of consolidation: specialized tools repeatedly emerge to solve new problems, the modern data stack, vector databases, MLOps, and the big platforms repeatedly absorb them, because customers prefer one integrated home over many stitched-together pieces. And the fourth, threaded through everything, is democratization: each advance takes a capability that belonged to a privileged few and makes it available to many, which is the quiet engine driving the entire transformation.
What It Means Going Forward
For enterprises, the state of play offers remarkable capability and genuine choice: powerful, competing platforms on which to build AI-driven operations, a rich supply of models to choose from, and infrastructure that makes sophisticated data and AI work accessible as never before. The benefit of this long competition is precisely this abundance; the relentless rivalry among the major players has produced better, cheaper, more capable tools than any monopoly could.
The risks have grown alongside the capabilities. The concentration of power among a handful of platforms, especially NVIDIA, raises real concerns about dependence and resilience. The agentic systems now acting on company data demand governance and trust that the industry is still learning to provide. The reliability problems of AI, hallucination, opacity, and the difficulty of verifying machine output are managed but not solved. And the entanglement of AI with geopolitics adds a layer of uncertainty no enterprise can fully control. The transformation is not finished; in many ways, the most consequential chapters, the ones about what a world saturated with capable, acting AI actually becomes, are still being written.
But the foundation is now unmistakable. What started as a quest to make data cheaper and more accessible has become the substrate of artificial intelligence, and the companies that built that substrate, from a Berkeley research project and a stealthy Oracle veteran startup to platforms worth tens and hundreds of billions of dollars, find themselves at the center of the most important technological shift of our time. The data era and the AI era are one and the same, and we are still very much living inside it.