The lakehouse exists because of an argument about storage. Data warehouses were structured, governed, and expensive. Data lakes were cheap, flexible, and ungoverned to the point of uselessness. The lakehouse proposed one thing that did both.

Forrester's guidance alongside its current evaluation of the market is that buyers should stop prioritising storage-centric capabilities.

That is a category being told to abandon its founding premise as an evaluation basis, and it happened inside two years. The reasoning is worth following, because it changes what the platform is for rather than merely what it does.

What the architecture actually merged

A data warehouse imposed schema before writing. That produced reliable, queryable, governed data and a bottleneck, because every new source required modelling work before anyone could use it.

A data lake inverted it. Write anything in any format, work out the structure when you query. That removed the bottleneck and produced the well-documented failure mode where nobody could find anything, nothing was trusted, and the term data swamp entered the vocabulary.

The lakehouse keeps lake storage, cheap object storage holding open file formats, and adds the properties that made warehouses trustworthy: transactional consistency, schema enforcement where you want it, time travel, and governance. Open table formats are the mechanism, sitting between the files and the engines and maintaining the metadata that makes a directory of files behave like a table.

The consequence people care about commercially is that multiple engines can operate on the same data without copying it. Cloudera's citation in the current evaluation leads with exactly that: zero-copy multi-engine access over governed data.

Copying data has been the hidden tax in enterprise analytics for two decades. Every copy is a divergence, a governance gap, and a storage bill.

Inside The Forrester Wave: Data Lakehouses, Q2 2024

Published on 30 April 2024, the evaluation scored thirteen providers against twenty four criteria.

Databricks placed as a Leader with the highest score in both the current offering and strategy categories, taking maximum marks in nineteen of the twenty four criteria.

Google also placed as a Leader with BigQuery, taking maximum scores in fifteen criteria.

One vendor topping both scored dimensions with nineteen of twenty four criteria maxed is a dominant result, and it reflected a company that had named the category and shipped the reference implementation of it.

Worth noting the adoption figure circulating at that point: roughly three quarters of global CIOs reported having a lakehouse in their estate, with most of the remainder intending to within three years.

A category at that penetration has stopped being an architectural choice and become the default. Which is usually the moment the basis of competition moves.

Inside The Forrester Wave: Data Lakehouses, Q3 2026

The current evaluation scores providers on current offering, strategy, and customer feedback.

Cloudera placed as a Leader, taking the highest possible score in the vision and roadmap criteria. Forrester's assessment of the platform covers zero-copy multi-engine access over governed data, real-time streaming, federated SQL, semantic and vector search, agentic AI workflows, autoscaling, and enterprise observability, positioning it for enterprises needing hybrid flexibility, strong governance, real-time analytics, and operational control across large, distributed, highly regulated environments.

Read that capability list against the 2024 framing and the shift is visible in the nouns. Semantic and vector search. Agentic AI workflows. Observability. Real-time streaming. None of those describe storage.

Forrester's own commentary makes the reframing explicit. The lakehouse, in its assessment, is no longer limited to storing and serving data for analytics. It must continuously provide trusted, governed, real-time context that AI agents can use to reason, decide, and act. As AI moves from generating insights to executing business processes, organisations should evaluate vendors on how effectively their platforms operationalise AI rather than on storage capability.

The phrase Forrester uses for what the lakehouse is becoming is an execution layer for agentic AI.

Why agents change the requirement

The distinction between serving analytics and serving agents is not marketing, and it is worth being precise about what actually differs.

An analytical query is asked by a person who will interpret the answer. If the number looks wrong, they notice. If the data is a day stale, they usually know and adjust. If two reports disagree, someone investigates. Human judgement sits between the data and the decision, absorbing a great deal of imprecision.

An agent acting on data has no such buffer. It receives a value and acts. Stale data produces a wrong action rather than a puzzled analyst. Ambiguous definitions produce confident errors. And because agents operate continuously rather than when someone opens a dashboard, the errors compound before anyone reviews them.

That raises the requirement on four things simultaneously.

Currency, because an agent deciding on yesterday's inventory position makes commitments the business cannot keep. This is why real-time streaming appears in the capability lists.

Lineage, because when an agent takes a wrong action, the question is immediately where the input came from and what produced it. A lineage graph that a human could reconstruct with effort is insufficient when the volume of agent actions exceeds what anyone can review.

Governance and permissions, because an agent querying on behalf of a user must not return what that user cannot see. This is the same failure described in retrieval systems: information leaking through an answer rather than through a link.

And semantics, because an agent has no informal knowledge of which of your seven revenue columns finance recognises. What a human analyst resolves by asking a colleague, an agent resolves by picking one.

Vector search appearing alongside SQL in a lakehouse capability list is the visible edge of this. The platform is being asked to serve both structured queries and semantic retrieval against the same governed data, because agents need both.

The format question underneath

Open table formats are the layer that makes the multi-engine promise real, and they carry a strategic decision that vendors present as a technical one.

The proposition of an open format is that your data sits in your storage in a documented format, and any compliant engine can read it. That is genuinely different from a proprietary warehouse where the data is only accessible through the vendor's engine, and it is the strongest argument in the lakehouse case.

The complication is that the formats have vendor lineages, the ecosystems around them are not equally mature, and interoperability in practice is less complete than the specifications suggest. Writing from multiple engines concurrently, catalogue federation, and governance enforcement across engines all vary considerably in how well they work.

The practical question for a buyer is narrower than the ideology. If you leave this vendor, what actually moves? If the answer is that the files stay and you point a different engine at them, the openness is real. If the answer involves migrating a proprietary catalogue, rewriting governance policy, and rebuilding pipelines, the format is open and the platform is not.

That question matters more now than it did in 2024, because the capabilities Forrester is scoring have moved up the stack. Storage was the interoperable part. Semantic layers, agent workflows, and governance are where the lock-in has relocated.

What the category won, and lost

Return to the adoption figure. Three quarters of large enterprises had a lakehouse in 2024, and the remainder mostly planned one.

Categories that reach that penetration face a predictable transition. The architectural argument is settled, so competing on it stops working. Every vendor supports open formats, separates storage from compute, and handles both structured and unstructured data. The differentiation that built the category evaporated into table stakes.

What replaces it, on the evidence of the current evaluation, is operational capability at the edges: real-time streaming, observability, governance in distributed and regulated environments, and the agentic workflow layer.

Cloudera's positioning is instructive precisely because it is unfashionable. Hybrid flexibility and operational control across distributed, highly regulated environments is not a pitch about elegance. It is a pitch to organisations whose data cannot all move to one cloud, for regulatory or sovereignty or practical reasons, and who have been underserved by an industry that assumed consolidation.

That is a real and durable segment, and its existence tells you something the aggregate adoption numbers do not: the lakehouse won as an architecture, and the assumption that everything would consolidate into one of them did not.

Where this leaves the decision

The uncomfortable implication of Forrester's guidance is that the lakehouse decision has become downstream of an AI decision most organisations have not made explicitly.

If your AI programme is analytical, generating insight that people act on, the storage-centric evaluation criteria that this category was built on remain adequate. Query performance, cost per terabyte, and format openness will serve you.

If your AI programme involves agents acting on data without a person in the path, the requirements change to the ones described above, and they are harder to assess. Currency, lineage, semantic consistency, and permission enforcement under machine-speed query load are not things a proof of concept on clean data will reveal.

Most organisations are somewhere in between and moving in one direction. The reasonable position is to evaluate against where the AI programme is heading rather than where it is, while recognising that the vendors are describing a destination the industry has not reached either.

Forrester has been unusually direct that the basis of comparison has moved. That instruction is more useful than any tier placement, because it tells you which questions produce a decision that survives the next two years rather than one that answers the last two.

Analyst Source

Forrester Research

Category definition, vendor inclusion, and evaluation findings in this article draw on Forrester's coverage of data lakehouses. The Q2 2024 Wave scored 13 providers against 24 criteria across current offering and strategy. The Q3 2026 Wave scores providers on current offering, strategy, and customer feedback, and is accompanied by Forrester guidance that buyers should evaluate platforms on how effectively they operationalise AI rather than on storage-centric capabilities.

Source research

Forrester does not endorse any vendor named here, and tier placement should not be read as a recommendation to buy.