In 2024, Forrester did something no major analyst firm had done before: it scored the large language models themselves. Not the platforms that served them, not the consultancies that deployed them, the models. Ten of them, against twenty one criteria, ranked head to head.
Two years later, the firm has more or less concluded that this was the wrong thing to measure.
That reversal is the most useful thing to understand about this category, and it is why anyone building a shortlist today should be careful about which analyst document they are reading. The category has not been abandoned. It has been split into two, and the split tells you something about how enterprises actually buy AI.
The thing being evaluated
A foundation model is a large model trained on broad data that can be adapted to many downstream tasks rather than built for one. In practice, for enterprise buying purposes, it means the language models: the thing you send a prompt to and get text back from.
The reason a market formed around this at all is that the models became purchasable independently of everything else. You could pick a model the way you pick a database. For a while, that framing held.
What made it a hard market to evaluate was pace. Forrester itself described this as one of the most inscrutable markets a buyer could face, driven by the rate of innovation and the choice between well capitalised startups and established platform vendors. A ranking published in June is describing a field that has already moved by September.
What The Forrester Wave: AI Foundation Models For Language, Q2 2024 actually measured
The evaluation covered ten providers: Amazon Web Services, Anthropic, Cohere, Databricks, Google, IBM, Microsoft, Mistral AI, NVIDIA, and OpenAI.
Google took the highest scores in both current offering and strategy, with Gemini singled out for multimodality, context window length, and how tightly it connected to the surrounding cloud services. Databricks placed as a Leader on the strength of a complete generative AI stack rather than a standalone model. IBM landed as a Strong Performer, positioned almost entirely on training data transparency and the licensing risk that transparency removes.
Read those three summaries together and something stands out. Only one of them is really about model quality.
The report's own guidance pointed the same direction. It told buyers to look past incremental benchmark gains and toward roadmap fit, the ability to govern and configure models to reduce hallucination, respect for intellectual property in training data, and whether the thing stays up under load. Benchmarks were the least of it.
Why the category broke
Model capability turned out to be a poor predictor of whether an enterprise deployment succeeded.
Forrester has since put this bluntly, observing that raw capability predicts little about enterprise success and pointing at the behaviour of the model labs themselves as evidence. Anthropic and OpenAI both stood up large services operations that place engineers alongside customers to build the actual solutions. The most sophisticated software ever written needed people sitting next to the buyer to produce value from it.
That is not a criticism of the models. It is a statement about where the difficulty lives. Context, orchestration, evaluation, workflow integration, and governance are the expensive parts, and none of them are properties of a model.
So the unit of analysis moved up a layer.
The Forrester Wave: AI Platforms, Q3 2026 and the two paths that replaced it
The successor coverage arrived in August 2026 as The Forrester Wave: AI Platforms, Q3 2026, and it looks nothing like a model ranking. Fifteen vendors: Amazon Web Services, C3 AI, Databricks, Dataiku, DataRobot, Google, IBM, Microsoft, Oracle, Palantir, Pegasystems, Salesforce, SAS, ServiceNow, and UiPath.
Workflow automation vendors. RPA vendors. Enterprise SaaS. Data management. Cloud infrastructure. A category that used to mean data science workbenches now includes companies that would not have appeared in an AI evaluation five years ago, because agentic AI redrew the boundary of what an AI platform has to do. Agents do not stop at producing an insight. They navigate systems and take action, and that pulls workflow and orchestration into the definition.
Running alongside it, Forrester has declared a second and genuinely new category: frontier AI model platforms. The qualifying test has three parts. The vendor develops and owns state of the art models rather than fine-tuning or reselling someone else's. It carries the capital burden of advancing them. And it wraps them in the tooling, context, evaluation, and agent infrastructure that converts model capability into business capability.
Frontier, in this definition, describes a commitment to building models rather than a position on a leaderboard. That distinction matters, because it excludes a large number of companies currently marketing themselves as AI vendors.
The Frontier AI Model Platforms Landscape is scheduled for Q4 2026, with a full Wave evaluation following in Q1 2027.
What this means if you are buying now
The two categories represent different bets, and Forrester's framing is that you may reasonably make both.
A frontier model platform is a bet that riding one rapidly improving model beats optimising around a stable one. You accept lock-in to a single lab's trajectory in exchange for being early to whatever it ships next. A model-agnostic AI platform is the opposite bet: your data context, your orchestration layer, and your ability to swap models matter more than any single model's current lead. Several vendors sit on both paths and will happily sell you either.
Which means the practical question is not which model is best. It is which of those two bets your organisation is structurally capable of making. An enterprise with deep proprietary data and long integration cycles is a poor candidate for chasing model releases. A product team shipping AI-native features has the opposite profile.
One caution on the 2024 evaluation, since it is still the most cited document in this category and it is now two years old. Every model in it has been superseded, several by multiple generations. The vendor list remains a reasonable map of who was serious about enterprise language models at that point, and the criteria are still a defensible checklist. The scores are history.
Analyst Source
Forrester Research
Category definitions, vendor inclusion, and evaluation findings in this article draw on Forrester's coverage of AI foundation models and AI platforms. Its Wave methodology scores providers on current offering, strategy, and customer reference interviews, and publishes scorecards buyers can reweight against their own criteria.
Source research
Forrester does not endorse any vendor named here, and tier placement should not be read as a recommendation to buy.