Forrester published its evaluation of this market alongside a prediction that the following year would be another year of customer experience mediocrity.
In the same commentary it noted that improving digital experience is the single most common action organisations take when they decide to improve customer experience, according to its own priorities survey.
Put those together and the picture is uncomfortable. The most popular remedy is digital experience improvement. The tooling for digital experience improvement is mature and widely deployed. And the expected outcome is continued mediocrity.
Something in that chain is not working, and it is not the software.
What the category covers
Forrester defines experience optimization solutions as enabling the ongoing delivery of relevant, timely, and optimised digital experiences that meet evolving customer needs, using cross-channel customer interactions.
The core use cases are experimentation, next best product or offer or action, next best experience, and experimentation combined with recommendations or personalisation.
Around those sit extended uses: audience and segment work, understanding user experience and customer behaviour, cross-channel optimisation, automatic discovery of optimisation opportunities, customer feedback combined with experimentation, and feature experimentation.
That last one is worth noting because it belongs to a different function. Feature experimentation, meaning feature flags and gradual rollouts, is engineering practice rather than marketing practice, and its presence in the same category as campaign design tells you these platforms are now sold to two audiences with different vocabularies for the same mechanism.
From A/B testing to a platform
Forrester's own framing is that this market evolved from basic online testing.
That origin still shapes it. A/B testing was a simple, honest proposition: show version A to half your visitors and version B to the other half, measure which performs better, ship the winner. It required no modelling, no assumptions, and produced a causal answer rather than a correlation.
The expansion happened in three directions.
Server-side testing, because client-side approaches introduce latency and flicker, and cannot test anything that happens before the page renders. SiteSpect's maximum scores in implementation and deployment reflect an architecture built around this, spanning client-side, server-side, and hybrid with multiple deployment models including on-premise.
Personalisation, which replaces the question of which version is best with which version is best for whom. That is a different statistical problem and a considerably harder one.
And orchestration, connecting the optimisation layer to content, data, and delivery so that a winning variant becomes the default experience rather than a result in a report.
Inside The Forrester Wave: Experience Optimization Solutions, Q4 2024
Published on 10 December 2024, the evaluation scored providers across current offering and strategy.
Optimizely placed as a Leader and was the top-ranked solution in both categories, taking maximum scores in sixteen current offering criteria including campaign design, techniques of online testing, techniques of personalisation, and collaboration. Forrester credited its unified platform vision as a competitive advantage, letting customers plan, create, and optimise experiences end to end, and noted investment in generative AI features including generating variations of experiences. Forrester positioned it for marketing and business leaders and product teams needing rapid results and wanting to scale optimisation using generative AI and data beyond behavioural signals alone.
SiteSpect achieved maximum scores in implementation and deployment and in techniques of online testing, on a patented approach supporting client-side, server-side, and hybrid environments across private cloud, public cloud, on-premise, and API deployment.
One vendor topping both scored categories is a dominant result, and the collaboration criterion is worth noticing in that list. Collaboration is not a testing capability. It is workflow, and its presence as a scored criterion reflects something the market learned expensively: experimentation programmes fail on process far more often than on statistics.
The problem nobody puts in a demo
Here is the structural issue underneath this entire category, and it is the reason the mediocrity prediction and the tooling maturity can both be true.
An experiment detects an effect only if it has enough traffic to distinguish that effect from noise. The smaller the effect, the more traffic required, and the relationship is unforgiving. Detecting a large improvement takes modest volume. Detecting the two or three percent improvements that most real optimisations produce takes a great deal.
Most organisations do not have that traffic on most pages. They run a test on a checkout flow with a few thousand weekly visitors, watch the numbers for a fortnight, see one variant ahead, and ship it. The result is frequently noise, and the programme accumulates a portfolio of changes that were never validated at all.
Three specific failures follow, and none of them are the platform's fault.
Underpowered tests, as above, producing confident conclusions from insufficient data.
Peeking, where someone watches the dashboard daily and stops the test when it looks significant. That practice inflates false positive rates substantially, because a random walk will cross a threshold eventually if you keep checking. Sequential testing methods exist to handle this correctly, and most programmes do not use them.
And multiple comparisons, where a team tests twenty variants or slices results across fifteen segments and reports the ones that reached significance. At conventional thresholds, a portion of those will be false by construction.
The uncomfortable summary is that a mature experimentation programme with poor statistical discipline produces a documented history of decisions that feel evidence-based and are not. That is arguably worse than no programme, because it carries the authority of measurement.
Which generative AI makes worse before better
Forrester credits Optimizely specifically for developing features that generate variations of experiences, and that capability is genuinely useful. Producing variants was the expensive, slow part of experimentation, and removing that constraint lets teams test things they previously could not afford to build.
It also removes the natural discipline that scarcity imposed.
When each variant cost design and development time, somebody had to justify testing it. That justification was a crude form of hypothesis-setting, and it kept the number of comparisons low. When variants are free, the temptation is to generate fifty and let the system find the winner.
Fifty variants split across the same traffic means each receives a fiftieth of the sample. The statistical power to distinguish them collapses, and the winner that emerges is more likely to reflect random variation than a real effect.
Multi-armed bandit approaches address part of this by allocating traffic dynamically toward better-performing options, and most serious platforms offer them. They optimise for outcome rather than for learning, which is the right trade in some situations and the wrong one when you need to know why something worked in order to apply it elsewhere.
The practical consequence is that generation capability raises the requirement for statistical sophistication rather than lowering it. A team that could previously get away with informal practice because it ran six tests a quarter cannot when it runs six hundred.
Personalisation is a different epistemic claim
Worth separating clearly, because the two capabilities sit in the same platforms and are frequently discussed as one thing.
An experiment answers a causal question with a controlled comparison. Version B outperformed version A for this population, and randomisation means the difference is attributable to the change.
Personalisation answers a predictive question. Given what we know about this individual, which version will they respond to best. That is a model, and its accuracy depends on data quality, on the stability of behaviour over time, and on the assumption that the patterns learned from past visitors apply to this one.
Both are legitimate. They fail differently. A bad experiment gives you a wrong answer you can detect by replication. A bad personalisation model quietly serves worse experiences to some segments while the aggregate metric looks fine, because the improvements to well-modelled segments mask the degradation to poorly-modelled ones.
Forrester's positioning of Optimizely for teams wanting to scale using data beyond behavioural insights alone points at the input side of this. Behavioural data describes what someone did on your site. Contextual and declared data describes who they are, and models built on the first alone are working from a narrow window.
Why mediocrity persists
Return to the opening tension. Good tooling, widespread adoption, and an expectation of continued mediocrity.
The explanation is not that the platforms do not work. It is that optimisation operates on a fixed surface. A/B testing a checkout page makes that checkout page better. It cannot tell you that the checkout page should not exist, that the product is confusing, or that customers are irritated by something that happens after purchase and outside the digital experience entirely.
Forrester's own broader position on customer experience has consistently been that the differences that matter are organisational and cross-functional rather than interface-level. Optimisation tooling is exceptionally good at local improvement and structurally incapable of the other kind.
Which produces the pattern many organisations recognise: years of positive test results, a documented cumulative conversion lift, and a customer experience score that has not moved. Both facts are real. The tests improved the pages. The experience is determined by things the pages do not control.
That is not an argument against the category, which does what it claims efficiently. It is an argument for locating it correctly. Experience optimisation is a tool for making decisions well within a defined space. It is not a strategy for deciding which space to be in, and organisations that treat it as the latter get exactly the outcome Forrester predicted.
Analyst Source
Forrester Research
Category definition, use case framing, vendor inclusion, and evaluation findings in this article draw on Forrester's coverage of experience optimization solutions. The Q4 2024 Wave scored providers across current offering and strategy, and was published alongside Forrester commentary on the market's evolution from basic online testing and on the role of generative AI in managing personalisation at scale.
Source research
Forrester does not endorse any vendor named here, and tier placement should not be read as a recommendation to buy.