ARTICLE

AI can't tell you which test mattered

AI

In 2026, AI can do most of the mechanical work of conversion optimisation. It generates variants on brand, spins up the test, watches micro-behaviours a manual setup would miss, and shifts traffic toward the winner in real time. The pitch is seductive: stop guessing, test everything, let the machine sort it out while you sleep.

The pitch has a hole in it. Speed was never the thing holding most testing programs back.

Look at where AI actually helps and where it does not. It replaces the repetitive parts well: session review, variant generation, statistical crunching, traffic allocation. What it does not replace is the judgement that decides what is worth testing in the first place, and whether a result that is technically significant is actually true. Those were always the bottleneck, and they are exactly the parts that do not scale by adding compute.

The failure mode is already common enough to have a shape. Teams deploy an AI testing tool with no structured hypothesis behind it, and the algorithm does what algorithms do: it optimises whatever it can measure. Point it at the wrong objective and it will find efficient, confident ways to hit the wrong target. It will happily generate a hundred button variants for a page whose real problem is a checkout flow that makes no sense, a problem no one told it to question. Speed applied to the wrong question just gets you to the wrong answer faster.

The numbers back this up. Linear’s 2026 review found expert-guided AI testing delivered conversion lifts of 28 to 34%, against 4 to 7% for self-serve AI tools run without a practitioner. Same technology. The difference was the human deciding what to feed it. The tool was never the lever; the thinking was.

And AI does nothing for the problem underneath all of this: data you can trust. An algorithm optimising on noisy, half-broken tracking will confidently steer you toward conclusions built on sand. It has no way to know the "lead" event fired twice, or that half your journeys are fragmented across devices. Garbage in, but now at machine speed and with a confident dashboard on top. The traffic-quality and data-quality problems that decided outcomes before AI still decide them now. AI just raises the stakes on getting them right.

So where does that leave the analyst? In a better place than the hype suggests. When generation and number-crunching get cheap, the scarce skill is not doing more of them. It is knowing what to test, reading behaviour to form a hypothesis worth testing, and having the judgement to tell a real result from a flattering one. The role shifts from execution to interpretation. That is not a downgrade. That is the part that was always the actual job.

Use AI. It is genuinely good at the mechanical work, and refusing to touch it is its own kind of stubbornness. Just do not mistake it for the strategy. A hundred tests run at speed still cannot tell you which one mattered, or why. That question is still yours.

18

Share It

El Mahdi Khiyat

Digital analyst who reads the data, shapes the experience, and builds the page that answers it.

PARIS · GMT+1

© 2026 El Mahdi Khiyat. All rights reserved.

I care about the person behind the click.