Mida's agent reads your own results and tells you what is worth testing next.

Try it for FREE now

How we built these benchmarks

Every rate comes from tests real teams ran on Mida — not a survey, not a literature review, not an estimate. This page states exactly how the numbers are produced, including the parts weaker than we would like.

Where the data comes from

Concluded and stopped experiments run on Mida between August 2023 and August 2026, across thousands of tests and hundreds of accounts. Each variant contributes one record: what changed, how much traffic it saw, its conversion rate against its control, and how it resolved.

Which records reach a published figure

Not every recorded variant is publishable, and the largest cut is for traffic rather than anything to do with the result. Of every variant we hold, 61% met the sample floor, and 55% survive every filter to reach a published rate. The rest are dropped for the reasons below.

The sample floor

A variant has to have seen enough traffic for the result to mean anything. Below that bar a flat result is not evidence the change failed — it is evidence the test never ran long enough to say. 39% of non-control variants are excluded on this basis alone, which is the main reason our win rates are higher than a raw count over all tests would give.

Removing measurement artefacts

Some recorded wins are setup faults. The clearest signature is a variant converting above 50% while its control converts under 5% — which happens when a conversion goal fires on the variant's own page, so arriving is converting. Split URL tests are where this concentrates, and the worst examples record lifts in the tens of thousands of percent. Only records were removed this way. They barely move win rates but badly distort lift, so they go first.

What counts as a win

A variant that beat its control on that test's own primary goal, with a statistically significant result. Flat and losing variants are both in the denominator.

We report the verdict each test's own settings produced rather than re-adjudicating everything at a bar of our choosing. It is what the team running it concluded and acted on; a figure we recomputed afterwards is one nobody ever saw.

Multiple variants in one test

Just under a third of variants share a test with at least one sibling. A test with four variants has more chances for one to clear its bar, so the rates carry some inflation from that. We report each variant as one observation rather than correcting for it, because these are descriptive rates rather than decisions.

What these tests measured, and why it matters

Every test here counts something. What it counts turns out to predict whether it "wins" almost as strongly as what was changed.

measured outcomeshare of testswin rate95% interval
Purchase / order complete45%12%10–13%
Signup / registration12%10%8–13%
Lead / form submission11%15%12–18%
On-page click / engagement9%19%16–23%
Product or content page reached5%17%13–22%
Booking / appointment4%7%4–11%
Add to cart4%14%10–19%
Reached a later funnel step4%11%7–16%
Outbound click to a provider3%15%10–22%
Scroll or time on page1%20%12–31%
Subscription1%3%0–14%

Read the bottom of that table against the top. A test measured on scroll depth or an on-page click wins around one time in five. A test measured on a completed purchase wins about one time in eight, and one measured on a booking or a subscription less often still. The intervals for on-page clicks and purchases do not overlap, so this is a real gap and not sampling noise.

None of that means the softer metrics are wrong to use. A click can be the right thing to optimise, and an early-funnel test often has no purchase to point at. It does mean a win rate is only worth as much as the outcome behind it, and that two rates measured on different outcomes are not comparable. Every page here shows what its tests measured for exactly that reason.

Outcomes are classified from the goal's configured value and the name the customer gave it, in two passes: deterministic rules where the signal is unambiguous, and a language model over the remainder, which is often not in English. 0.4% could not be placed and are excluded from the table. The goal names themselves are customer content and are never published — only the category.

Why you should not compare categories directly

This is the biggest caveat on this data, and it is one we got wrong in our first publication.

Every rate here is partly a statement about what those tests chose to measure, not only about what they changed. Goals divide into business outcomes such as a purchase, funnel steps such as reaching a checkout page, and proxies such as a click or a scroll. They do not clear the bar equally often:

what the test measuredwin rate
A proxy — a click, a scroll, time on page13%
A funnel step — reaching a later page10%
A business outcome — a purchase, a signup9%

If the mix were the same across categories this would not matter. It is not: proxy goals account for anywhere from 2% to 41% of a category's tests. Some kinds of change are simply more often measured against a soft signal than others.

The consequence is direct. Comparing only tests measured against a business outcome — like for like — the gap between the highest and lowest category falls from 9 percentage points to 6, and at that point the confidence intervals overlap. Layout and headline, the comparison our pages leaned on hardest, go from 15% versus 8% to 12% versus 10% — a difference we cannot distinguish from noise at these sample sizes.

So the per-category rates are sound as descriptions and weak as a ranking. "15% of layout tests won" is true. "Layout beats headline, so test layout" is not supported once you account for what each was measured against. We have softened those recommendations across the site rather than leaving them standing, and every comparison table now carries this note.

The honest version of the advice: pick the change you have a real hypothesis for, and if you want to compare categories yourself, compare ones measured the same way.

What counts as a change type

Labels are assigned automatically from the variant definition, not by hand.

how the label was derivedrecords
Code diff3617
Text only2120
Redirect target1658
Screenshot comparison533
Structural diff6

Records the labeller could not classify are dropped rather than filed under a catch-all. Tests changing several things at once are labelled by their dominant change, so a "headline" test here may have carried minor styling changes too. Single-variable discipline is a property of the teams running the tests, not something we can enforce retrospectively.

Industry comes from the verified domain

Our first publication had no industry breakdowns, because the only industry field available was a self-reported onboarding answer. Industry is now derived from the domain the Mida pixel actually fired on, covering 97.8% of published variants.

That change was worth making: where we hold both values, the self-reported industry and the verified one disagree 65% of the time. The substitutions are not near misses — sites self-reporting as Professional Services that are plainly e-commerce, sites self-reporting as SaaS that are agencies. Any industry benchmark built on the survey field would have been wrong about roughly two rows in three.

Domains are classified in two passes with an explicit "unclear" option at each stage. The first reads the domain and its real page paths; the second fetched the unplaced homepages and classified them from each site's own title and description. A small remainder is still unresolved and are excluded from every industry figure — some no longer respond, some sit behind a bot wall, some return a page that still does not say what the business does. A bot wall is recorded as blocked rather than fed to the classifier, since an interstitial would otherwise be labelled SaaS every time. This is a classification, not a registry lookup: treat a single label as good-not-certain.

What we do not publish

  • No customer copy. No headline, button label or body text from any account appears here. Descriptions of what changed are machine-rewritten to remove verbatim wording, brand names, product names, prices, selectors and domains.
  • No URLs, no account names, no test names.
  • Nothing unreviewed. 58 descriptions the rewrite left with residual identifying detail are withheld from every example. They remain in the denominators, because a description we cannot quote is still a test that ran.
  • No cell under 30 tests. Thinner categories are not published at all rather than published with a caveat.

What these numbers cannot tell you

They do not tell you what will work on your site. A 15% win rate for layout changes is a base rate across hundreds of sites, not a prediction about yours. The useful move is comparative: if layout wins roughly twice as often as headline copy, spend a scarce test slot on layout.

Goal quality varies. Roughly a third of these tests measured a business outcome such as a purchase, a third a funnel step, a fifth a proxy such as a click. A proxy win is weaker evidence than a purchase win, and the headline rates do not separate them.

Survivorship runs through everything. These are tests that got set up, launched and left running long enough to read. Abandoned tests, and ideas nobody built, are invisible.

The corpus skews to what people test, not what works. Styling is the largest category by a wide margin. That tells you it is cheap to build, not that it is important.

Updates

Figures are recomputed periodically and every page carries its own last-updated date. We do not revise past numbers silently — where a recomputation moves a rate, the page says so.

Revision, August 2026. Industry re-derived from the verified domain the pixel fired on, including a homepage fetch for domains a name alone could not place. That made industry benchmarks publishable for the first time, at 97.8% coverage.

Who ran these tests

This is every Mida account that ran a readable test — in-house marketers, founders, product teams, and agencies working on client sites. Nothing here is filtered by who ran the experiment or how experienced they are.

Low win rates are normal in experimentation, including at the top end. Microsoft's experimentation team, reporting on its own platform, found that only about one third of ideas improve the metric they were designed to improve — and that roughly another third actively hurt it. That is a dedicated experimentation organisation with research, prioritisation and review behind every test.

A mixed population like this one runs below that. The gap is roughly what disciplined practice buys you: ideas grounded in research rather than opinion, one variable at a time, and tests built so the result can actually be read.

A win is a variant that beat its control on that test's primary goal with a statistically significant result. Tests that never got enough traffic to say anything either way are excluded. The methodology has the full detail, including what these numbers cannot tell you.