Mida's agent reads your own results and tells you what is worth testing next.

Try it for FREE now

How we built these benchmarks

Every rate comes from tests real teams ran on Mida — not a survey, not a literature review, not an estimate. This page states exactly how the numbers are produced, including the parts weaker than we would like.

Where the data comes from

Concluded and stopped experiments run on Mida between August 2023 and August 2026, across thousands of tests and hundreds of accounts. Each variant contributes one record: what changed, how much traffic it saw, its conversion rate against its control, and how it resolved.

Which records reach a published figure

Not every recorded variant is publishable, and the largest cut is for traffic rather than anything to do with the result. Of every variant we hold, 61% met the sample floor, and 55% survive every filter to reach a published rate. The rest are dropped for the reasons below.

The sample floor

A variant needs at least 100 visitors. Below that, a flat result is not evidence the change failed — it is evidence the test never ran long enough to say. 39% of non-control variants are excluded on this basis alone, which is the main reason our win rates are higher than a raw count over all tests would give.

Removing measurement artefacts

Some recorded wins are setup faults. The clearest signature is a variant converting above 50% while its control converts under 5% — which happens when a conversion goal fires on the variant's own page, so arriving is converting. Split URL tests are where this concentrates; the worst example in the corpus recorded a lift of over 75,000%. Only records were removed this way. They barely move win rates but badly distort lift, so they go first.

What counts as a win

A variant that beat its control on that test's own primary goal, resolved at the confidence threshold the team running it had set.

We report the verdict those settings produced rather than re-adjudicating every test at a bar of our choosing. It is what the team concluded and acted on; a figure we recomputed afterwards is one nobody ever saw. In practice the threshold is one of three:

confidence thresholdshare of variants
95%43.1%
80%32%
90%22.8%
99%1.8%

97.8% of variants sit at 80%, 90% or 95%, and 95% is the single most common setting. A long tail of other values accounts for the remainder. So these are not one uniform bar, and you should know that before quoting a rate — but nor are they scattered. 97% of these tests used Mida's Bayesian engine; the rest a frequentist test.

A note on what this means for comparison. A rate resolved at a mix weighted toward 95% is not interchangeable with one computed strictly at 95%, and it will read a little higher. If you need the stricter figure for a particular category, ask us and we will run it.

Multiple variants in one test

Just under a third of variants share a test with at least one sibling. A test with four variants has more chances for one to clear its bar, so the rates carry some inflation from that. We report each variant as one observation rather than correcting for it, because these are descriptive rates rather than decisions.

What counts as a change type

Labels are assigned automatically from the variant definition, not by hand.

how the label was derivedrecords
Code diff3617
Text only2120
Redirect target1658
Screenshot comparison533
Structural diff6

Records the labeller could not classify are dropped rather than filed under a catch-all. Tests changing several things at once are labelled by their dominant change, so a "headline" test here may have carried minor styling changes too. Single-variable discipline is a property of the teams running the tests, not something we can enforce retrospectively.

Industry comes from the verified domain

Our first publication had no industry breakdowns, because the only industry field available was a self-reported onboarding answer. Industry is now derived from the domain the Mida pixel actually fired on, covering 97.8% of published variants.

That change was worth making: where we hold both values, the self-reported industry and the verified one disagree 65% of the time. The substitutions are not near misses — sites self-reporting as Professional Services that are plainly e-commerce, sites self-reporting as SaaS that are agencies. Any industry benchmark built on the survey field would have been wrong about roughly two rows in three.

Domains are classified in two passes with an explicit "unclear" option at each stage. The first reads the domain and its real page paths; the second fetched the unplaced homepages and classified them from each site's own title and description. A small remainder is still unresolved and are excluded from every industry figure — some no longer respond, some sit behind a bot wall, some return a page that still does not say what the business does. A bot wall is recorded as blocked rather than fed to the classifier, since an interstitial would otherwise be labelled SaaS every time. This is a classification, not a registry lookup: treat a single label as good-not-certain.

What we do not publish

  • No customer copy. No headline, button label or body text from any account appears here. Descriptions of what changed are machine-rewritten to remove verbatim wording, brand names, product names, prices, selectors and domains.
  • No URLs, no account names, no test names.
  • Nothing unreviewed. 58 descriptions the rewrite left with residual identifying detail are withheld from every example. They remain in the denominators, because a description we cannot quote is still a test that ran.
  • No cell under 30 tests. Thinner categories are not published at all rather than published with a caveat.

What these numbers cannot tell you

They do not tell you what will work on your site. A 12% win rate for layout changes is a base rate across hundreds of sites, not a prediction about yours. The useful move is comparative: if layout wins roughly twice as often as headline copy, spend a scarce test slot on layout.

Goal quality varies. Roughly a third of these tests measured a business outcome such as a purchase, a third a funnel step, a fifth a proxy such as a click. A proxy win is weaker evidence than a purchase win, and the headline rates do not separate them.

Survivorship runs through everything. These are tests that got set up, launched and left running long enough to read. Abandoned tests, and ideas nobody built, are invisible.

The corpus skews to what people test, not what works. Styling is the largest category by a wide margin. That tells you it is cheap to build, not that it is important.

Updates

Figures are recomputed periodically and every page carries its own last-updated date. We do not revise past numbers silently — where a recomputation moves a rate, the page says so.

Revision, August 2026. Industry re-derived from the verified domain the pixel fired on, including a homepage fetch for domains a name alone could not place. That made industry benchmarks publishable for the first time, at 97.8% coverage.

Who ran these tests

This is every Mida account that ran a readable test — in-house marketers, founders, product teams, and agencies working on client sites. Nothing here is filtered by who ran the experiment or how experienced they are.

Low win rates are normal in experimentation, including at the top end. Microsoft's experimentation team, reporting on its own platform, found that only about one third of ideas improve the metric they were designed to improve — and that roughly another third actively hurt it. That is a dedicated experimentation organisation with research, prioritisation and review behind every test.

A mixed population like this one runs below that. The gap is roughly what disciplined practice buys you: ideas grounded in research rather than opinion, one variable at a time, and tests built so the result can actually be read.

A win is a variant that beat its control on that test's primary goal with a statistically significant result. Tests that never got enough traffic to say anything either way are excluded. The methodology has the full detail, including what these numbers cannot tell you.