Mida's agent reads your own results and tells you what is worth testing next.

Try it for FREE now

Do styling A/B tests work?

Sometimes — and this is the most-run change type by a wide margin. Styling changes win 11% of the time. That is around the middle of the pack. What is not in question is the volume: styling is the most-run category, and it has a median lift of +22.2% when it does land.

Most-tested by a wide margin, with a middling hit rate. Styling is often what gets tested when there is no hypothesis — easy to build and easy to justify after the fact.

The numbers

Won 11% No measurable change 74% Lost 16%
Tests analysed1,000+
Beat control11%
No measurable difference74%
Lost to control16%
Median lift when it won +22.2%

Typical winning range: +12.2% to +54.2% for the middle half of winners. Median traffic per variant was 3016 visitors. Most ran on homepages (596), product pages (296), other pages (142).

What counts as a styling test

Visual treatment shifts — spacing, weight, borders, colour, shadows — with the same content in the same order. Move the content and it is a layout test. Rewrite it and it is a copy test.

Control and variant wireframe for a styling A/B test
Control on the left, variant on the right. Only the changed element is highlighted.

Adjacent categories, and how often they win: layout (15%), hero image (12%), CTA copy (12%).

Styling tests by page type

page typewin ratetests
homepages 10%500+
product pages 11%250+
category pages 10%100+
lead capture pages 10%50+
landing pages 7%50+
checkout pages 6%47 tests

Only page types with at least 30 styling tests appear here.

Styling tests that won

  • The checkout button gained click tracking, and the hero CTA button’s background color changed without any text changes.

    Travel & Hospitality+22.2%

  • It hid the comparison modal panel and the two-column grid section from view.

    Product pageE-commerce+22.2%

  • The medium buttons changed from their previous background color to orange-red, while the wording stayed the same.

    HomepageSaaS / Software+22%

  • The review-product element changed from its previous display to a flex layout on single product pages.

    E-commerce+23.3%

  • The slideshow banner slides and related overlay links and images were hidden, and on mobile the hero image, heading, and link were resized and repositioned without changing any text.

    HomepageE-commerce+20.9%

Styling tests that did not

  • The calc banner button's background and border colors changed, and its hover color was updated, while the text stayed the same.

    Professional ServicesNo significant difference

  • The old element was hidden and the new element was forced to show, while the wording stayed the same.

    HomepageSaaS / SoftwareNo significant difference

  • The request info button changed color while its text stayed the same.

    HomepageEducationNo significant difference

  • The container element was hidden instead of being shown on the page.

    Professional ServicesNo significant difference

These are individual tests, not rules. Each ran on one site, with one audience, against one page we are not showing you. They are picked to be illustrative rather than sampled at random, and a change that won here can lose on your page for reasons none of this captures. Read them as prompts for what to test, not as findings to copy.

Why the rate looks like this

A styling change is usually the smallest real intervention available: the visitor sees the same things saying the same thing in the same order, only dressed differently. Small interventions produce small effects, which is what the lift figure shows.

How this compares

change typewin ratetests
CTA colour 16%32
Layout 15%351
Form 12%81
CTA copy 12%315
Hero image 12%160
Price framing 11%117
Styling (this page) 11%1414
Split URL 10%1165
Headline 8%918
Social proof 8%196
Body copy 6%337

What these styling tests measured

A win rate only means as much as the thing being counted, so this is what these tests were actually optimising for.

measured outcomeshare of tests
Purchase / order complete51%
On-page click / engagement11%
Lead / form submission8%
Signup / registration7%

Outcomes below 5% of the set are not listed, so these do not sum to 100. How the outcome was measured moves the win rate a lot — see the methodology.

Why this table is not a ranking

These rates are not like-for-like. How often a test wins depends partly on what it chose to measure: a goal that records a soft signal — a click, a scroll — clears the bar more often than one that records a purchase. Across this data, tests measured against a proxy goal win 13%, against 9% for tests measured against a business outcome.

Categories differ a lot in that mix — proxy goals account for anywhere from 2% to 41% of a category's tests, and styling tests sit at 20%.

Hold the metric fixed and the spread narrows sharply. Comparing only tests measured against a business outcome, styling wins 8% (6–10%) rather than 11%, and the gap between the highest and lowest category falls from 9 points to 6 — at which point the intervals overlap and the categories are not statistically distinguishable.

So read the table as a description of what happened, not as a ranking of what works. If you want the comparison to mean something, compare categories that were measured the same way. The methodology page has the full breakdown.

FAQ

Why do so many people test styling?

It is the cheapest variant to build and needs no research to justify. That is a statement about cost, not about expected return.

Is it worth testing at all?

At 11% it is not worthless. But styling is the most-run category and has one of the smallest median lifts, which is worth sitting with: it may be tested because it is easy to build rather than because it is important.

Who ran these tests

This is every Mida account that ran a readable test — in-house marketers, founders, product teams, and agencies working on client sites. Nothing here is filtered by who ran the experiment or how experienced they are.

Low win rates are normal in experimentation, including at the top end. Microsoft's experimentation team, reporting on its own platform, found that only about one third of ideas improve the metric they were designed to improve — and that roughly another third actively hurt it. That is a dedicated experimentation organisation with research, prioritisation and review behind every test.

A mixed population like this one runs below that. The gap is roughly what disciplined practice buys you: ideas grounded in research rather than opinion, one variable at a time, and tests built so the result can actually be read.

A win is a variant that beat its control on that test's primary goal with a statistically significant result. Tests that never got enough traffic to say anything either way are excluded. The methodology has the full detail, including what these numbers cannot tell you.