Skip to content
HomeAbout UsInsightsContact UsBook a free consultation

Shopify CRO, run as an experiment programme

Research first, then prioritised hypotheses, then tests powered well enough that the result means something. Most of what a CRO programme produces is losing and inconclusive tests — a programme that only ever reports wins is not measuring, it is marketing.

Why TLX

Most "CRO wins" are noise

A test on a page with 400 weekly sessions, stopped the moment it looked good, will report a 30% uplift roughly whenever you like. Underpowered tests stopped early are the single most common failure in ecommerce experimentation.

Powered, or not run

Before a test starts we calculate the sample size needed to detect an effect worth caring about, and the runtime that implies at your traffic. If that number is unreachable, we say so and do something else instead.

Tests run to their planned duration and whole weeks, because stopping when the line looks good is how noise becomes strategy.

How the programme runs
Powered, or not run

When you do not have the traffic

Plenty of stores cannot power a meaningful A/B test on anything but the homepage or cart. That is not a reason to fake it.

For those stores the programme is research, heuristic and accessibility fixes, session replay, and changes with such obvious rationale that measuring them costs more than shipping them. We will tell you which situation you are in.

Testing vs shipping
When you do not have the traffic
Technology

What the programme is built on

Evidence you can audit, not a dashboard that only ever goes up.

Analytics that agree

Before any testing, the data has to be trustworthy — deduplicated events, correct attribution, and a funnel that reconciles with Shopify’s own reporting.

Qualitative research

Session replay, on-site surveys, support ticket themes and user testing. Quantitative data says where; qualitative says why.

Properly powered experiments

Sample size calculated up front from your real traffic and baseline conversion, runtime set in whole weeks, and a stopping rule agreed before the test starts rather than negotiated once the numbers move.

Server-side rendering, no flicker

Variants rendered in Liquid rather than swapped by a client-side script after paint. Flicker costs conversion on the variant itself, which quietly biases the result you are trying to measure.

Segment analysis

New versus returning, device, and traffic source — because an aggregate win can hide a loss in the segment that matters most.

Shopify-native measurement

Checkout and order data from Shopify, so revenue-per-visitor is measured against the source of truth rather than a pixel.

Scope

What a CRO programme covers

Research and analytics come first. Testing is what happens once there is something worth testing.

Measurement audit

Event tracking, attribution and funnel reconciliation. If the numbers are wrong, everything downstream is decoration.

Research

Replay, surveys, support themes, heuristic review and usability testing, turned into a ranked list of problems.

Hypotheses ranked by expected value

Each hypothesis carries the evidence behind it, the pages it affects, the traffic those pages get and the build cost. Ranked by expected value rather than by how interesting it is, which is how the homepage stops absorbing a programme that should be fixing the product page.

  • Traffic on the page
  • Size of the problem
  • Cost to build

Reporting that includes the losses

Every test reported with its result, including the ones that lost and the many that were inconclusive. A programme reporting only wins is either not measuring properly or not telling you everything — and both cost more than a losing test.

  • Wins
  • Losses
  • Inconclusive

Implementation

Variants built server-side in Liquid so there is no flicker, and winners merged into the theme rather than left running in a testing tool forever.

Non-test improvements

Accessibility, performance and obvious usability fixes shipped without testing, because measuring them costs more than doing them.

Decide

A/B testing vs just shipping the fix

Testing is not always the right call. Traffic decides more than opinion does.

  Just ship it Test it
Under ~1,000 conversions/month The only realistic option Cannot be powered
Obvious accessibility or bug fix Ship it Wasteful
Cost of being wrong is low Ship it Overhead
Genuinely contested decision Guesswork Worth the runtime
High-traffic template Leaves value on the table Compounding gains
Result needs to persuade a board Anecdote Evidence
Change affects revenue directly Risky Measure it
Best when Traffic is low or the answer is obvious Traffic supports it and the answer is not
How we work

How a CRO programme runs

7 stages. Open any one of them.

Analytics audit

Trust the data before using it.

Event tracking, attribution and funnel reconciliation against Shopify’s own numbers. Programmes built on broken measurement produce confident nonsense.

TrackingAttributionReconciliation

Find the real problems

Quantitative says where, qualitative says why.

Replay, surveys, support themes, heuristic review and usability testing, turned into a ranked problem list with evidence attached.

ReplaySurveysHeuristics

Hypotheses

Ranked by expected value.

Each one carries evidence, affected pages, traffic and build cost. The ranking is arithmetic, not enthusiasm.

EvidenceTrafficEffort

Variant and test plan

Including the stopping rule.

Sample size, runtime in whole weeks, primary metric and stopping rule agreed and written down before anything is built.

Power calcPrimary metricStopping rule

Implementation

Server-side, no flicker.

Variants rendered in Liquid rather than swapped after paint, so the test does not degrade the thing it is measuring.

LiquidNo flickerQA

Test and analyse

To the planned duration.

Run to plan, then analyse in aggregate and by segment. Inconclusive is a legitimate and common result, reported as such.

Full runtimeSegmentsHonest result

Ship or revert

Then repeat.

Winners merged into the theme; losers reverted and documented so the same idea is not re-proposed in six months.

MergeRevertDocument
Engagement

Ways to engage

Three shapes, depending on how well defined the work is. Every one starts with a scoped written proposal — no work begins on a verbal estimate.

Fixed scope

Fixed price

against a written scope

  • Written scope and acceptance criteria up front
  • Fixed price against that scope
  • Staged delivery with review points
  • Change requests priced separately, never absorbed silently
Request a proposal

Time & materials

Tracked time

billed as used

  • Billed against tracked time
  • Suits discovery, R&D and migrations
  • Estimate per ticket before it starts
  • Stop or change direction at any point
Talk it through

Rates are quoted against a written scope rather than published as a tier, because the same service costs very different amounts on a five-template store and a five-hundred-template one.

Context

The categories we build for

Where conversion is lost is highly category-specific — sizing in apparel, specification in electronics, reorder friction in B2B.

Fashion & Apparel
Beauty & Cosmetics
Food & Beverage
Jewelry & Luxury
Home & Outdoor
Electronics
Health & Supplements
Subscription
Sports & Fitness
Multi-region
Marketplace & Multi-store
FAQ

Questions about Shopify CRO

An experiment programme that starts with a measurement audit, then research, then hypotheses ranked by expected value, then tests powered well enough that the result means something. Variants are rendered server-side in Liquid to avoid flicker, sample size and stopping rules are agreed before a test starts, and every result is reported — wins, losses and inconclusive.

It depends on your baseline conversion rate and the effect size worth detecting, but as a rough guide, meaningful tests on anything other than your highest-traffic templates get difficult below roughly a thousand conversions a month. We calculate the required sample size for your store before proposing a testing programme, not after.

Then a testing programme is the wrong purchase and we will say so. Research, heuristic and accessibility fixes, session replay and obvious usability improvements all still apply — they just get shipped rather than tested.

No, and anyone who does is either not measuring properly or not telling you about the losses. Most well-run CRO programmes produce more inconclusive and losing tests than winners; the value is in the compounding of the winners and in not shipping the losers.

Only where it earns its place. We prefer to render variants server-side in Liquid, because client-side tools swap content after paint and that flicker measurably costs conversion on the variant — biasing the result you are trying to read.

To its planned duration, in whole weeks, with the stopping rule agreed before the test starts. Stopping early because the line looks good is the most common way a CRO programme generates confident nonsense.

Server-side variants add essentially nothing. Client-side testing tools do add a blocking script, which is one of several reasons we avoid them, and performance is monitored throughout the programme.

The measurement audit and research phase produces actionable findings within the first cycle. Test results depend on your traffic and the power calculation, which we share before committing to a programme.

Fine. We often run experimentation alongside an agency handling acquisition or content, and the analytics audit — deduplicated events, correct attribution, a funnel that reconciles with Shopify's own reporting — tends to benefit both teams.

Want conversion work that reports its losses?

Tell us your traffic and conversion rate. We will tell you honestly whether you can run a testing programme, or whether the money is better spent elsewhere.