Metric agreed first
Named before the work starts, so nobody relitigates what counted once the result is in.
RESEARCH, HYPOTHESIS, TEST, SHIP
Research, hypothesis, test, ship — a continuous programme rather than a one-off redesign. We agree which numbers we are moving before any design work starts, and every test either ships or is documented as a dead end.
IN SHORT
Every store we audit already contains the answer to what is wrong with it. Session recordings show people hunting for delivery information. On-site search is full of terms the navigation does not serve. Support tickets repeat the same three questions the product page failed to answer.
None of that requires a test. It requires someone to read it. A surprising amount of what gets sold as CRO is really just looking at what you already have — and then shipping the obvious fix rather than running a four-week experiment to confirm that customers would like to know when their order arrives.
Testing earns its place for the genuinely uncertain decisions, where reasonable people disagree and the stakes justify the wait.
Six stages, repeating monthly. Stage one is where most of the value is.
Analytics, session recordings, on-site search, checkout drop-off and support tickets. Most stores already contain the answer to what is wrong — it is just unread.
Each idea written as a claim with evidence behind it and a metric attached: what we think is happening, why, what we will change, and what should move.
Scored on expected impact against build cost. The scoring is arithmetic, which makes the running order a calculation rather than an argument.
Built in theme code. Where traffic supports a split test we run one for at least two business cycles; where it does not, we ship the better-evidenced version and measure before and after honestly.
Winners ship permanently. Losers go into the backlog with what we learned, so the same idea does not get re-proposed every year by someone new.
What shipped, what moved, what was a dead end, what is next. If a month produced nothing, the review says so.
You need roughly 10,000 sessions and a few hundred conversions per variant per month for a test to conclude within two to four weeks. That is the cadence that makes a programme work.
Below it, tests take months, and the temptation to call them early becomes overwhelming. A variant that is 20% up on day three is usually noise; the same variant is often flat by day twenty-one. More bad decisions come from stopping early than from any other testing mistake.
If your traffic is under the threshold, we will tell you, and we will run a research-led programme instead. That is a smaller retainer, which is precisely why most agencies do not mention it.
Named before the work starts, so nobody relitigates what counted once the result is in.
Failed tests are documented with what we learned. They are the cheapest knowledge you own.
If a test lifted add-to-cart and dropped revenue per session, the report says both.
Winners become permanent theme code, not a rule inside a testing tool you pay for forever.
CRO runs best alongside speed work — a faster page lifts every test that follows it — and is usually delivered inside a development retainer.
Roughly 10,000 sessions and a few hundred conversions a month per variant to reach significance in a sensible window. Below that we optimise from research and qualitative evidence rather than split tests.
At least two full business cycles — usually two to four weeks — so the result is not a reflection of one payday or one campaign.
It varies so widely by category, price point and traffic source that a benchmark is close to useless — a considered-purchase furniture brand and an impulse accessories brand should not be compared. The number worth tracking is your own, segmented by device and traffic source, against itself over time.
Whatever the research points at, but in practice the highest-value areas are usually checkout and cart drop-off, product page clarity around delivery and returns, and mobile specifically, since that is where most traffic and most friction sit.
No, and anyone who does is either guessing or planning to pick a favourable metric afterwards. What we commit to is a documented programme, agreed metrics and honest reporting, including the tests that did not work.
App bloat stripped, scripts deferred, LCP and CLS measured on real devices rather than a lab score.
A senior engineer walks your storefront, theme and checkout and sends back a prioritised list of fixes.
Campaign and launch pages built to a section library, so marketing ships without a dev queue.
NEXT STEP
Storefront, theme performance and checkout reviewed by a senior engineer. No pitch deck, no obligation.