Skip to content
CRO

Last updated

How to Choose a CRO Agency: A 2026 Checklist

By Belen Crespo
How to Choose a CRO Agency: A 2026 Checklist

The short version: Choose a CRO agency on five criteria. Check test velocity, statistical rigor, research depth, production, and reporting. Ask one direct question about each area on the first call. Ignore the logo wall and headline win rate. A high win rate can reflect weak stopping rules or low-risk tests.

A logo wall shows who signed a contract. It does not prove that tests reached a valid result. This checklist helps you assess operating discipline in one conversation.

The Remarkable runs a conversion optimization practice, so we are one of the agencies you might score. Use the same criteria on us and our competitors.

How Do You Choose a CRO Agency?

Score every candidate on five criteria. They are test velocity, statistical rigor, research depth, in-house production, and reporting focus. Ask one question about each before the second call.

Use the question column during the first sales call. Specific answers reveal more than a tool list or case-study headline.

Ask for one recent example after each answer. The agency should explain the decision, method, and result without exposing client secrets. Concrete operating details are harder to fake than polished principles.

CriterionWhy it mattersQuestion to ask
Test velocityLearning compounds with shipped experimentsHow many tests did you ship last month?
Statistical rigorPeeking at results invents wins that vanishWhat threshold and sample size do you set first?
Research depthTests without research are guesses at scaleWhat research preceded your last hypothesis?
Design and devTests stall when they wait on your engineersDo you build the variants or do we?
Reporting focusConversion rate alone hides revenue per visitorWhat number leads your monthly report?

The rest of this post takes each criterion one at a time, with what a strong answer sounds like.

How Much Should a CRO Agency Test, and Does Statistical Rigor Matter?

Statistical rigor matters more than any other criterion. A disciplined agency sets the threshold and sample requirement before launch. It does not call a winner as soon as the dashboard looks positive.

VWO’s statistics roundup, citing Enterprise Apps Today, says 52.8% of CRO professionals lack a standard stopping point. Treat that figure as secondary reporting, not a fresh VWO study. Still, the question matters: ask how the agency decides when a test ends.

A strong answer names the method, threshold, sample requirement, and business-cycle rule. A weak answer is “when it looks clearly better.” Repeated peeking can increase false-positive risk. Sequential methods handle monitoring differently.

Ask what happens when traffic cannot support the planned test. A disciplined team may combine variants, extend the run, or choose another research method. It should not lower the standard after seeing an early result.

Velocity also needs a denominator. Four tests may be strong for a complex product flow and weak for a high-traffic landing page. Compare shipped tests with available traffic, engineering effort, and required sample sizes.

More tests help only when rigor holds. Ask how many shipped last month and how many reached the required sample. High velocity without a stopping rule is fast guessing.

What Research Methods Should a CRO Agency Use?

A strong CRO agency uses qualitative and quantitative research before writing a hypothesis. Funnel data shows where users leave. Interviews, surveys, and recordings help explain why. A/B testing without research turns opinion into a larger experiment.

Ask how the team formed its latest hypothesis. “We thought the headline was weak” is a guess. A stronger answer connects observed behavior to a specific test.

Research takes time, and the scope should include it. A proposal that funds only test production may skip diagnosis. Check the promised research hours and methods.

Ask how the agency ranks hypotheses. A useful process considers expected impact, supporting evidence, effort, and traffic. The exact scoring model matters less than a visible reason for choosing one test over another.

Research should continue after launch. Flat and losing tests create evidence for the next hypothesis. The agency should store those lessons instead of treating each experiment as an isolated campaign.

Want to see what disciplined CRO looks like?

Our conversion optimization engagement leads with research, sets significance thresholds before launch, and reports on revenue per visitor. Score us against this checklist yourself.

Book a Free Strategy Call

Should a CRO Agency Bring Its Own Design and Development?

Usually. A testing program can stall when approved experiments wait on your engineers. An agency that builds its own variants controls more of the schedule. An agency that sends tickets depends on your roadmap.

Ask who builds each test variant. If your team owns production, reserve capacity before the engagement starts. Otherwise, the first crowded product sprint can delay the testing plan.

Also ask who reviews quality before launch. The team should test analytics, devices, browsers, and important user states. A broken variant can waste the sample and create a false business conclusion.

A strategy-only agency can work when dedicated in-house designers and developers are available. Confirm their capacity and service-level expectations in writing. Do not assume spare production time exists.

How Should a CRO Agency Report Results?

A strong CRO agency reports on revenue per visitor, not conversion rate alone. A discount can lift checkout rate while reducing average order value. Revenue per visitor shows both effects.

Ask what leads the monthly report. If the answer is conversion rate, ask about revenue per visitor and order value. The reporting model should match your business economics.

For lead-generation businesses, revenue per visitor may arrive too late. Use qualified-lead rate, pipeline per visitor, or another downstream measure. The key is connecting experiments to value, not choosing one universal metric.

Watch how the agency frames wins. An unusually high win rate may reflect early stopping or low-risk ideas. Healthy programs include flat and losing tests, then record what each result changes.

How Do You Make the Final Call?

Score each candidate on all five criteria. Give statistical rigor and reporting the most weight. First, check whether you are at the right stage. When to hire a CRO agency covers traffic and readiness. Then compare pricing with what a CRO agency costs.

The right agency produces results that remain credible three months after launch.

New to the discipline itself? Start with what is conversion rate optimization.

Want to run this checklist on us in real time? Book a Free Strategy Call and ask all five questions. We will answer every one without a slide.

Like this? Get the next one.

Short emails. New posts as they ship.

B
Belen Crespo

Growth Strategist

Focused on helping B2B companies build scalable acquisition systems.

Frequently Asked Questions

How do you choose a CRO agency?

Evaluate a CRO agency on five criteria: test velocity (how many experiments they ship a month), statistical rigor (do they set a significance threshold and sample size before launch, and never peek), research depth (qualitative and quantitative methods, not only A/B tests), whether they bring design and development, and what their reporting leads with. The strongest signal is reporting on revenue per visitor instead of conversion-rate alone.

What questions should I ask a CRO agency before hiring?

Ask one question per criterion. How many tests did you ship last month? What significance threshold and sample size do you set before a test goes live? What research did you run before writing the last hypothesis? Do you bring your own designers and developers or wait on ours? What number leads your monthly report? Specific questions expose whether a team runs disciplined experiments or only tests button colors.

What is a good CRO win rate?

There is no universal good win rate, and an agency that brags about a high one is often peeking at results or testing only safe changes. Healthy experimentation produces plenty of flat and losing tests, because those still teach you what your buyers ignore. Judge an agency on rigor and learning velocity, not a headline win-rate percentage that is easy to inflate.

Get the next post in your inbox

I write about growth, AI performance creative, and what's actually working in 2026. New posts when I have something real to say.

Or book a strategy call →