
What Is a CRO Hypothesis?
A hypothesis is not a wishlist. Here is the shape we use: belief, change, metric, timeframe, and why we test one at a time.
If you are reading a Convfetti proposal and wondering what we will actually test, this is the answer.
A hypothesis is not a guess. It is not a wishlist. And it is definitely not let's test the button color and see what happens. I'll say it plainly. Most stores that call something a hypothesis are just describing a change. A real hypothesis has a shape, a reason, and a way to know if it was right.
Here is how we think about it at Convfetti, and why it matters for your store.
What a hypothesis is not
Founders usually confuse three things with a hypothesis. I get it. The language is sloppy everywhere.
First is a wishlist. We should try a new homepage, new PDP, new checkout, plus a faster theme. That is a list of things to build. There is no belief behind it. No metric. No timeline. You could do all of it and still not know what worked.
Second is a single-element guess. Test the button color. Sure, you can change a color. But why? What do you believe about the buyer that makes that color the lever? Without that, you are not testing a hypothesis. You are testing a hunch and hoping the number moves.
Third is best-practice cargo. Everyone says sticky add to cart helps, so we should add it. Best practices are not beliefs. They are patterns that worked somewhere else, for someone else's buyer, in a different context. Sometimes they help here. Sometimes they hurt. You do not know until you frame it as a belief about your buyer.
None of these is a hypothesis. All three feel productive. That is the trap.
Anatomy: belief → change → metric → timeframe
A real hypothesis has four parts. Always in this order.
One, belief. What we think is true about the buyer right now. Not about the site. About the person. Example shape: visitors do not understand what makes this product different in the first six seconds, so they scroll past the value.
Two, change. What we will do on the page to address that belief. One focused change that follows from the belief, not three things at once.
Three, metric. What we will measure to know if the belief was right. We pick one primary metric. Could be add to cart rate. Could be checkout start rate. Could be purchase conversion. But we pick one. Not five.
Four, timeframe. When we will judge it. For us that is one sprint, which is three to four weeks. Not one day. Not whenever we feel like it.
Here is the skeleton we use. This is Example structure, not a past result:
We believe [buyers do not see the offer clearly above the fold], so we will [make one change to the first screen that states the offer in plain language], and we expect [add to cart rate] to move within [one sprint].
That is it. Belief, change, metric, timeframe. If any piece is missing, it is not a hypothesis yet.
You will notice the funnel language in that example: Example: add to cart → checkout → purchase. We use that shorthand a lot because it maps to where buyers actually drop. We will label it every time so you know when we are illustrating versus reporting.
Where hypotheses come from (baseline, funnel, sessions)
We do not pull hypotheses from opinion. We pull them from three sources, in order.
Baseline comes first. That is the before picture. Where the store leaks right now, by funnel step, by device, by traffic source. If your mobile add to cart is half your desktop, that tells us something. If paid traffic bounces faster than organic, that tells us something else. Baseline is not a judgment. It is a map. We do not skip it even when the brief feels obvious. I have seen obvious hypotheses fail because baseline showed the real leak was one step earlier.
Funnel comes second. Where the drop is steepest. We often frame this as Example: Meta → landing → PDP → checkout, illustrative. That sequence helps because it names the handoff. Meta promises something, the landing page has to keep that promise, the PDP has to close the gap, checkout has to not leak. When the drop lives between Meta and landing, the hypothesis lives there too. Not on the checkout.
Sessions come third. What people actually do. Hesitations. Back and forth. Rage clicks. Scrolling past the thing you thought was clear. We do not claim we watch every session or that we have a tool no one else has. We look at enough to see the pattern, then we tie that pattern back to baseline and funnel.
If a hypothesis cannot be traced to at least one of these three, we do not run it. That is the filter.
A note on numbers. If we name a baseline number in a readout, we will label it. If it is a placeholder for shape, it will say TBD, insert approved baseline example when available. If it is an illustrative funnel, it will say Example: Meta spend vs site conversion split, illustrative. You will always know which is which.
Why one hypothesis per sprint
You can change two things and get a lift. You will not know why. You can change three things and get nothing. You will not know what to fix next. That is the cost.
One hypothesis per sprint is learnable. Whether the number moves or not, you learn something you can use next time. Three hypotheses at once gives you a result you cannot explain and a backlog you cannot prioritize. The learning is muddled.
One at a time also compounds. Each readout makes the next hypothesis sharper. We go from buyers do not get the offer to buyers get the offer but do not trust delivery to buyers trust delivery but hesitate at the guarantee. That sharpening only happens if you isolate the signal.
Does that mean we only change one element on the page? No. We change what is needed to test the hypothesis. If the belief is about clarity above the fold, the change might touch headline, subhead, and visual together. But it is still one hypothesis, to one metric, in one timeframe. You judge the belief, not the element count.
A sprint is three to four weeks. We will cover that cadence in detail in how sprints work. For now, just know this: one hypothesis per sprint is how we keep the learning clean.
How a hypothesis connects to the readout
The hypothesis sets what the readout will judge. That is its whole job.
At the end of the sprint, the readout answers three things. Did the metric we tied to the hypothesis move in the timeframe. What else moved or did not. And what is the next hypothesis.
That loop is always the same: hypothesis → implement → measure → readout. The hypothesis is the start, the readout is the accounting, and the next hypothesis is the output.
We will sometimes say a result is directional. That is not hand waving. It means the signal points the same way across related steps or segments but is not at a threshold we would call conclusive yet. It is still useful for picking Sprint 2. We define that properly in what you get in a readout, so you are not left guessing what we mean by it.
What you should take from this page. A hypothesis is a belief about your buyer, turned into one change, judged by one metric, in one timeframe. It comes from baseline, funnel, and sessions. We run one at a time so the learning is clean and the next sprint gets better. The readout is where we prove it or learn from it.
Next, see how sprints work for the full loop, or jump to what you get in a readout to see how we report it.
And if you are comparing plans, see engagement plans on our pricing page. One line, no pitch. Just context for where these sprints live.