Pages

Friday, 24 July 2026

Part 3: How to Decide Which CRO Test to Run First


Because the loudest idea in the room is not always the best experiment.

In the last post, I wrote about the main experiment types: A/A, A/B, Split URL, and multivariate testing. I also introduced the formula that quietly changed how I think about experimentation:

Success = Chance × Frequency

Once that clicked, I assumed the next step would be simple.

Pick an idea > Run a test > Measure it > Repeat.

Then I actually tried to do it. Because once people around you start thinking in CRO terms, ideas come from everywhere.

The designer wants to test the hero image. The founder wants to rewrite the pricing page because “it does not feel premium enough.” Someone from sales heard that a competitor changed their CTA and now wants to copy it immediately. Someone read a blog at midnight and suddenly wants to add a chatbot, a quiz, a floating banner, and possibly a dancing mascot.

Before you know it, your experiment backlog looks like a Jira board after a chaotic sprint planning session.

  • Lots of tickets.
  • Very little structure.
  • Everyone convinced their idea is the one that will change the company’s destiny.

This is where prioritization comes in. Not because prioritization makes you smarter, but because it protects the experimentation program from becoming purely opinion-driven. And in CRO, opinions can multiply faster than credit card debts.

The Hidden Cost of Running the Wrong Test

At first, “let’s just test this” sounds harmless. CRO is about experimentation, right? Just run everything. But every test has a cost.

Even when the testing tool makes setup easy, a test still consumes traffic, time, design effort, analysis attention, and usually a couple of weeks of your roadmap. If your site has enough traffic to run two meaningful tests this month, and your backlog has twenty ideas, running the wrong two means you delayed the better ones. That is not a small cost.

This felt immediately familiar from Devops work. You do not deploy every feature request just because someone suggested it in a meeting with confidence and a nice slide deck. You triage. You look at impact, risk, effort, dependencies, and what the system is actually telling you before you start making changes.

CRO needs the same discipline. An idea is not valuable because it sounds smart in a meeting. It becomes valuable when it is connected to evidence and scored against other ideas in the backlog.

The HiPPO Problem

Before getting into frameworks, there is a force you need to understand. It quietly kills many CRO programs. It is called the HiPPO: the Highest Paid Person’s Opinion. Every team has one. In fact there is a rude saying about opinions. “Opinions are like…” (well… Google it)

The senior leader who says, “I think the button should be green,” and suddenly the next sprint is about button colors. The founder who saw a competitor’s landing page and wants to rebuild yours to match it by Thursday. The VP who has “a feeling” about the checkout flow. The stakeholder who says, “Can we just test it?” in the same tone people use before creating six months of technical debt.

To be fair, these inputs are not useless. People close to the business often have genuine intuition worth exploring. Sales teams hear objections. Support teams hear confusion. Founders understand positioning. Product teams know what users struggle with. The problem is not the idea. The problem is when the idea skips the queue without being scored.

In DevOps, we do not prioritize incident response based on who is shouting loudest in Slack. At least, we should not. We prioritize based on system impact and what the metrics are actually saying.

CRO needs the same structure. A prioritization framework is your polite defense against the HiPPO.

You are not saying: “No, your idea is bad.”

You are saying: “Great, let’s score it the same way we score everything else. If it ranks high, it goes first.”

That small shift changes the conversation. The framework depersonalizes the argument. The number does the heavy lifting. Very convenient, because numbers do not get awkward in meetings.

Start With the Funnel, Not the Idea

Before scoring individual test ideas, figure out where to look first. Not every page deserves equal attention. A page with very little traffic will take forever to produce statistically meaningful results. A page with high traffic but no meaningful business impact may still not be worth optimizing ahead of a page that sits directly in the conversion path.

Press enter or click to view image in full size

So before asking: “What should we change?” ask, “Where are we losing people?”, and then. Map your funnel.

For an e-commerce site, it might look like this:

Landing page > Product page > Cart > Checkout > Purchase

For SaaS, it might look like this:

Homepage > Pricing page > Signup > Activation > Paid conversion

For lead generation, it might look like this:

Landing page > Form view > Form start > Form submit > Qualified lead

Then look at where the biggest drops are happening. If 70% of users are abandoning at checkout, testing a homepage headline is probably not the first move.

  • It might be interesting.
  • It might even win.
  • But it is probably not where the biggest leak is.

Analogy in DevOps: If the database is causing 90% of your latency, shaving 20ms off a frontend asset is not the priority. Sure, the frontend improvement may look nice in a performance report. But the database is still sitting there like an overloaded elephant in the server room.

Find the bottleneck first. Then fix it. In CRO, high-traffic pages close to revenue, signup, checkout, or lead generation are your hot paths. Start there.

Small disclaimer: Read at your own risk. This may make you hungry

The Two Frameworks: ICE and PIE

Once you know which pages to focus on, you need a way to score the ideas sitting in your backlog. Two common prioritization frameworks in CRO are ICE and PIE. Both are simple, both use overlapping dimensions and both are easier to remember if you stop thinking about frameworks and start thinking about food. Which, honestly, improves most business concepts. ;)

ICE: Ordering Food at Midnight

ICE stands for:

  • Impact
  • Confidence
  • Ease

Think of ICE as ordering food on delivery app at 11 pm. You are hungry. You are tired. You are already in bed. This is not the time for adventure. You are not trying to discover a hidden culinary gem. You are trying to avoid sleeping angry. You see a biryani place.

Before you tap “Order,” three thoughts happen automatically.

First:

“How good is this actually going to be?”

If it is bad biryani, why bother?

That is Impact.

Second:

“Can I trust this place right now?”

Will it arrive hot, on time, and with raita? Or will it arrive cold, late, and missing the one thing that emotionally holds the whole meal together? You check the rating, reviews, delivery time, and whether someone recently complained that they received only rice and sadness.

That is Confidence.

Third:

“How easy is this?”

Is the restaurant open? Is delivery available? Will it reach in 25 minutes, or will the app say “arriving soon” until you lose hope? Will it ask you to verify your phone number again even though it has known you for five years?

That is Ease.

ICE is practical and slightly paranoid. It wants evidence before committing.

In CRO terms:

Impact asks: If this test wins, how much could it move the business?
Confidence asks: How sure are we that this is a real problem worth testing?
Ease asks: How difficult will this be to build, QA, launch, and analyze?

Score each dimension from 1 to 5. Then average the three. The highest-scoring ideas usually go first. ICE works especially well when you already have research data: heatmaps, recordings, analytics, surveys, support tickets, or past experiment learnings. Low confidence should lower the score, no matter how exciting the idea sounds.

Because “I saw this on a competitor site” is not research. It is browsing with ambition.

PIE: Choosing What to Order When You Have Options

PIE stands for:

  • Potential
  • Importance
  • Ease

Now stay in the same food delivery app. You are still hungry. It is still 11 pm. You still should have eaten earlier like a responsible adult, but here we are. This time, you are not evaluating one biryani place. You are deciding what category of food deserves your attention first.

Biryani? Pizza? Rolls? South Indian? Chinese? That one “healthy bowl” option you add to cart and then remove after seeing the price? PIE helps you decide where the biggest opportunity is.

First:

“How good could this get?”

Could this be the kind of meal that fixes your mood, your evening, and possibly your belief in humanity?

That is Potential.

Second:

“How important is this meal?”

Is this a casual snack, or have you skipped dinner and now your stomach is sending production-level alerts?

That is Importance.

Third:

“How easy is it to get?”

Is it nearby and deliverable in 25 minutes, or is the restaurant far away, closing soon, and likely to cancel after making you wait?

That is Ease.

PIE is more optimistic than ICE.

ICE asks, “Can I trust this specific choice?”

PIE asks, “Where is the biggest opportunity?”

In CRO terms:

Potential asks: How much improvement could this page or funnel step realistically have?
Importance asks: How valuable is this page or funnel step to the business?
Ease asks: How simple or difficult is the test to implement?

PIE is useful when you are comparing broader areas of opportunity.

For example, if you are deciding whether to focus on the homepage, pricing page, checkout, signup form, or onboarding flow, PIE can help you decide where the biggest opportunity may be.

So, in food delivery terms:

ICE helps you decide whether to trust one restaurant.
PIE helps you decide which food category is worth exploring first.

ICE vs PIE: The One-Line Difference

Press enter or click to view image in full size

Here is the easiest way I remember them:

  1. ICE helps you decide whether to trust one restaurant. PIE helps you decide which food category is worth exploring first.
  2. ICE is about confidence. Can I trust this specific option to deliver? PIE is about opportunity. Which area has the biggest upside if I focus there? Same food app. Different thinking.
  3. With ICE, you are asking: “Should I order from this biryani place?” With PIE, you are asking: “Should I even be looking at biryani first, or is pizza, rolls, South Indian, or Chinese the better opportunity tonight?”
  4. Both end with Ease, because nobody wants unnecessary complications at 11 pm. Especially not when they are hungry and the app is already saying, “Restaurant is closing soon,” like it is adding pressure to your life.

Which Framework Should You Use?

Use ICE when your team already has research data and wants to prioritize based on evidence. ICE is useful when you are evaluating a specific hypothesis. For example:

“Should we move the CTA above the fold on the pricing page?”

Learn about Medium’s values

You already have heatmap data. You know users are not scrolling far enough. You want to know whether this specific idea deserves to be tested next. That is ICE.

Use PIE when you are earlier in the process and still deciding where the biggest opportunity is.

For example:

“Should we focus on the homepage, pricing page, checkout flow, signup form, or onboarding?”

That is PIE.

At Convert.com, the Compass feature includes both PIE and ICE inside the hypothesis builder. When you create a new hypothesis, you can choose a framework, score each dimension on a scale of 1 to 5, and use that score to sort your backlog.

That helps because the score becomes visible to everyone. It is much harder to push your pet idea to the top when the number is sitting right there on the screen, judging you quietly. A visible score is like the delivery rating on a food app. You can still order from the suspicious restaurant. But now everyone can see that it has 2.3 stars and seven reviews mentioning “never again.”

Turn Ideas Into Hypotheses Before You Score Them

Press enter or click to view image in full size

One thing both frameworks require is that your ideas are written properly before scoring.

“Change the button color” is not a hypothesis.

“Move the CTA above the fold” is not a hypothesis.

“Make the page look more premium” is definitely not a hypothesis.

That is a mood.

A proper hypothesis looks like this:

If I change X, then Y will happen, because Z.

The “because” is the most important part. Without it, you are not testing an assumption. You are just making a change and hoping the graph goes up, which is basically the CRO version of deploying to production and whispering, “Please work.”

Or, to stay with the food delivery analogy, it is like ordering from a restaurant with no ratings because the photo looked nice. Could it be amazing? Maybe. Could it be regret in a plastic container? Also maybe.

For example:

If we move the main CTA above the fold on the pricing page, more visitors will start the demo request flow, because heatmaps show most users are not scrolling far enough to see the current CTA.

Now you have something worth scoring.

You know:

  • The page
  • The change
  • The expected outcome
  • The evidence behind it

Even if the test loses, you learn something specific. A backlog full of vague ideas becomes a dumping ground. A backlog full of hypotheses becomes a learning system.

Or in food terms:

“Order something nice” is not useful.

“Order from a highly rated nearby restaurant because delivery time is low and reviews mention fresh food” is much better.

Specificity saves you from bad decisions. And occasionally, bad paneer.

The One Thing ICE and PIE Get Wrong

ICE and PIE are useful, but they share a blind spot. They often treat every idea as if it exists in a vacuum. But it does not.

Every idea exists inside the history of your previous tests. If you have run five tests on your pricing page and four of them lost, your confidence score for the next pricing-page idea should reflect that history. Something on that page may be resisting your assumptions.

  • Maybe the traffic quality is different from what you expected.
  • Maybe the real problem is upstream in the funnel.
  • Maybe people already understand the pricing page, and the actual friction is in signup.
  • Maybe your pricing page is innocent and has been framed by bad hypotheses.

The same applies to food delivery. If you have ordered from the same restaurant four times and three times they forgot the raita, you should not treat the fifth order like a fresh mystery.

You have history. Use it. Past experiment results should feed into how you score new ideas. They should not sit inside a spreadsheet named something like: final_results_v3_really_final_updated_NEW.xlsx

This is why documentation matters as much as the framework itself.

After every experiment, write down:

  • The hypothesis
  • The result
  • The confidence level
  • The primary metric
  • The key learning
  • What it suggests about the next test

Even a losing test is useful if it tells you where not to look. Without this history, your backlog resets to zero every quarter. With it, your scoring gets more accurate over time because you are building a picture of what your specific audience actually responds to. A CRO backlog without documentation is like a food app that forgets every bad order you ever had. You keep making the same mistake. And somehow, the raita is still missing.

What to Do When Everything Scores the Same

Sometimes three ideas all score 4.2 and suddenly you are back in the same argument. This is normal. Frameworks help, but they do not eliminate judgment. A few tiebreakers can help. Pick the page with the most traffic. More visitors usually means faster results and less time waiting for statistical significance.

In food delivery terms, this is like choosing the restaurant with enough recent orders and reviews. If only two people have ordered from it since 2021, the rating may not tell you much.

Pick the test that is fastest to build. A high-scoring idea that takes three weeks of development work may be worth doing after a slightly lower-scoring idea you can ship in a day. Remember the frequency side of the success formula.

This is the difference between ordering something that arrives in 25 minutes and choosing a dish that says “preparation time: 95 minutes.”

Technically, it may be great. Emotionally, you may not survive. Pick the test that teaches you the most regardless of outcome. Learning value is underrated. A test that answers a fundamental question about your audience is worth more than a test that optimizes something already working reasonably well.

In other words, do not only ask:

“What might win?”

Also ask:

“What will we learn even if this loses?”

That question saves a lot of time. And occasionally, a lot of ego.

Define Success Before You Launch, Not After

One mistake that shows up in beginner backlogs is choosing the primary metric after the test has already started. That is dangerous. If you look at enough metrics, something will always appear to have improved.

  • Clicks went up.
  • Scroll depth changed.
  • Form starts increased.
  • Time on page looked different.
  • Someone in Nebraska spent 14 minutes reading the footer.

But what was the experiment actually supposed to improve? Every test needs a primary metric before launch.

  • Demo requests.
  • Purchases.
  • Trial signups.
  • Form completions.
  • Revenue per visitor.

Pick one. Commit to it. Let that metric make the call. This is like deciding what “good food” means before ordering. Are you optimizing for taste? Delivery speed? Price? Portion size? Not waking up with regret? If you decide after the food arrives, you can justify anything.

“The biryani was cold, but the container was sturdy.”

That is not success. That is coping. You can still watch secondary metrics for learning, but the decision should rest on the primary one. And alongside your primary metric, define your guardrail metrics. These are the things you are not directly trying to improve, but cannot afford to break. For example:

  • If you are optimizing demo signups, lead quality should not drop.
  • If you are optimizing checkout completion, average order value should not fall.
  • If you are optimizing form completion, spam submissions should not explode.
  • If you are optimizing clicks, downstream conversions should not suffer.

This is the CRO version of SLOs. You define what success looks like and what failure looks like before you are emotionally invested in the outcome. A test that increases clicks while damaging qualified leads is not a win. It is noise wearing a party hat.

Or in food delivery terms: A restaurant that delivers in 12 minutes but sends the wrong order did not “win on speed.” It failed dinner.

A Practical First Backlog

If you are starting from zero, here is the workflow I would follow.

  • Map your funnel and identify the biggest drop-off point.
  • Use heatmaps, session recordings, analytics, surveys, and any sales or support feedback you can access.
  • Write down every test idea without judging too early.
  • Convert each idea into a proper hypothesis using the “If X, then Y, because Z” format.
  • Score each one using ICE or PIE.
  • Pick one high-scoring, reasonable-effort test.
  • Define the primary metric and guardrail metrics before launch.
  • Run the test.
  • Document the result and the learning.
  • Then re-score the backlog based on what you now know.

That last step is what most teams skip. And it is what separates a backlog that compounds from one that stagnates. A CRO backlog should not be a parking lot where ideas go to quietly disappear. It should be more like a living system.

Or, to stay on brand, a well-maintained CI/CD pipeline for learning. Every test should make the next decision smarter. Just like every food order teaches you something.

  • Restaurant is reliable.
  • Delivery is slow on weekends.
  • The “extra spicy” option is not a personality test you need to pass.
  • Never trust a place where every menu item has the same stock photo.

Learning compounds when you write it down.

The Real Takeaway

Prioritization is not about finding the cleverest idea. It is not about the biggest redesign. It is not about the most senior opinion in the room. It is about running the tests that have the best combination of meaningful impact, strong evidence, reasonable effort, and clear learning potential. Done consistently, this turns CRO from a collection of random experiments into a learning system. You find the biggest leaks, collect evidence, write hypotheses, score ideas, run the most valuable tests first, document what happened and then repeat.

The DevOps parallel is hard to miss. Good infrastructure teams do not just ship more changes. They improve how they decide what to change, how they measure it, and how they learn from it.

Good CRO teams do exactly the same thing. And the food delivery parallel is also hard to miss. The goal is not to order everything on the menu. The goal is to make better choices each time, based on what you know, what matters, what is realistic, and what you learned last time.

That is prioritization.

In the next post, I want to go deeper into reading results without fooling yourself: statistical significance, confirmation bias, and why the most dangerous moment in CRO is when a test is almost winning.

I work with Convert.com and write about the intersection of DevOps and experimentation. If you are building your first CRO backlog or trying to move your team away from opinion-driven testing, drop a comment or reach out. Happy to compare notes.

Wednesday, 24 June 2026

From a zip file in 2020 to an AI-powered wiki in 2026. Here's what happened in between

 AI generated image

Someone on the team once asked about a feature deprecated back in 2020.

45 minutes of digging through backups, old slack threads, a google doc nobody had touched in years, and the answer eventually showed up inside a zip file. Like a digital archaeological dig. Except less cool and more frustrating.

That moment made it clear to me that something needed to change. This post is about what the team figured out together.

Part 1: An old hero called MDWiki

Before getting to the AI part, here's a tool that ended up being central to the whole thing.

Years ago, there was a situation many people would remember. I was in a pilot batch at a new office location for an MNC. Fresh start, no internal tools, no documentation platform, no proper handover process. Knowledge transfer from HQ was slow, questions were piling up, answers were scattered everywhere.

MDWiki turned out to be the answer. Its an old tool, almost ancient by tech standards, but brilliant in its simplicity. plain markdown files in a folder, and it renders a full documentation site. no database, no server setup, no hosting complexity. just files.

Hosting it on OneDrive meant the whole company could sync it locally. suddenly every team had a shared wiki that worked offline, updated automatically when anyone saved a file, and needed zero infrastructure to maintain. Lightweight, fast, and it worked. sometimes old tools solve new problems better than anything shiny.

Part 2: The scale problem every growing team eventually hits

Fast forward to now. At a certain point, documentation at our company had simply outgrown the way it was being managed. Knowledge contributed by multiple team members over years. text files, images, screenshots, video recordings, architecture diagrams, slack conversations, meeting notes, product threads. Rich, honest knowledge built up over time by people who genuinely cared.

But here is what happens at scale. When knowledge lives across too many places, it becomes hard to navigate. some topics had multiple versions, each slightly evolved from the last. naturally, over time, some overlap crept in. This is not a people problem. Its a growth problem. every team that ships consistently runs into this. knowledge compounds faster than anyone can manually organise it.

And when someone moves on, a little undocumented context moves with them. not because they were careless, but because there was no system to capture it before it walked out the door. that gap needed closing, and honestly it was on us as a team to fix it sooner.

Part 3: Andrej Karpathy walked in (virtually)

That's when we came across Andrej Karpathy's LLM-Wiki idea. if you don't know Karpathy, he was a founding member of OpenAI and Tesla's AI director. when he writes something, the internet pays attention.

His idea is elegant. Stop treating your knowledge base as a place humans write into. treat it like a git repo. let the LLM own the writing layer. humans write exactly one file by hand, a schema file that defines the rules, structure, and conventions. Everything else? the LLM reads raw sources and builds the wiki, maintaining it as new content comes in.

The wiki becomes a persistent, compounding artifact. it gets richer with every source added. cross references are already there. contradictions get flagged. the synthesis reflects everything that has been added, not just the last thing someone remembered to update. Brought it to the team. We decided it was worth trying.

Part 4: What we built together

We combined LLM-Wiki with MDWiki and added a CI/CD layer on top. Here is how the whole thing works.

Three layers.

Raw folder. all source material goes here. old docs, slack exports, meeting notes, screenshots, video transcripts, architecture diagrams. untouched, immutable. source of truth.

CLAUDE.md. the one file written by hand. it teaches Claude how MDWiki works, what the folder structure looks like, title formats, how to handle overlapping topics, what goes where, and when to flag something for human review instead of auto-placing it. the rulebook for the entire system.

The wiki itself. Claude reads whatever lands in the raw folder and writes, organises, and cross-links the markdown files in MDWiki format. every document lands in the right place, correctly formatted, connected to related pages.

And the most important part, the human review layer. Any team member can drop a document into the raw folder and raise a pull request on GitHub. a reviewer steps in before anything gets merged. they check:

  • Is this information still accurate or has the process changed?
  • Does this conflict with something already in the wiki?
  • Is there any sensitive information, credentials, or customer data that shouldn't be here?
  • Is this still relevant or does it describe something no longer in use?
  • Does the framing match our internal style and audience?

Once the PR is approved, Claude adds the knowledge to the wiki, formats it for MDWiki, cross-links it with related content, and updates the index. A GitHub Actions pipeline deploys the whole thing to Google Firebase. every approved merge triggers a deployment. the team sees the updated wiki within minutes, no manual publishing needed.

Part 5: The Guardrail reminder

This one bears repeating, and it came from a hard conversation within the team early on. Claude has write access to the knowledge base. and we made sure it cannot merge its own PRs.

Not because we don't trust the model, but because nobody should. a recent research paper tested 19 LLMs on long document editing workflows and found that even the best models, GPT, Claude, Gemini, silently corrupt an average of 25% of content over time. not crashing, not warning anyone. just quietly getting things wrong as the workflow gets longer.

The human review step is not a formality. its the most load-bearing part of the whole pipeline. we almost skipped it to move faster. Glad we didn't. Automation without a review layer is not a knowledge base. its a very confident source of misinformation.

If you're building something similar or thinking through your own documentation architecture, drop a comment. happy to share more on the setup.

Ref

#Documentation #DevOps #LLMWiki #KnowledgeManagement #GitHub #MDWiki #AIAgents #GitHubActions #Firebase

Note for student readers

If you're in college and hunting for internships, stay with me for one more minute. This entire system works because of one thing: markdown. MDWiki renders it. Claude writes it. GitHub versions it. Firebase serves it. the whole pipeline has markdown at the center.

Markdown is not a niche skill. its the common language of documentation across DevOps, software engineering, product teams, and now AI pipelines. When you write a README.md on GitHub you're not just leaving notes for yourself. you're showing every hiring manager who lands on your profile that you think about the person coming after you. that's rarer than it sounds.

Most students push code and skip the README. don't be most students. Simplest thing you can do this week: go to your best GitHub project and write a proper README.md. not just "this is my project." write what it does, what problem it solves, how to run it, and what you learned building it.

That one file will do more for your internship hunt than a fancy resume template. Markdown today. Opportunity tomorrow.

✏️ drafted with ai assist, because practicing what I preach.

Monday, 22 June 2026

Part 2: The Experiments That Actually Move the Number

 

If you read the first post in this series, you now have a working definition of CRO, a sense of why conversion rate matters, and hopefully a free account on some combination of Clarity, GA4, and Convert.com(sshhh… I work for them).

Maybe you even opened a heatmap, stared at the colors for a while, and thought: “Okay. Now what?”

This post is that now what.

Once you understand the basics of CRO, the next challenge is figuring out what kind of experiment to run. A/A test? A/B test? Split URL test? Multivariate test? They all sound similar at first, but they solve different problems at different stages.

And if you are coming from a technical background, the temptation is to overthink the setup before shipping anything. I know, because that is exactly what I started doing.

So let’s break down the main experiment types, when to use each one, and how to think about them without getting stuck. I will also share a formula from a CRO course I have been working through that quietly changed how I think about the whole thing.

The Four Experiment Types, Explained Without Jargon

1. A/A Testing

Before you trust your first experiment result, run an A/A test.

In an A/A test, both groups see the exact same page. No headline change, no button change, no layout change. Nothing at all. The purpose is not to improve conversions. The purpose is to check whether your measurement setup is actually trustworthy.

If your tool is configured correctly, both groups should show roughly similar conversion rates over time. If they do not, something is off with how traffic is being split, how visitors are being tracked, or how conversions are being recorded.

The DevOps parallel is straightforward. You would not set up Prometheus alerts on a misconfigured scrape target and then make infrastructure decisions based on that data. An A/A test is the same idea: confirm the instrument is working before you start using it to make calls.

A lot of teams skip this because it feels boring. There is no variation, no redesign, no winner to announce. But if your measurement layer is broken, every A/B test after this becomes questionable. You may spend weeks debating a result that was never reliable in the first place.

On Convert.com, setting up an A/A test takes about three minutes. Do it before anything else, and leave it running for at least a week.

Press enter or click to view image in full size

2. A/B Testing

Once your A/A test passes, you are ready to actually experiment.

An A/B test compares two versions of something. Version A is the original, called the Control. Version B is the changed version, called the Variant. You split traffic between the two and measure which one performs better on your chosen metric.

The element you change could be a headline, a CTA button, a hero image, a form layout, a pricing message, anything on the user journey. The constraint is this: change one meaningful thing at a time.

If you change the headline, button color, pricing layout, and image all in the same test, and the variant wins, you have no idea what actually caused the lift. You got a result but no learning.

The DevOps analogy that clicked for me: a clean commit. If one commit touches ten unrelated files and something breaks, debugging is painful. If one commit changes one logical thing, the cause is much easier to trace.

A good A/B test starts with a hypothesis in this format: “If I change X, then Y will happen, because Z.”

For example: “If I change the CTA text from ‘Buy Now’ to ‘Get Your Copy Today,’ then more visitors will click through to checkout, because the new copy feels less transactional and creates a stronger sense of ownership.”

The “because” is the part most people skip, and it is the most important part. Without it, you are not testing a hypothesis. You are just nudging things and hoping the graph goes up.

Press enter or click to view image in full size

The confusion between A/A testing and A/B testing

I was always confused as to why A/A and A/B testing are needed. If you are also in the same state, here is a crude example that may simplify it.

A/A Testing: The Paranoid Caterer

You are catering a wedding. You make one giant pot of biryani. You serve it to Table 1 and Table 2 from the same pot, same quantity, same everything.

Table 1 finishes 95 plates. Table 2 finishes 30 plates.

You have not changed anything. Both tables got identical food. So why is Table 2 eating so little?

You investigate. Turns out the waiter assigned to Table 2 was busy on his phone and forgot to refill the serving bowls.

The biryani was not the problem. The delivery system was broken.

That is an A/A test. Before you start testing new recipes, you confirm that the serving system itself is not lying to you.

A/B Testing: The Sensible Caterer

Same wedding. You make two versions of biryani. Version A has the usual spice level. Version B has a little more pepper because your cousin who went to Bangalore said people there like it spicy now.

You serve Version A to the left side of the hall and Version B to the right side. You count which side finishes faster, asks for seconds more often, and complains less to the mother of the bride.

One change. One comparison. One winner.

You do not also change the raita, the naan, the serving bowl size, and the background music at the same time. Because if the right side loved it, you will never know if it was the pepper or the fact that the DJ on that side was finally playing something decent.

So:

A/B test: “Which version is better?”
A/A test: “Can I trust the test setup?”

3. Split URL Testing

A standard A/B test usually changes something on the same page. The URL stays the same and the testing tool modifies what users see. Split URL testing is different: users are sent to entirely different URLs.

Control: /pricing. Variant: /pricing-v2. Traffic is split between them.

Use this when the change is large enough that modifying the existing page gets messy. Major redesigns, completely different landing page structures, new checkout flows, testing a long-form page against a shorter one. If it is a minor copy tweak, A/B is faster and simpler. If you are rebuilding the page from scratch, Split URL is cleaner.

The infrastructure analogy is blue-green deployment: two versions running simultaneously, traffic routed between them at the experimentation layer instead of the load balancer. Same principle, different domain.

One thing to watch carefully: tracking consistency. Both pages need to fire the same analytics events and conversion tracking correctly. A common mistake is having GA4 or pixel tracking set up on the original page but missing or misconfigured on the variant. When that happens, the variant looks like it is losing when the actual problem is your observability. As DevOps people know well, bad telemetry makes healthy systems look broken.

Press enter or click to view image in full size

4. Multivariate Testing

A/B testing changes one element at a time. Multivariate testing (MVT) changes multiple elements and tests every combination of them simultaneously.

Say you want to test two headlines and two hero images. MVT creates four combinations: Headline 1 + Image A, Headline 1 + Image B, Headline 2 + Image A, Headline 2 + Image B. Traffic is split across all four.

The benefit is finding the best combination without running four separate sequential tests. The downside is traffic dilution: more combinations means each version gets a smaller slice of traffic, which means it takes longer to reach a reliable result.

This is why MVT is better suited for high-traffic pages. On a low-traffic site, running multivariate tests is like spinning up a four-node Kubernetes cluster to host a static HTML page. Technically possible. Practically painful.

Join The Writer's Circle event

If you are just getting started, stick to A/B tests until you have enough traffic and enough process confidence to justify MVT.

Press enter or click to view image in full size

Confused between Split URL and Multivariate?

Here’s the crude example:

Split URL Testing: The Ambitious Caterer

This caterer is not making a variation of the existing menu. He has built two entirely separate stalls on opposite ends of the lawn.

Stall A is the traditional setup: biryani, raita, gulab jamun, done.

Stall B is the “fusion experience”: biryani bowls, deconstructed raita in shot glasses, and gulab jamun on a stick with a QR code to leave a Google review.

Half the guests are directed to Stall A and half to Stall B. You measure which stall has a longer queue, higher plate counts, and fewer people quietly throwing food in the dustbin.

This is a Split URL test. You are not tweaking one element on an existing page. You have built two completely different pages and you are routing real traffic to both to see which one works.

The infrastructure cousin of this is blue-green deployment. Same idea, different catering budget.

Multivariate Testing: The Caterer Who Has Lost the Plot

This person is testing:

  • 2 types of biryani (chicken vs. mutton)
  • 2 types of raita (boondi vs. cucumber)
  • 2 types of bread (naan vs. rumali roti)
  • 2 types of dessert (gulab jamun vs. rasgulla)

That is 16 combinations. You need 16 separate tables, each getting a different combination, and enough guests at each table to statistically prove which combination is the winner.

The wedding has 200 guests. You need at minimum 1,600 guests for the data to mean anything.

You have invited 200.

You wont get that big venue on any of the wedding seasons, but lets just say… you did.

This is multivariate testing on a low-traffic website. The math does not work. The experiment runs for six months. By the time you have a result, the wedding is over, the couple has moved cities, and nobody remembers what the original research question was.

Run multivariate tests only when you have enough traffic that dividing it sixteen ways still leaves each slice statistically meaningful. Otherwise, stick to A/B.

Heatmaps Are Research, Not Experiments

This was one of my early mental shifts, and it is worth being explicit about it.

A heatmap does not tell you what will win. It tells you what might be wrong.

Click heatmaps show where users are clicking (and where they are not). Scroll heatmaps show how far down the page users travel before leaving. Session recordings are anonymized playbacks of real user journeys, where you can watch someone hesitate over your pricing table, try to click something that is not a link, or abandon a form halfway through.

None of these tools give you a winner. They give you a hypothesis.

Here is how the workflow actually connects: you open a heatmap for your pricing page and notice that most users never scroll past the first fold. Your main CTA button is near the bottom. Now you have a possible problem and a possible hypothesis: “If I move the CTA above the fold, more users will see it and click it.” You run an A/B test to confirm or deny it.

The heatmap gave you the question. The A/B test gives you the answer.

That distinction matters because it stops you from treating research tools as decision-making tools. Heatmaps, recordings, analytics, and surveys help you decide what to test. Controlled experiments help you decide what to change.

The Formula That Changed How I Think About All of This

Success = Chance × Frequency

Success is the number of winning experiments you produce in a year. Chance is your win rate, the percentage of experiments that produce a meaningful positive result. Frequency is how many experiments you run per month.

Now run the numbers with me.

Team A wants every test to be excellent. They spend weeks on research, debate every design decision, and get thorough sign-off before anything ships. Their win rate is high, say 100%. But all that effort means they can only run one test per month. Twelve winners a year.

Team B is less precious about each individual test. Their research is decent but not exhaustive. Their win rate is only 25%. But they run eight tests a month. Twenty-four winners a year.

Worse hit rate. Double the output.

This is not an argument for carelessness. Bad experiments waste time, and random button-color testing is still random button-color testing. But it is a clear argument against perfectionism as a strategy. The team that ships eight decent experiments a month will outlearn and outperform the team that ships one perfect experiment a month, almost every time.

If this sounds familiar, it should. DevOps made the same argument against waterfall delivery years ago. Smaller releases, shipped more often, create faster feedback loops. You learn sooner, recover faster, and stop treating every release like a quarterly ceremony. CRO has the same rhythm: define the hypothesis, design a clean test, ship it, measure it, document the learning, and move on. The goal is not to win every test. The goal is to learn fast enough that your wins compound.

A Note on Statistical Significance

Every testing tool gives you some version of a confidence percentage as your test runs. This tells you how likely it is that the result you are seeing is real and not just random variation. Most practitioners aim for 95% confidence before calling a winner.

But one of the most common beginner mistakes is stopping a test early because the numbers look exciting on Day 2. Early data is noisy. Weekday and weekend behavior can differ. Email campaigns, seasonality, and traffic source changes can all skew early results. A variant that looks like it is winning after 200 visitors may look very different after 4,000.

Let tests run for at least one full week, preferably two, before drawing conclusions.

And here is the question I now ask before trusting any result: not just “Did the tool say 95%?” but “Do I trust how we got to 95%?” The number matters. So does sample size, test duration, traffic quality, and whether the tracking was set up cleanly to begin with.

What Should You Run First?

If this feels like a lot, start small. Here is the beginner loop I would follow:

  1. Run an A/A test to confirm your measurement setup is reliable.
  2. Pick one high-traffic page with a clear conversion goal.
  3. Use heatmaps, session recordings, or analytics to find one point of friction.
  4. Write a hypothesis using the “If I change X, then Y, because Z” format.
  5. Run one clean A/B test.
  6. Let it run long enough to collect meaningful data (at least one week).
  7. Document what happened, win or lose.
  8. Repeat.

That loop is what CRO actually is, more than any specific tool or technique. Observe, hypothesize, test, learn, repeat. The experiment types are just the mechanics. The real skill is building a habit of running the loop consistently, and gradually running it faster.

In the next post, I will get into prioritization: how to decide which test to run when you have more ideas than time, and how to build a backlog driven by data rather than whoever spoke loudest in the last meeting.

I work with Convert.com and write about the intersection of DevOps and experimentation. If you are a technical person finding your footing in CRO, drop a comment or reach out. Happy to compare notes.