Three A/B tests on the wrong problem

One step in your checkout flow converts eleven points below the one before it.
You know the numbers: segment splits, device breakdowns, trend lines, and the last six months of history.
And you still have no idea why.
Most programs respond to this by running an A/B test on something plausible. You’re guessing. It’s an educated guess and you’re trying to validate it with data, but it’s still a guess. The real cause could be completely different from where your intuition pointed you first.
Five different causes for the same drop
Akash Doshi led part of the digital strategy for Delta Vacations for over three years. He joined us on Unite Voices to talk with Katie Green about what can sit behind a single drop-off.
{{quote}}
Doshi diagnoses five different causes for the same drop, all of which would look exactly the same in an analytics dashboard.
Worse, each would have a different solution: API failure means engineering, while add-to-cart calls for a form fix. If it’s behavioral, there’s no fix needed, but if the experience is confusing or the UI isn’t clear, you may need copy changes or UX restructures.
But the data you have records what happened, rather than why. So how do you close the gap?
Qualitative data alongside dashboards
{{quote1}}
Doshi’s insight is a good one, but “voice of the customer” can mean a lot of things. For example:
- Session replays show where hesitation, backtracking, and repeated clicks cluster.
- Support tickets and chat transcripts tell you what customers identify as obstacles.
- Surveys triggered at friction points catch intent while it still exists.
- Usability sessions show you what someone tries in the moment.
This is why it makes sense to use your funnel as a locator and let qualitative work pick up the natural shortcomings in that approach.
Dividing labor between qualitative and quantitative data
In Doshi’s own words:
{{quote3}}
Remember: qualitative insight is directional, but is rarely conclusive. Once you’ve reviewed your customer support tickets indicating that the label is confusing, you still don’t know that a rewrite is the right move or what specifically to say.
Customers are notoriously bad at predicting their own behavior. So you collect qualitative data to tell you which problem is real, and then experiment to find out whether your fix works. Rather than a better hunch, you’re left with real answers at the end of the process.
Why this matters as testing gets cheaper
For most of the last decade, this problem has been fairly self-limiting. When the average variation took two or three sprints to build, the queue itself was a filter, forcing teams to prioritize tests, leading to what Doshi describes as “a giant backlog of great hypotheses that never make it to market.”
He further notes that “AI’s ability to insert itself into experimentation lets you clear that backlog.” Today, with tools like Kameleoon’s Prompt-Based Experimentation (PBX), anyone can build a variation in minutes. Weak hypotheses can now reach production as easily as strong ones, which means a program can triple its test count.
{{pink-block-1}}
Faster testing is usually a good thing, and improved velocity, done right, means improved learning. As Doshi notes:
{{quote4}}
But this is also why qualitative data is crucial. Let’s say you are tripling your test count using PBX, up from one test per month to three. If you’re trying to solve the checkout flow from earlier, you could easily run all three tests on the wrong issue and solve nothing. Your learnings amount to what the problem wasn’t.
Of course, that’s still valuable. But had the tests been based on real customer feedback instead, you might have used the same time to test three different solutions to a real problem and found the one that best solves it.
When your A/B tests find the solution for a proven customer problem, you can scale your program significantly and greatly improve your customers’ experiences.
{{cta-block}}
“If it's the payments page, maybe an API is failing, or maybe customers are having trouble adding something to cart. Maybe the experience is confusing, or the UI doesn't make it clear how to proceed. It could also be a behavioral problem—the customer just isn't ready to convert, maybe they're checking prices. There are all these possibilities.”

"When we have voice of the customer, we’re able to really zoom in on the actual issue, rather than doing a lot of guesswork and extrapolating. We get a very firm indicator of what’s going wrong.”

“The first half is the voice of the customer identifying the correct problem, and the second half is A/B testing to make sure we have the right solution in place.”

“When you get to market quicker, feedback loops activate. You hear from the customer, which validates your guesswork instead of you just saying ‘we think this is the right path forward.’”

As A/B tests become cheaper and faster to build, these program expansions become increasingly realistic for all teams. Using Kameleoon’s Contentsquare integration transforms Sense Analyst insights into test-building prompts that bridge the gap between feedback and variation in a single step.
See how Contentsquare insights become Kameleoon experiments here.
Want to hear more? Akash Doshi discusses how to balance vibe coding with product management, the order of operations on data, and qualitative sources on Unite Voices, Kameleoon’s podcast featuring real stories from the people behind today’s most innovative experimentation programs.
Want to hear more? Akash Doshi discusses how to balance vibe coding with product management, the order of operations on data, and qualitative sources on Unite Voices, Kameleoon’s podcast featuring real stories from the people behind today’s most innovative experimentation programs.



