Experimentation is not a scoreboard

How many experiments did your team run last quarter?
Collin Crowell recently argued that velocity reveals more about your experimentation program than rigor. He’s got a point: most teams run one or two tests per month against traffic that could support three or four times as many.
No amount of statistical discipline can fix a program that won’t try hard enough to find its real obstacles.
But, he notes, if you can hit your maximum, “rigor and per-capita experimentation is exactly where your attention should go.”
Few people understand what that looks like better than Ilya Izrailevsky, who has built experimentation platforms at Intuit and Robinhood and now leads the experimentation team at DoorDash. Each of these programs operate at a volume most teams will spend years reaching for.
Recently, he chatted with Katie Green and Makram Mansour (Head of Marketplace, ID.me) on Unite Voices to describe what that experience has taught him.
{{quote}}
Transitioning from experiment quantity to decision quality
What Izrailevsky found at scale is that the metric that built the program eventually stopped measuring anything useful:
{{quote1}}
Decision quality here means asking a simple question for every experiment: did a real decision depend on this result? Was the feature shipping regardless of the outcome? If so, the test was simply reporting and not useful, because you didn’t really learn from it.
For this reason, Izrailevsky’s team at DoorDash is called “decision systems,” rather than “experimentation” or “CRO.” It’s a helpful reminder that the output of a team that runs tests is decisions. At this scale, they’re past the point where raising velocity is useful.
Reaching the limit of test velocity
For many companies, reaching the point where increased velocity is a hindrance, rather than a benefit, is a distant dream. As techniques like prompt-based experimentation appear and improve, however, the cost of building tests is collapsing.
Today, any team can turn an idea into a live experiment at a pace once unthinkable. Now, backlogs are clearning and programs can reach the limits of their traffic years ahead of when they used to.
So the constraint moves. “Can I build this?” Becomes “should I test this?” Izrailevsky and Mansour both run programs where testing is not just the job of the specialist, and highlight the importance of approval workflows and guardrails when testing is so open; Mansour noted that as experimentation “democratizes,” “everybody can write content, everybody is now able to launch an experiment … with the right checks and balances.”
The human at the helm
The best guardrail to implement is one Izrailevsky and Green both agreed on: humans make the decisions.
{{quote3}}
{{quote4}}
Remember: if your program optimizes for launches, AI will produce more launches. If it optimizes for decisions, AI will help make decisions that actually grow businesses.
Whether you’ve reached your maximum velocity or not, this is the kind of rigor that allows experimentation programs to thrive.
Start counting decisions
Fortunately, none of this requires a reorg and none of it means slowing down. If your program can support higher velocity, look at how you can intelligently expand your testing. But, as Izrailevsky warns, don’t lose sight of what you’re testing for. Record the call that informs every test and what changed as a result of it, even if the answer is “nothing.”
This will lead to better conversations and honest reflections around your testing culture, the kind that turn a fast program into a decision engine.
“Earlier on, I focused on scalability. How can we run more and more experiments? The key metric was volume, because the more experiments you run, the better things will be. But what I learned is it’s not about the quantity, but more about quality.”

“Instead of pursuing this magic number of 10,000 experiments, 20,000 experiments, 100,000 experiments, focus more on decision quality.”

“Humans are always at the helm. Humans should be making the final calls.”

“AI is an amplifier, not a strategy, and you are the strategy that employs AI.”

Want to hear more? Ilya Izrailevsky and Makram Mansour discuss experimentation as a scoreboard, consolidating experimentation stacks, and testing with lower traffic on Unite Voices, Kameleoon’s podcast featuring real stories from the people behind today’s most innovative experimentation programs.
Want to hear more? Ilya Izrailevsky and Makram Mansour discuss experimentation as a scoreboard, consolidating experimentation stacks, and testing with lower traffic on Unite Voices, Kameleoon’s podcast featuring real stories from the people behind today’s most innovative experimentation programs.



