Login
English

Select your language

English
Français
Deutsch
Platform

PROMPT-BASED EXPERIMENTATION

Optimize any website by chatting with PBX, Kameleoon’s AI. Learn more

SOLUTIONS
Experimentation
Feature Management
KEY Features & add-ons
Mobile App Testing
Recommendations & Search
Personalization
specialities
AI in Experimentation
Single Page Application
Data Security & Privacy
Data Accuracy
Integrations & APIPartners ProgramSupportProduct Roadmap
Solutions
for all teams
Marketing
Product
Engineering
Data Scientists
For INDUSTRIES
Healthcare
Travel & Tourism
Financial Services
Media & Entertainment
E-commerce
B2B
Automotive

Samsung gains autonomy and speed in its most advanced tests with PBX

read the success story
PlansResourcesCustomers
Book a demo
Book a demo
Pick metrics that surprise you: the economist's rule for experimentation

Katie Green x Zach Flynn

0:00
--:--
Also available on:
home
EPISODE
9

Pick metrics that surprise you: the economist's rule for experimentation

Pick metrics that surprise you: the economist's rule for experimentation

Katie Green x Zach Flynn

0:00
--:--
Also play on:
Published on
August 31, 2026

About the episode

What if the metric you're tracking is guaranteed to go up, no matter what you build? That's not a win. It's a warning sign that you picked the wrong thing to measure.

This episode breaks down why the metric you choose matters more than the test itself, and how an economist's instinct for trade-offs changes what "success" looks like in an experiment. Zach shares the three pillars he used to scale experimentation at Udemy, and how to keep marketing, engineering, and leadership aligned around the same evidence.

About our guest

Zach Flynn is an economist who has spent his career finding the numbers hiding inside a business, from pricing work at Amazon and Udemy to his current role as a pricing and monetization scientist. He walked into experimentation with zero prior experience and ended up rebuilding an entire program from the ground up.

‍

‍

Zach Flynn
Pricing and Monetization Scientist
Fivetran
Katie Green
Principal Advocate & Host of Unite Voices
Kameleoon

Key takeaways

  1. Choose metrics you genuinely don't know the outcome of in advance, since a metric that always moves in the expected direction teaches you nothing.
  2. Decide your significance level and your decision criteria before the experiment starts, so no one reads the tea leaves after the fact.
  3. Remove process and guardrails first when a program has stalled, since teaching best practices only works once people are actually running experiments.

Transcript

Introduction: meet Zach Flynn

Katie Green: All right, and we’re live. Okay, Zach, thank you so much for joining us on Unite Voices, hosted by me, Katie Green. I’m principal advocate at Kameleoon. My favorite way to explain what I do is that I’m a community connector, and I sell experimentation: the practice of experimentation, the culture of experimentation. I love connecting people who are solving really complex problems, and one of those people is Zach, our guest today.

Zach, you were nominated for an experimentation thought leadership award in the data and engineering industry, analytics innovator category. That’s a lot of words to say all at once, but you have incredible experience with numbers, including a background in economics. Can you introduce yourself and tell people who you are?

Zach Flynn: Yeah, absolutely. I’m Zach. I have a background in economics. I’ve worked at places like Amazon and Udemy, and currently at Fivetran. I’ve done a lot of experimentation work, as well as pricing and other related work. My background is really anything that involves data, numbers, and usually some connection to economics.

Katie Green: I think that’s why people are going to tune into this episode in particular. I’ve run tests myself. I’m an experimentation practitioner, and I’ve run pricing tests, but I don’t have the background to measure that kind of experimentation rigor when it comes to testing pricing, pricing models, and subscription models. I relied heavily on our analytics team to help me understand the impact of those tests.

So I know a lot of people are going to be tuning in to hear your take, because they’re probably struggling with the same thing. They don’t have access to someone with your background in economics. Can you tell us a little bit about how that intersects with testing, since that’s the nature of the show? With your background in economics, how does your definition of a metric compare to that of a traditional product manager?

Choosing metrics like an economist

Zach Flynn: I think that’s a good example of how thinking economically helps you choose metrics. We want to pick metrics that will surprise us. We don’t want to pick metrics where we already know what’s going to happen after we run the experiment.

A classic example: say you replace an element on a web page and make it more prominent. If your metric is how often people click on that element, of course they’re going to click on it more often. That doesn’t surprise you, and it’s not really what you’re trying to learn. What you want to understand is the trade-off: what could go wrong, and what could go right? That’s what I want to choose as my metric for the experiment, so I can test which of those two things will happen.

In a pricing or monetization context, think about promotions. If people run a promotion with a discount, it’s not surprising that people buy more. That’s not what we’re ultimately trying to learn. We want to learn whether it drives the outcome we’re actually trying to reach, and we do that by picking metrics where we genuinely don’t know the direction of the answer before the experiment runs.

It also helps to have some statistical understanding that certain metrics move faster than others, usually the ones closer to the actual intervention. If you have a funnel and you’re intervening at step three, you want to look at metrics that happen after that, but closer to step four rather than step five. You’ll get a lot more power and a lot more impact by looking at metrics closer to the intervention.

Katie Green: I love that. I think power and impact are what people are looking for, and it’s definitely why they’re tuning into this episode. I want to get into the fact that you completely reworked an experimentation program in your past. I think that’s going to be the big chunk that people are looking for, because maybe they’re in a similar position.

But first, in the context of your current role as a pricing and monetization scientist, can you tell us a little bit about how you’re working with metrics and data today, and what that means for your understanding of experimentation at large? Then we can get into the nitty gritty.

Pricing, monetization, and the case for long experiments

Zach Flynn: I think the way to connect the pricing and monetization work to experimentation is: as much as possible, you want to be able to test things. One thing I’ve learned from my current role is to not be scared of running long experiments. Sometimes experiments take a long time, but if the impact is large enough and it’s a significant business direction, you can justify it.

This is an interesting context because it’s the B2B enterprise world, so there are fewer data points than you might have in an e-commerce setting. You might need to run the experiment longer, but one of the nice things is that you can ultimately get insights that are grounded in reality, rather than the kind of vibes-based decision-making you’d otherwise have to rely on.

People usually aren’t running experiments directly on price itself. They can, in some contexts, but a lot of the time you’re running experiments on the periphery, like how often people get discounts for various things.

Katie Green: Yeah, I like to call it the presentation of price. It’s the emotional presentation of how much something costs.

Zach Flynn: Right, exactly.

Rebuilding Udemy’s experimentation program from scratch

Katie Green: That’s great. And when it comes to experimentation, I know you have such deep, robust experience here, and you were working on a program where you completely restructured it. That’s the project I’d love for you to share with people. Can you tell us a little bit about the infrastructure you were working within, and what changes you made to scale up experimentation? We have a ton of practitioners tuning in, and their biggest question is: how do I take it to the next step?

Zach Flynn: Yeah, so this is back when I worked at Udemy. What I did there was, first off, I knew nothing about experimentation before starting that job. I had never run an experiment or done any kind of experiment analysis, so I was in that exact boat, but I got a job to work in experimentation there.

The goal was to rewrite the entire experiment analysis process: how we do experiment analysis there. One of the nice things about that was I got to think about a lot of problems from first principles. I didn’t have any preconceived ideas about how to do it correctly. There are three key pillars, I’d say, to getting that right.

Pillar one: doing experiments beats doing them perfectly

The first pillar is that always doing experiments is better than not doing experiments. If you can get people to run experiments, that’s a positive thing. Getting caught up on whether they’re running it correctly, whether every i is dotted and every t is crossed, is not really a good way to go. We want to make sure people look at some data about what happened when customers saw A and what happened when customers saw B. The details are good and important to get right in a mature situation, but the first part is just doing that: getting people to look at both those things.

Pillar two: choose a significance level that fits your business

The other piece is making sure we’re okay with making decisions without necessarily using the statistical significance levels that are often used in academia, like the 5% level. In a business context, you’re generally willing to take more risk than you would with major government policy, drug approvals, or airplane design. Here, we’re talking about e-commerce or something similar. No one crashes if you get this wrong. So the key thing is to choose a significance level that makes sense for your industry and the projects you’re working on.

The reason for that is twofold. First, if people always get an insignificant result, they’ll stop looking at the data and stop thinking about the size of the effect. You want to make sure people actually see meaningful results often enough, in a reasonable timeframe, before they have to make decisions.

Pillar three: metric selection

The other key piece, which I touched on earlier, is metric selection. A standard reflex is that everyone is trying to drive revenue, so their default is to pick revenue as the metric. The problem is that everything in the world affects revenue: macroeconomics, demand for that product class, competitors entering the market, a social media post going wrong, whatever. Given enough time, your experiment will be able to distinguish whether your change increased revenue or not, but it will take a really long time to get there.

A better approach is to pick metrics that are closer to your actual intervention. If you improve some part of the funnel, look closely at that part and find metrics with real statistical power. That’s where running a power analysis comes in: looking at realistic sample sizes for your experiments.

I also like to encourage people to think carefully about how often they check, or “peek at,” their experiment results. The correct answer isn’t zero peeking, and it isn’t peeking every day either. It’s something in the middle. The way to structure that is to pre-specify how many times you’re going to peek during the experiment, and then calculate the correct statistical adjustments for that when you’re making decisions.

You want to get to a world, though you won’t get there right away, where people are mostly deciding to ship things that are statistically significant, and not shipping things that aren’t. That takes a while to get to, and there are obviously practical things that come up, but that’s the end state goal you want to work toward.

Katie Green: I’m guilty of peeking. My bad. I would do it daily, mostly because what I struggle with when I’m running tests is figuring out what percentage decrease in a KPI I’m comfortable with before deciding something really isn’t working. I know you have to wait until you hit statistical significance, but when you have a hundred thousand people in a test and the conversion rate is down sixty percent, that’s obviously dramatic. Having some level of statistical rigor to understand when it’s premature to call it is something I’ve really struggled with in my career.

I hope somebody listening is going to want to DM you on LinkedIn about this. I think experimentation resolves questions. That’s what we’re here to do: we ask questions, and we provide answers.

Navigating organizational conflict with evidence

Katie Green: But what that creates in a corporate environment, particularly in B2B, is friction, or we’ll call it conflict. Philosophy clashes when you bring data evidence into the room. I know a lot of people in experimentation struggle with that. I’ve had calls with CMOs and CEOs where I’ve had to say, “This idea you were certain was going to work, we tested it every which way, and it doesn’t.” You have a leadership position and have done so well creating rigor and a culture of learning and failing forward. I’d love to hear how you handle those deep organizational conflicts when evidence counteracts a philosophy that’s already in place.

Zach Flynn: Yeah, absolutely. I’d say the first thing is that one of the powerful things about experimentation is that it gives you the most ammunition to say something isn’t working. If you think an executive’s idea isn’t going to work because of some descriptive statistics or observational analysis, there’s going to be a lot of back and forth about whether you checked the right thing or whether an assumption really holds. One of the good things about experimentation is that it’s much more direct: we showed half the people this, half the people that, and the people who saw the new thing did not do so well. I think that’s very helpful in shaping those conversations.

Sometimes, when people aren’t convinced by that evidence, the truth is there’s something else behind their thinking that the data hasn’t addressed. It’s often good to try to extract exactly why the evidence doesn’t convince them. Sometimes it’s a local optimum versus global optimum situation: they want to move the business in a certain direction and understand they’ll take a hit initially until the new thing gets optimized.

Think about an e-commerce site that’s highly optimized under one pricing model, and then leadership decides to make a dramatic pivot, say from transactional to subscription. That’s going to initially hurt revenue, because you haven’t optimized for it yet and you haven’t built anything for it. In that case, the leader is essentially saying, “We’re making this pivot because we think it’s the right structural move.” The question then becomes how to make that pivot good, not whether to make it, so it’s important to understand the context behind the concern.

In general, though, I think experimentation is one of the best ways to convince people. Most people I’ve found are convinced by experiments. If they thought something would work beforehand, and then you ran an experiment and it didn’t, in my experience they’ll generally agree that particular version of the idea isn’t convincing enough to move forward with. A lot of times they’ll come back with an iteration: they still like the general framework and want to try a different variation on the theme. But experimentation is a very convincing way to talk with people.

Katie Green: I talk a lot about the HiPPO, the highest-paid person’s opinion, with our executives. I feel like I have a similar story to a lot of people coming into experimentation: we fell into it because we had a question and wanted to know if we were right. I remember my first test was actually an email test, not website experimentation. I was working on lifecycle and multi-channel testing through paid channels, and I remember thinking, “How do we know this is going to work?” And then, “Oh yeah, let’s A/B test it.” That was the first time I had that thought, so being able to have that conversation backed by evidence is really important.

There’s one other thing you mentioned earlier that came up at our Unite Summit London recently, where I had the pleasure of emceeing our sold-out event. One of the things that came up on a panel was exactly what you’re describing: changing how you think about your metrics, from test-specific metrics to that global change. How are these metrics affecting your business metrics, your business KPIs? So many people build decision trees and understand that metrics have to ladder up, but it’s really hard to have that conversation about decision-making if the evidence isn’t speaking the language of the person making the decision.

Balancing marketing, engineering, and product needs

Katie Green: Before we get to our last question, marketing versus engineering, versus data, versus design. There are all these different teams that touch experimentation, and given your experience rebuilding a program from scratch, how do you create evaluation criteria that work for marketers, developers, and analysts alike? How do you manage the metrics while balancing each team’s approach to experimentation?

Zach Flynn: There are definitely different needs across those teams. Marketing usually needs something that’s fairly easy to deploy, since they run a much larger number of experiments than teams touching the back end of the website. They might test different variations of an email, a text message, or an ad. So they want a platform where they can run a lot of things quickly, almost templated: here’s the metric, here’s the variant, and there’s one clear metric to make a decision on.

A lot of marketing experiments also aren’t necessarily decision drivers in the traditional sense. You run one, learn a general principle, and the next piece of marketing copy is probably going to be different from the last one anyway. So it’s a lot of quicker, information-gathering tests that usually don’t last very long.

On the product side, you need the ability to change things on the back end and be deeply integrated into the site, which matters a lot for engineering too. You want an easy way to do that, and you want to make sure everything is instrumented. That’s one of the more technical realities of experimentation: you need to make sure the metrics actually exist and can be measured.

That usually takes more time than people think, because when a product is just starting, people are focused on building the business and generating revenue. The details of making sure a button click triggers something in a database aren’t usually top of mind early on. So one of the first steps toward becoming more experimentation-forward is simply instrumenting those metrics so they exist. That’s a fixed cost: you pay it once, and it makes engineering’s life much easier going forward, even as you add more metrics over time.

Products also tend to run longer experiments with a wider variety of metrics, often specific to whatever new feature or product is being tested, which requires flexibility on the back end. So the way to satisfy all these groups is to build a setup that can run fast experiments and long experiments, and can change the back end, the front end, and the mailers, while still having one unifying structure underneath: this person is mapped to variant one, this person is mapped to variant two. That way the analysis becomes automatable, straightforward, and easy to make decisions from.

Decide your decision criteria before the experiment starts

One last thought on that front: a key thing is to decide how you’re going to decide before the experiment starts. Otherwise, you end up in an ex post situation where everyone looks at the metrics after the fact, squints, cuts them every which way, reads the tea leaves, and lands on a story that usually just confirms their prior beliefs.

The powerful move is to say, before the experiment, “If this metric goes up and this other metric doesn’t go down, we’re going to launch it,” and decide that in advance. That’s a good way to get everyone on the same page, because you can refer back to it afterward: here was our decision criteria, and here’s what we learned that would make us change it.

Katie Green: I think that’s the most important piece I want to underline: what is our decision-making criteria, and if we’re divorcing from it, what’s the reason why? Having that in advance, along with the flexibility point you made earlier, is really critical. No experimentation program is one size fits all, and you have a really good breadth of experience speaking to that. Thank you for sharing your knowledge on this.

Monday morning advice: remove the guardrails first

Katie Green: With our final few minutes, I ask our guests all the same thing before I let them go, which I call the Monday morning advice. Say someone is building a brand-new experimentation program, or their program has stalled and they’re looking to scale up and make a dramatic change. What is the one thing you think they should start with?

Zach Flynn: The one thing you should definitely start with, for getting a program off the ground or expanded, is to go out and preach the gospel of experimentation: get people to just run experiments. Don’t put guardrails up. A lot of times, the reason people don’t run experiments is that there’s too much process. With good intentions, teams create requirements like documents everyone has to write and forms everyone has to fill out. That’s a good practice to eventually get to, but if the current problem is that no one’s running experiments, those documents don’t matter yet. That’s a secondary concern.

The first thing is getting people into the practice of running experiments, even if they do it by the seat of their pants at first. That’s still going to be more helpful for the business in figuring out which direction to go. So my first recommendation is to look at what guardrails or processes are currently in place that create friction, and take them away, even if they have good reasoning behind them. If the problem is that no one’s running experiments, remove those barriers, get people running experiments, and then teach best practices while they have actual experiments in front of them to learn from. That becomes a much easier conversation than trying to teach best practices for something people don’t currently do at all.

Katie Green: I love that. We’ll clip that every which way for sure. It’s practice what you preach: if you’re preaching that you want a culture of experimentation and learning, you have to take a hard look in the mirror and ask whether your own processes are actually pushing that culture forward. That’s a wonderful piece of advice to end on.

Zach, thank you so much for being a part of Unite Voices.

Zach Flynn: Absolutely. Thank you so much for having me.

‍

Read THE FULL TRANSCript
hide transcript

Build experiments in minutes by chatting with AI

Describe what you want. Kameleoon's Prompt-based Experimentation (PBX) will generate and launch tests instantly.

Try it for free
Try it for free
Experiment your way

Get the key to staying ahead in the world of experimentation.

[Placeholder text - Hubspot will create the error message]
Thanks for submitting the form.

Newsletter

Platform
ExperimentationFeature ManagementPBX Free-TrialMobile App TestingProduct Reco & MerchData AccuracyData Privacy & SecuritySingle Page ApplicationAI PersonalizationIntegrations
guides
A/B testingPrompt-Based ExperimentationFeature FlaggingPersonalizationFeature ExperimentationAI for A/B testingClient-Side vs Server-Side
plans
PricingMTU vs MAU
Industries
HealthcareFinancial ServicesE-commerceAutomotiveTravel & TourismMedia & EntertainmentB2B & SaaS
TEAMS
MarketingProductDevelopers
Resources
Customers StoriesAcademyDev DocsProduct RoadmapCalculatorWho’s Who
compare us
OptimizelyVWOAB Tasty
partners
Our Partner EcosystemBecome a PartnerIntegrations DirectoryPartners Directory
company
About UsCareersContact UsSupport
legal
Terms of use and ServicePrivacy PolicyLegal Notice & CSUPCI DSS
© Kameleoon — 2026 All rights Reserved
Legal Notice & CSUPrivacy policyPCI DSSPlatform Status