Login
English

Select your language

English
Français
Deutsch
Platform

PROMPT-BASED EXPERIMENTATION

Optimize any website by chatting with PBX, Kameleoon’s AI. Learn more

SOLUTIONS
Experimentation
Feature Management
KEY Features & add-ons
Mobile App Testing
Recommendations & Search
Personalization
specialities
AI in Experimentation
Single Page Application
Data Security & Privacy
Data Accuracy
Integrations & APIPartners ProgramSupportProduct Roadmap
Solutions
for all teams
Marketing
Product
Engineering
Data Scientists
For INDUSTRIES
Healthcare
Travel & Tourism
Financial Services
Media & Entertainment
E-commerce
B2B
Automotive

Samsung gains autonomy and speed in its most advanced tests with PBX

read the success story
PlansResourcesCustomers
Book a demo
Book a demo
Why do people ignore good A/B test results? A statistician's perspective

Katie Green x Cristina McGuire

0:00
--:--
Also available on:
home
EPISODE
11

Why do people ignore good A/B test results? A statistician's perspective

Why do people ignore good A/B test results? A statistician's perspective

Katie Green x Cristina McGuire

0:00
--:--
Also play on:
Published on
September 25, 2026

About the episode

What do you do when the data says one thing and a stakeholder refuses to believe it? Most experimenters never get taught how to handle that moment.

This episode is for anyone who has watched a good hypothesis get lost in politics, pressure, or too many metrics. Cristina walks through how she raises the bar on rigor without becoming the office math teacher, and how she gets stakeholders to accept results they don't want to hear.

‍

About our guest

Cristina McGuire has spent more than five years building experimentation programs from the ground up, first at Expedia Group, then Chewy, and now on the Windows experimentation team at Microsoft. She came into the field as a data scientist with a statistics background and learned along the way that the math is only half the job.

Cristina McGuire
Experimentation Data Scientist
Microsoft
Katie Green
Principal Advocate & Host of Unite Voices
Kameleoon

Key takeaways

  1. Anchor every test to one or two primary metrics before it launches, so you are not tempted to swap in a flattering metric halfway through.
  2. When a stakeholder pushes back on a result, quantify the trade-off in dollar terms instead of arguing the statistics.
  3. Revisit your experimentation strategy roughly once a month rather than per test. A good plan is replicable across many experiments.

‍

Transcript

Meet Cristina McGuire: from data science to experimentation at scale

Katie Green: Welcome to another episode of Unite Voices, hosted by me, Katie Green. I am serving as the principal advocate at Kameleoon, and I am joined by Cristina, who has incredible experience in experimentation. She has built programs at companies like Chewy and Expedia, and is now most recently joining the team to do experimentation at Microsoft. So, Cristina, can you introduce yourself and tell everyone who you are?

Cristina McGuire: Yeah, I'm Cristina. I entered the world of experimentation more than five years ago as a data scientist. As Katie said, I'm currently working with the Windows experimentation team at Microsoft. Before that, I was at Chewy, and before that, at Expedia Group, where it all really started.

I came in as a data scientist in a role where the main goal was to drive experimentation, and I've never turned back since. I really love being in this space. A few years ago, it felt like a niche, but recently experimentation feels like a much bigger community, and I'm excited to be part of it.

Katie Green: I feel that too. We're a tight-knit group. We do podcasts like this, we go to events, and it just feels like it's getting bigger. Your name came to me through a connection of a connection, which shows in real time how people get connected in this community. That's one of the takeaways I want from this podcast: somebody listening and thinking, I want to follow this person, connect with them on LinkedIn. That's success in my eyes. So for anyone who wants to connect with Cristina, or with me, we're always here for you.

You just name-dropped some huge brands. A lot of people listening probably work at smaller or medium-sized businesses, or on smaller teams, even within big brands. Can you tell us what it's like to work at scale? Everywhere you've worked has had a huge experimentation program.

Building experimentation programs across different company cultures

Cristina McGuire: I have a background in statistics, and I started my career doing general data science. What really got me fascinated with experimentation is that I joined a team and realized this is the most practical application of statistics, scaled in the real world. You're technically running t-tests hundreds, thousands, or more times. Getting started with that, I thought, wow, this is really interesting. I became deeply involved in the experimentation community: attending conferences, networking events, and learning from other practitioners.

I was lucky when I started my experimentation journey. I joined Expedia Group, where they already had a mature experimentation program, but they were rebuilding it to make things even better. So it still felt like we were building the experimentation culture, but now powered by a platform. Having a really good platform is what makes experimentation scale even better, because as statisticians we know the right thing to do, how you should run experiments like a scientist, but you can't personally watch every single person and check that they're running good experiments. So working hand in hand with the platform is really what drives the scale that makes your experimentation culture strong.

When I moved from Expedia to Chewy, what got me interested is that Chewy had a smaller experimentation program. I was hoping to bring in the foundations I'd built at a more mature company like Expedia, and build that foundation from scratch, so people could run more experiments with statistical rigor, at a level where they could really use it for decision-making.

And Microsoft is another big company. Each organization has a different culture, and the way you scale is very different too.

Bringing statistical rigor to a team without becoming "the math teacher"

Katie Green: Your experience is so rich, and I'm curious about something I didn't prep you on. You have a master's in statistics, and a lot of experimenters don't have that kind of rigor on their teams, or in themselves. How do you teach this? When you're working with teams, like at Expedia specifically, how do you show people how to have that rigor without being a math teacher? Maybe you are a math teacher.

Cristina McGuire: I did teach math for undergrads as part of my teaching assistantship, so I do have some of that experience. But that's a good question, and I might take a different route to answer it.

When I started, my manager knew I had a statistics background, and within my team I'd worked with people who had PhDs in statistics, really scientifically driven people. When I started my career at Expedia, my role was to raise the bar of experimentation, and I took that seriously. I'd been a few years out of school, but I still carried the mindset of, this is what I learned, and this is how people should run experiments. So I went in thinking my goal was to identify how teams were experimenting, find what was wrong, and teach them the right way.

That didn't last long. I realized everybody wants to run good experiments. It's just that they're dealt different limitations and constraints: limited time, incomplete data, organizational pressure to prove conversion is up, or that the product isn't working when the business wants to push it anyway. There's a lot more to experimentation than statistics can solve. Statistics helps, but it's not the full story. That's one of the big learnings for me: coming in with a strong statistics background, but layering in more business understanding and empathy for how people are actually doing the work.

What I really like about having a statistics background is that it makes some conversations less emotional. Say a customer wants to launch an experiment. You can bring the conversation back to a neutral zone by asking, did you power your experiment? You go back to the structure, because if you have the structure, when they talk to their leaders, it's not about them, it's just the math and the data.

That's kind of how I layer in statistical rigor. But you can't just tell people, you have to do a power analysis, or your hypothesis is wrong. Having that conversation and understanding the business need is what really makes the bigger impact.

Measuring the immeasurable: turning brand impact into real metrics

Katie Green: So much of testing, in my experience, involves things that are almost immeasurable. How do we measure brand impact? You have to create the right hypotheses and track the right metrics. What experience does a statistician bring to that part of the workflow, taking something amorphous and turning it into black-and-white numbers, so you can help people make good, easy decisions? What's it like translating from emotion to the practical?

Cristina McGuire: When you have a statistical background, it's very structured. There's a path: let's design the experiment. I need to understand the goal, whether that's identifying or measuring something like brand, and what data sets are available that could act as a proxy for how well customers feel about your brand. Do we actually have that data?

It always starts with having a good hypothesis, even if that sounds like a cliche, and understanding what's measurable. What metrics can we use to help the customer answer, can we quantify brand? From the business's perspective, that might be revenue impact. From the customer's perspective, it might be retention rate, or engagement with the product. A lot of the time, it's about grounding people back to what's measurable and what we can actually answer, because if we had unlimited power, we'd want to know the long-term impact of a feature on the brand. But experimentation has real limitations. You might only have two to four weeks to measure something, so you may not get that long-term answer. There are more advanced statistical methods to work around that, but is there a simpler way to go back to the basics: designing a good hypothesis, identifying the right metrics, understanding what engineering can actually launch, and deciding what decisions we can make from the insight?

At the end of the day, you want to maximize learning with the resources you have. As a statistician or data scientist, you ground people in what's measurable, how it can be implemented, and what decisions can be made. If that's not good enough, you go talk about it more.

Why a good hypothesis is the most replicable part of experimentation

Katie Green: A lot of people are going to resonate with that. I think we say, oh, it's so obvious, start with a good hypothesis. But there are so many people running so fast, with such high expectations and stress, that they just want to get tests out the door, no matter what they are. I once saw a great conference presentation that said teams shouldn't measure success by how many tests they run. They measure by how many decisions they make.

Cristina McGuire: Right. And I think sometimes in the life cycle of an experiment, someone will ask a question in the middle, like, should I run this experiment longer? You wouldn't even need to ask that if you'd spent enough time developing your plan earlier.

Maybe I'm biased, but I think most of the time you only need to develop that plan or strategy once every month or so and reassess it, because it's replicable. That's the fun part about experimentation: you run one good experiment, get good at it, and then run it over and over again with different features. Maybe I'm simplifying it, but I think people get overwhelmed by the details. If you spend time strategizing your experimentation program up front, you can scale it.

How to prioritize primary, secondary, and guardrail metrics

Katie Green: I love that, and it is replicable, that's part of what makes it fun. Something I personally struggle with, and I'm kind of selfishly asking you this, is that it's easy to find an insight. It's like looking for a needle in a stack of needles when you have hundreds of metrics available. I want to get a little more tangible here. Is it the strategic piece up front where you say, this is the only decision-making metric for this test? Or do you flex into other metrics based on performance? Can you tell us how you prioritize the hierarchy of your metrics, and what that does to the impact of the test?

Cristina McGuire: This is where my statistician mindset kicks in. I think sticking to your hypothesis in terms of identifying your primary metric, secondary metric, and guardrails is key. You fill those in before you start the experiment, so you need to use them wisely. Having a good primary metric is really important to make sure you're sizing your test well. If, in the middle of the experiment, you decide a different metric looks good and want to make that your new primary metric, you might just be observing it because the test is short and you didn't have enough power to confirm the change isn't exaggerated.

Anchoring to one, or maybe two, primary metrics really helps remove a lot of decision paralysis once the data comes in, because there will be plenty of data outside the metrics you originally listed. If you're a statistical purist, you want as few metrics as possible, because that reduces your multiple-hypothesis errors and lowers your false positive rate if you're intentional. But we also live in a generation where data is abundant, so there's a pull to give users access to everything. There are pros and cons to that. When you have all the data, some of what you're seeing is likely just randomness. But if you become too restrictive and limit yourself to, say, ten metrics, experimenters feel like they don't have enough.

What I've seen work is allowing people to track a lot of metrics, which depends on your engineering capacity to calculate them, but making sure users clearly specify their real primary, secondary, and guardrail metrics, and educating people that those are the metrics you use to make the launch decision. All the other metrics are available to help generate new hypotheses or surface an insight you hadn't considered. So there's value in both: the set of metrics you use to make the launch decision, and the metrics that enrich your learning about the experiment or feature. They serve different purposes, and you can use both. Just always ask, will this make it easier to make a decision, or will I just confuse myself in the end?

What to do when stakeholders don't want to accept the data

Katie Green: Confusing myself in the end is probably what I've done a lot in my career, if I'm being honest. That's why I selfishly asked. As I continue my career, don't be surprised if you get a DM from me asking for help with a problem I'm having.

I think we've done a great job talking about how to be rigorous, and you've mentioned it takes the emotion out of it. But what do you do when you run into people who don't want to accept a reality, when the test outcome says one thing but somebody is trying to poke holes in it? I've seen this myself: someone really wants a feature to go live, and we're showing that conversions are down, but next-step clicks are up. That's not what the hypothesis was measuring. How do you navigate a situation where somebody doesn't really care about the math?

Cristina McGuire: That's a good question, and I've definitely seen that. That's where you need to hit the brakes and not freak out, put yourself in the other person's shoes, and understand what stakeholders and organizational pressure they're dealing with. Most of the time, I'd suggest agreeing with them first: okay, I see what you're saying, but let's also quantify the trade-off. For example, engagement is really high, but we're not seeing enough change in the primary metric. Make sure you're presenting the whole story, not filtering it.

My approach is yes, but, or yes, and: let's show the full story. If you have a good hypothesis and clear metrics, you present those no matter what other insights come up. How deep you go depends on the customer and the problem they're trying to solve. If a metric is improving, the question becomes, can we convert that into revenue? If someone says the click-through rate on a specific button is really high and asks if we should launch, I'll ask, how many dollars does that mean for more people to click that button? The rule of thumb for opportunity sizing is to put a dollar value against your metric and ask, is this really important? Is it practically significant? Will the business actually care? And how does that compare to the effort and cost of launching the feature? So I battle it with another data point.

Katie Green: That's perfect. At Unite the Summit in London, our conference there, I led a panel, and one of the main points was that every single metric on your experiment should ladder up to a business objective. That's another way of saying, what is the actual value of a click-through? Does it outweigh a decrease in conversion? That could be the case, who knows, I've seen crazier things when it comes to experimentation.

I really wanted to underline that: stick to your primary metric, because that's your hypothesis. And going back to your very first point, you've done the work and had the discussions to make sure your strategy and hypothesis are solid. When you carry that work into your metrics, it filters into the final decision, and you're able to stand firm in what you're presenting and say, I know this is the right read.

Monday morning advice: leveling up without a statistics background

Katie Green: So for everyone listening, there's your recap. We're running out of time, but there's one question I like to ask at the end, what I call the Monday morning advice. For someone listening who wants to make their experiments more reliable, but doesn't have a statistics background, what can they do to level up and start implementing some of the rigor and strategy you've maintained throughout your career?

Cristina McGuire: I'll answer it in two parts: from the perspective of a data scientist, and from the perspective of an everyday experimenter.

If you're working in an experimentation platform, or you're a data scientist embedded in a product team, my advice is to always understand the customer's actual problem, and that's usually not the literal question they ask you. Most of the time, if someone asks, is this the right minimum detectable effect, and you start calculating the MDE, the real question they're trying to ask is more like, how long do we need to run this test, or what's the cost. I think a lot of data scientists jump straight to, what's the metric, what method do we apply, what's the data design, when really you need to zoom out and understand where the question is coming from. It's also worth recognizing that the barrier is rarely the same across different teams. Sometimes people need to improve their experimentation knowledge, sometimes the real problem is engineering or data quality, and no amount of statistics will fix that. Because experimentation is such a cross-functional field, you need to figure out the actual problem and what the customer needs before jumping to a conclusion.

For the everyday experimenter, the good news is you don't need to know all the statistics. If you have a good experimentation platform that you trust, the rigor and quality of experimentation has gotten better and better over the years, especially with platforms driving that scale. I've done a lot of benchmarking of different experimentation platforms, and the basics are there. As a starter, don't get overwhelmed thinking you need to know the statistics, because most of the time you don't. Find the knowledge repository for the platform you're using, understand the steps, and there's usually a generic process for how to run an experiment. Do it once, then do it over and over until you become the expert.

Most experts in this field only had to run one or two experiments before they really understood the process, because a lot of people shy away, thinking they need more statistics or more engineering knowledge before they can run an experiment. But really, you just need to go through the process once or twice, and you'll gain enough knowledge to start seeing experimentation as a tool for making bigger, better decisions, rather than a process that delays your product development. So, you don't need to know the statistics. Just go for it, and lean on the data scientists on your team. They'll help you.

Katie Green: I love that. It's not the answer you typically get from someone in your role, telling you that you don't really need them. But I love it, and I think it's something everyone can take away. It was wonderful talking to you. I realize we're at time, so thank you so much for being on Unite Voices, Cristina.

Cristina McGuire: This was so fun. Thanks, Katie.

Read THE FULL TRANSCript
hide transcript

Build experiments in minutes by chatting with AI

Describe what you want. Kameleoon's Prompt-based Experimentation (PBX) will generate and launch tests instantly.

Try it for free
Try it for free
Experiment your way

Get the key to staying ahead in the world of experimentation.

[Placeholder text - Hubspot will create the error message]
Thanks for submitting the form.

Newsletter

Platform
ExperimentationFeature ManagementPBX Free-TrialMobile App TestingProduct Reco & MerchData AccuracyData Privacy & SecuritySingle Page ApplicationAI PersonalizationIntegrations
guides
A/B testingPrompt-Based ExperimentationFeature FlaggingPersonalizationFeature ExperimentationAI for A/B testingClient-Side vs Server-Side
plans
PricingMTU vs MAU
Industries
HealthcareFinancial ServicesE-commerceAutomotiveTravel & TourismMedia & EntertainmentB2B & SaaS
TEAMS
MarketingProductDevelopers
Resources
Customers StoriesAcademyDev DocsProduct RoadmapCalculatorWho’s Who
compare us
OptimizelyWingify
partners
Our Partner EcosystemBecome a PartnerIntegrations DirectoryPartners Directory
company
About UsCareersContact UsSupport
legal
Terms of use and ServicePrivacy PolicyLegal Notice & CSUPCI DSS
© Kameleoon — 2026 All rights Reserved
Legal Notice & CSUPrivacy policyPCI DSSPlatform Status