A/B Testing, True Knowledge, and Power Analysis
Over the years, I've run into a few scenarios that piqued my interest in experiment design.
To be honest, when some of my stats friends were taking classes on experimental design in college, I thought those classes sounded boring as hell. Something about about how to frame survey questions, or maybe calculating margin-of-error for public opinion polls.
In practice, understanding the limitations of experiments is extremely useful for both planning and interpreting A/B tests!
Stat-Sig A/B Test != True Knowledge
A/B testing is widely used at tech companies and beyond. Meta is no exception. Even though it's so widely used, many, many people don't always fully understand it's nuances; and occasionally that causes issues when people see confusing results and want an explanation. In the sciences, similar issues have surfaced with regard to the "reproducibility crisis."
To start off, the hardest lesson I've learned about statistics is that we don't really have "true knowledge" of our A/B test results until we can get (1) very tight confidence intervals or (2) run multiple replications. By "true knowledge" I mean that I have a "good idea" of the effect size with "high confidence."
For example, if you see a *barely* *stat-sig* results with 95% confidence (p=0.05), you might think that you've met the golden standard of science. In fact, exact replication would fail almost 50% of the time, even if your experiment measured the *true effect size* perfectly.
To get some intuition for this, we can remember that A/B test results are drawn from a distribution. Let's say that our experiment result was +10k page views +-10k. We will also follow the approach of the paper above and assume we know the true effect size (in practice, it could be larger or smaller), and we'll assume that experiment results can be drawn from a normal distribution using the Central Limit Theorem.
As we can see in the diagram above, based on the assumptions we've made about the experiment, we will only replicate successfully half of the time. We would need to look at a large body of experiments to see if this assumption is too generous or too harsh in practice, but using the Maximum Likelihood Estimate as our hypothesized distribution is certainly a reasonable place to start.
To increase our chances of replication above 50%, we would need to increase the sample size significantly (informed by power analysis).
Tight Confidence Intervals as an Alternative to Replication
If you have the luxury of running larger experiments, one alternative to re-running experiments is conducting a large enough experiment to get very tight confidence intervals in the first place. For example, if we ran a large enough experiment to get a result like "+10k page views +- 5 page views" with 95% confidence, we would know that it's extremely unlikely for the true effect size to be far away from +10k page views.
The problem with the +10k +- 10k result in the previous section was that the confidence interval contains a lot of small results. For example, we might only consider it worth rolling out our experiment if it increases page views by at least 1k. If it is only neutral, then we don't think it's worth shipping.
If we get a tighter confidence interval that excludes low values, we can have more confidence in the true effect of our experiment. Of course, we still need to deal with the fact that our confidence intervals are noisy (e.g. 5% will be wrong with 95% confidence intervals). This essentially the route that equivalence testing takes.
For decisions where we need to make few mistakes, we can increase the significance level of our confidence intervals (e.g. to 99%). When we do this, the confidence intervals will widen. In order to tighten the less-noisy confidence intervals again and increase the "power" of our experiments we will then need to collect much more data.
Takeaways
1. Power analysis is your Friend
The process of analyzing how much data we need for an experiment (e.g. given the minimum effect size we want to detect, and how noisy our data is) is called power analysis. If we don't do this before running an experiment, we just have to hope the samples we collected are enough to give us the information we need.
I won't go into the details of power analysis right now, but hopefully this is enough motivation for you to think about how noisy experiments can impact your work -- at least for cases when you want to make important decisions.
If you are looking for resources on power analysis, this article might be an okay start. IMO, the wikipedia article isn't very accessible.
2. Sometimes, it's good enough to make noisy decisions
For most of us, we probably think of "true knowledge" as result we can trust -- probably because it's able to be replicated reliably.
For ourselves or for an organization, we might choose a lower standard for shipping experiments. Why? 1) Because true knowledge is hard to obtain (it can require a lot of samples or multiple replications) and 2) Because sometimes it's enough to take noisy steps in mostly the right direction, and, on the whole, this is a better use of our time.
In the past, I've seen similar arguments made about promotions as a type of "gradient descent" at companies. Not every promotion decision is perfect, but on the whole, they usually move the company in the right direction over time. A similar logic might apply to organizations and experimentation.
In a future blog post, I'll briefly review how we can investigate organization-wide experimentation more rigorously. Here's a teaser.

Comments
Post a Comment