A Guide To The Good Posts
A list of the posts that had a good thesis, effort, and, generally, didn't involve me being an old man, saying, "Things are different now, and I don't like it."
I have written this blog for three years this month. First, on Medium before it had a name; then, on Medium when it was called Good Enough Statistics and briefly flirted with the idea of accepting submissions; and now, a shiny Substack, where I plan to stick around. Because it has LaTeX embeds… [and I can write where I live vicariously through the blogs of the kinds of people that know the fashionable places to eat in New York City.]
The nature of the idea distribution is that the good ideas are much better than the so-so ideas, which are astonishingly better than the bad ideas, and the frequency of such ideas is inversely proportional to the quality. My rare—but extant!—good ideas since writing this blog are spread across the years, and so, they’re hard-to-find in the chronological Archive.
I decided to make a guide, so you don’t have to sift through the shit.
The criteria were simple. If I thought the post was pretty good when I scrolled by it in mid-August 2026, then I added it. I call this criterion “Taste (August 2026)” because I’m hearing from all the well-connected folks that Taste is the future of humanity.
So, here are the good posts. They are in a rough, chronological order because I couldn’t think of a better way to organize them. Enjoy!
I look at how we can use experiments to optimize a product over time when the product design decision can be modeled as a continuous variable. For example, prices, discount percentages, the length of a free trial, or the period of time a one-time purchase remains in effect. The key idea is to treat experimentation as evaluating a function and its derivative at a given parameter value, and updating as if you were solving an optimization problem via gradient descent.
I’ve made this point in a few places, but I think this is the most succinct version: we should choose metrics to maximize surprise. We don’t want to goal our experiment on metrics where we know what will happen. We want to maximize what we learn. We want to think, prior to the experiment, that it could go either way. For example, conversion rate should never be the primary metric for an experiment that gives some users a discount. Of course, larger discounts will increase the conversion rate. We don’t need to run the experiment to know that. We need a metric that measures both the potential costs and benefits of the change we’ve introduced. This shows up more subtly in many experiment designs discussed in the post!
A simple production model with uncertainty about how best to do a thing and a technology to automate doing it makes surprisingly strong predictions that match at least my experience with LLMs. The idea of the model is to treat LLMs as an automation technology (of course, in reality they have other uses) and study the implications. A neat result here is that workers will use LLMs for tasks they can do faster/better without the tech for moderate-uncertainty tasks (perhaps, LLM-generated docs…).
There’s this common objection to running an experiment: we need to move fast here. I argue that, properly understood, experiments do not slow us down. They increase product velocity. Velocity, of course, is the product of speed and direction. The steering wheel affects velocity at least as much as the gas pedal. It will take us longer to change direction if the change doesn’t work out if we don’t run the experiment because we won’t have evidence there’s a problem for a while.
I reference this post all the time. It’s one of the things it’d be cool to have more general buy-in for. I use insights from a paper that derives frequentist hypothesis testing as a statistical decision rule to form the basis for a simple, practical way to choose significance levels, instead of defaulting to 0.05.
The context for this post was the prevalence of blogs around this time saying statistical significance was bad, and I just had the thought: is it really? I didn’t find most of the alternatives in these posts obviously better, and I started thinking, “There are all these posts about how bad statistical significance is, and yet everyone uses it. Curious.” So, I started thinking about the revealed-preference gap, and I concluded that statistical significance is a pretty good way to make decisions. It’s better than the alternatives usually proposed anyway. It has the two main elements: effect size and how uncertain we are about it and a single-parameter to tune how much weight to place on both of them. Clearly, any decision rule will be some version of those two features, and statistical significance has an intuitive weighting scheme (the significance level). So, I came away with the sense that the criticism of statistical significance is about something else, not the statistical significance framework, and discussed creating a Cult/Church, depending on your perspective.
We don’t test ideas we think will do poorly. We spent engineering and design time building out the feature because we think it's a great idea. So, without experiments, we would just launch our great, new idea. If our new idea is actually great, then all the experiment did was waste time. If we hadn’t run an experiment, we would have launched earlier. But if, to our surprise, our new idea is actually bad (sadly, this is usually the case—at least in its initial form), then the experiment is highly valuable. It prevented us from making a mistake and blindly pursuing the idea we were once so gung-ho about. I include a calculation of the value of an experimentation program based on the losses it prevents.
This is a really common situation: so many experiments run on a site simultaneously in a mature experimentation program. So, the question is: do interactions between the experiments matter? I do a little math and show that—quantitatively—the interaction effects need to be very large before the effects could actually lead to you making the wrong decision. And that’s the main point of experimentation: decision-making. Interaction effects do impact estimated effects, though, if you’re relying on the exact effect value in some way. There’s no avoiding interaction effects—or environmental effects present during your experiment but not afterward—from biasing the impact estimates. Online experimentation is not a clean-room laboratory, but it is an amazing way to make the correct decisions about which direction to go.
Sometimes, you’ll get the advice not to “peek” at your experiment results early. But peeking is good! It can reduce the expected time of experimentation! Sometimes good things happen, and we should roll with it. You just have to make inferences in a way that accounts for peeking, and it’s easy.
This is one of those problems that rarely shows up in academia because the dataset is more or less fixed, so the issue is under-discussed in standard statistical training. But it’s very common when running online experiments. We have a metric for a unit we’ve randomized, and the question is: “Over what time period should we compute the metric after the unit enters the experiment?” The default in many experimentation platforms is from the time the unit entered the experiment to today. This causes a lot of problems for inference. It’ll make your experiments take much longer to run. It is even possible that you would never find a significant result—even if one exists—if you continued to run the experiment forever!
Thank you for reading, and for a delightful diversion from work for the past three years,
Zach
Connect at: https://linkedin.com/in/zlflynn






