Modeling and Statistics
We’ve got limited time, limited data, and infinite questions. How should we choose which questions to ask?
The first rule of a good model is that it’s explicit. It goes step-by-step.
Isn’t a model. It doesn’t say anything. It doesn’t tell me how the world works. Good models that change people’s minds explain things. They tell us how the whole world works, step by step, tracing the path from point A to point B. The players have incentives, goals, and structures they have to operate within.
Simple equations, although usually more robust (e.g., regressions, treatment effects) in the sense that they are true in more or more plausible states of the world, lose to models because they purchase their credibility at the cost of failing to provide meaning—“it made revenue go up” not “customers who value feature X have weaker price elasticities and…”—and we want to understand.
For example, we might compare two price vectors p and p’ in an ideal experiment. We can pick which one is better by comparing, say, average revenue, but you’ll be surprised how unconvincing that is unless there is some model for why p’ ought to do better than p. If you can say p’ takes into account the substitution patterns between the products you’re selling, while p prices each product as if it were independent, you’ll get immediate buy-in and even a thought that we should apply this everywhere, to other product areas, etc. If it’s just… empirically better, we’ll ship the treatment, but we’ll be wondering why constantly. If things go south, doubt will creep in, and we’ll start to suspect it was a fluke.
This is why religion crushes simple non-belief in the marketplace. It explains. It tells us how the world works, a real model of things. Non-belief doesn’t do that. You have to add to it to get a model of the world. Credibility doesn’t stand a chance against a model. No one’s overthrowing the bourgeoisie because they don’t believe in God. They have to believe in something else first. The dialectic. It’s important to consider these things when designing models to pick the optimal promo depth for your e-mail campaign.
The point is that building a model of what you’re doing matters. Customers choose among these alternatives. They value such-and-such. And, look, if that’s true, then the right price is $53.92 because… etc, etc.
So many good insights fall out of models. You can figure some things out deductively. The data is limited, and if you make it solve everything, it loses the power to tell you the answers to the truly unknown questions. For an extreme example that’s hopefully illustrative:
Suppose you have two products:
You can read an unlimited number of books from a certain library, and the reading app’s interface is Gold.
You can read two books per month from the same library, and the reading app’s interface is Blue.
Look, the product’s got two features, and some people might like Blue and other people might like Gold. So we’ve got both horizontal and vertical differentiation. Who’s to say which product’s price should be higher?
But we know, right? We have a model of the world that says readers care more about the number of books they can read than the color of the UI. We have another model of the world that says, in fact, we should just let people choose the color. No one’s willing to pay anything for that.
We could test this, and statistical noise would give us various customer valuations for Blue and Gold. It really would. We’d start talking about how people value Gold at around $2.31/month. There’d be confidence intervals. It’d look like analysis. We could make plans to drive Gold upsells.
Or: we can admit we know things, and if we use statistics to try to learn what we already know, we’ll just be drawing random shocks and tap-dancing around the truth.
The thing we don’t know here is how much people value reading two books per month versus the unlimited option. The signal we get from the data here is worth the noise. We can’t just know this from having a basic model of the world.
If instead of two books, it was a million, then we’d be back in the world of modeling again. No one could physically read that many books in a month. We don’t need to test. We have a model of the world—linear time—that collapses these two products into one.
When you get more explicit with even very simple modeling, surprising things fall out. Like, we have a suspicion that it takes us longer to use the Robots than it used to take us to do some tasks, but we use them anyway. Why? Well, it turns out a very simple production theory makes exactly that prediction…
We’ve got limited time, limited data, and infinite questions. We have to choose what to ask. Models help us decide what’s worth asking and what’s not.
Thank you for reading!
Zach
Connect at: https://linkedin.com/in/zlflynn

