When Models Disagree, Causal Testing Helps Data Leaders See Which Results Will Hold At Scale
Emily Nightingale, Director of Marketing Data Science at The Knot Worldwide, on how disagreement between sources becomes its own decision-grade signal.

If each of your models is telling you a different thing, that means you have no idea what the truth is. If you just choose one of those and make a decision, you may be spending money very inefficiently.

A dashboard can hand a leader a clean, confident answer, and that's what makes it dangerous. A model built on observational data will report that a channel or initiative is performing and attach a high confidence score to the result. The number starts to feel like the truth. That confidence reflects the math. It doesn't test the causal assumption underneath it, and a tidy output usually means something was left out. Data teams that read their models well treat a suspiciously clean answer as a reason to look harder.
That discipline sits at the center of how Emily Nightingale approaches her work as Director of Marketing Data Science at The Knot Worldwide, the wedding-planning company whose demand is tied tightly to life events and a sharply seasonal engagement cycle. That structure makes her category an unusually strict test of any signal: the peaks are predictable, so the hard part is separating genuine marketing impact from demand that was always going to arrive. Her answer is to resist the single clean number wherever it appears, starting with the tool most teams lean on hardest.
"If each of your models is telling you a different thing, that means you have no idea what the truth is. If you just choose one of those and make a decision, you may be spending money very inefficiently," she says.
Why one attribution source can skew the picture
The first place Nightingale wants leaders to lose confidence is any setup running on a single attribution source, especially a click-based one. "Whatever your CDP is, it's useful for decision making, but they are heavily biased to the bottom of the funnel," she says. Lean on them alone and the credit flows to paid search and other last-click channels, subtly steering spend toward the end of the journey while the demand-creating work upstream goes unmeasured.
Her prescription is deliberately harder to live with than the alternative. She wants leaders using several sources at once and triaging between them, precisely because they will conflict. One will call a channel a winner; another will tell a different story. "I know that's so frustrating as a busy leader who just wants to know what to do next, but I actually think some of that noise is really important," she asserts. The friction is the point, and it produces the through-line of her whole approach: the disagreement is information, not noise to be resolved away.
Where causal methods earn their place
The reason Nightingale pushes past correlation is that only a causal answer tells her whether a result will hold up when she acts on it. A pattern pulled from observational data can be a real and interesting finding and still collapse at scale. "Any other non-causal model may not scale," she says. "If you roll it out to 100%, you may see no additional impact. You may see negative impact." A causal read, by contrast, tells her how a change will behave when it goes company-wide, which is the only version of the answer a budget decision can safely rest on.
For a business where seasonality swamps almost everything, she gets that read through match-market testing, which is a form of causal inference. She illustrates with an intentionally oversimplified example comparing Los Angeles and New York. "Let's say LA and New York act very similarly and you run a marketing campaign in LA only. The only difference between those two cities is the campaign. So, you know that any difference in performance between New York and LA is due to your marketing campaign." Run such a test during the January engagement-season surge, and the method still isolates the marketing lift from the seasonal wave that would have arrived regardless. It's how she answers the question her category makes hardest: how much of what happened was actually the marketing?
The blind spots you can name but can't measure
The harder problem is the signal she knows exists but can't see cleanly in the data. New competitors and the rise of LLMs are both reshaping demand, and both are nearly invisible in her inputs. "It's a huge thing that we know is disrupting, but we really can't quantify it well," she says of AI search, where the platforms surface almost nothing about who's looking for a brand and third-party data is spotty at best. A confounding variable can be real, in other words, without being something you have a usable measurement of.
Rather than paper over the gap with synthetic data or push ahead on a model she doesn't trust, Nightingale designs around it. She favors methods that don't require every confounding variable to be specified in the first place, like match-market testing, A/B testing, and propensity matching, because each one holds the unknowns steady instead of demanding she enumerate them. Propensity matching, for instance, controls for a user's underlying intent, so a campaign doesn't get credit for high-intent people who were coming back anyway. The choice of method is the safeguard against the blind spot.
No model is finished, and that's the honest starting point
Even a fully specified model offers no escape from uncertainty, because there's no version that accounts for everything. Nightingale points to the econometrics literature and the book Mastering Metrics, for the underlying truth: one economist builds a causal diagram capturing every input they believe matters, and another can always arrive to insist a different variable belongs. "Everyone is right and everyone is wrong at the same time. Schrödinger's model," she says. There are technical ways to test whether adding a variable improves the fit, but the search for a perfect, complete model is a category error.
This is why she keeps returning judgment to the process rather than trying to engineer it out. Experimentation covers the high-stakes calls, the ones that could tank revenue if they go wrong. Everything below that threshold, where testing isn't worth it or isn't possible, comes down to decision-making, and she's blunt that the modeling never removed the human from it. "All of these inputs, all these assumptions, are just decisions about what the data scientist believes to be correct."
Turning uncertainty into a signal instead of a distraction
The instinct under pressure is to treat unresolved uncertainty as something to omit so a decision can be made. Nightingale takes the opposite view. "That uncertainty is a huge data point for what you should be doing, or how you should be thinking about it, or what you should prioritize," she says. When her sources diverge, that divergence tells her how confidently she can move, whether a call needs a test before rollout, and where to aim scarce experimentation. The gain is not resolution, but definition. "It's a known unknown now instead of an unknown unknown."
Nightingale is clear-eyed about where the broader adoption of AI is heading, and she folds that into the same framework. Her expectation is a cycle: heavy, enthusiastic use, followed by visible failures born of over-reliance, followed by a swing back toward control, process, and oversight. The point isn't to sit out the efficiency AI offers. It's to know which signals to trust it with and which still need a person who can look at a clean, confident answer and ask what it left out.




