The Fiddle Factor Nobody's Applying
In 2009, my co-founders and I created Winkle to bring to market some of the first online research tools for innovators. The reason we built it was simple: the standard tools for forecasting whether a new product would sell were slow, insanely expensive, and very poor at predicting the actual market performance of launches. Nielsen BASES, the market leader, is also a black box, where proprietary fiddle factors try to make measures of purchase intent correspond at least vaguely with real life. The fiddle factors are needed because everyone in the industry already knew that a five out of five on a purchase intent scale did not mean a five out of five chance of a sale. The whole model was built around a number everyone had agreed, quietly, not to trust on its own.
We didn't solve that problem, but we made the calibration a lot more transparent: benchmarked against live concepts already in the market rather than a fixed normative database, and fast and cheap enough to test far earlier and far more often. Alongside it we built tools to pull real consumer communities into shaping ideas from the start, rather than just rating them at the end. That's a genuinely different problem from the one this article is about. But it left me with a lasting suspicion of any purchase intent number that arrives without someone able to tell you exactly what it was corrected for.
I've been thinking about that fiddle factor a lot this week, because it's just disappeared from a conversation that badly needs it.
A post went round recently claiming that a toothpaste company had "quietly killed the entire market research industry." The claim rests on a real paper, from PyMC Labs and Colgate-Palmolive, and it's a genuinely clever piece of work. The problem researchers had run into with AI-simulated consumers was that if you ask a model directly for a rating out of five, it gives you safe, middle-of-the-road answers that don't look anything like a real distribution of human opinion. Their fix was to ask the model to write a free-text reaction first, in character, and then map that text back onto a rating scale using a similarity model. Tested against 57 real personal care product surveys and over nine thousand human responses, the synthetic panel matched what a second human panel would have said to within 90 percent of normal test-retest reliability. That's a real result. It solves a real technical problem, and the qualitative detail the models produce is genuinely richer than what most human panels bother to write.
But read that sentence again. Ninety percent reliable at reproducing what a *human panel* would say on a purchase intent survey. Not ninety percent accurate at predicting what people actually buy. The paper is careful about that distinction. The headline it's travelling under is not.
The number everyone already knew was soft
This isn't a new problem that AI has stumbled into. It's the oldest problem in market research, and the industry has known about it for sixty years.
The original study on this, from 1966, found that intentions surveys are weak predictors of purchase rates specifically because they can't pick up movement among the people who say they have no intention to buy, who go on to account for most actual purchases. Decades of follow-up work has confirmed it from different angles: intentions drift between the survey and the purchase; other people in the household weigh in on the decision the survey never asked about; the simple act of being surveyed changes what people go on to do, inflating the very correlation researchers are trying to measure. None of this is controversial inside the industry. It's why BASES runs its fiddle factors at all, and why every serious forecasting model built on stated intent has some version of the same correction sitting underneath it, usually invisible to the client who just wants a number for the board deck.
So the real story of what's happened is not "AI can now predict what people will buy." It's this: a very well-built piece of engineering has learned to reproduce, with impressive fidelity, a number that the market research industry has spent sixty years explicitly not trusting on its own.
The story isn't "AI can't be trusted." It's much subtler than that, and more sinister. The model isn't failing. It's succeeding, faithfully, at a task that was never a good proxy for the thing everyone actually wants to know. The failure, if there is one, belongs to us, for building a calibration step against the survey answer instead of against the sale.
Where the risk actually sits
The old system's weakness was visible, even if the fiddle factor itself was a black box. Everyone knew stated intent needed correcting, because the number came with an obvious asterisk attached: a respondent ticking the top box on a questionnaire because they want to be helpful, or because saying yes costs them nothing in the moment, in a way that buying the product later will not. You could feel the softness in it, even when you couldn't see exactly how Nielsen adjusted for it.
The new system's weakness is the same weakness, wearing a 90 percent reliability score. And the black box hasn't gone anywhere. Ask the people currently building these synthetic panels and, credit to them, plenty are candid about it: the methods are proprietary, the outputs are hard to independently verify, and clients are largely trusting a vendor's word for what's happening between the prompt and the score. That number doesn't feel soft. It feels like ground truth, and it will get treated like ground truth two or three steps downstream, by people who never read the paper and have no reason to know what it was actually measured against. A synthetic panel's output feeds a pricing model. The pricing model feeds a launch recommendation. The launch recommendation gets three lines in a board deck next quarter. By the time it lands in front of a decision-maker, the original asterisk, calibrated against a survey answer rather than a sale, has quietly disappeared.
That's the part that reminds me of 2008, and I want to be precise about why, because the comparison is easy to overreach. The 2008 crisis wasn't caused by anyone being surprised that mortgages carried risk. Everyone knew that. What made it catastrophic was that the risk got repackaged and re-rated enough times that the people holding it several steps downstream had stopped checking the thing underneath the rating, and trusted the rating instead. It was also catastrophic because the risk inside all those packages turned out to be far more correlated than the models assumed. Nobody had modelled the case where everything failed at once.
The synthetic consumer version of that second point is the one worth watching closely. Most of the vendors building these tools are working from a small number of shared foundation models. If there's a blind spot in how those models represent a particular demographic, a category, or a cultural moment, it doesn't show up as one company making one bad call. It shows up as a lot of companies making the same miscalibrated call at the same time, each one holding a report that says 90 percent reliable, none of them positioned to notice the pattern because from inside any single company it just looks like a normal, well-supported decision.
What this doesn't mean
It would be easy, and lazy, to turn this into "don't trust the AI." That's not the argument, and it undersells what the PyMC Labs and Colgate team actually built. A tool that can generate rich, demographically varied, less positivity-biased qualitative reaction at a fraction of the time and cost of a human panel is genuinely useful, especially early, when you're trying to rule out the worst ideas quickly rather than make a final call. The researchers themselves frame it as a complement to human research, not a replacement, and that framing is correct.
The argument is narrower and, I think, more useful: whatever correction factor your organisation used to apply to purchase intent data, when it came from a room full of humans, needs to survive the move to a synthetic panel. If anything it needs to get louder, not quieter, because the new number arrives with more apparent authority attached to it than the old one ever did.
And underneath all of this sits the question that no amount of calibration solves, because it was never a calibration problem. A synthetic panel, however faithfully it reproduces stated intent, is still answering a question you gave it. It cannot hand you the consumer insight that should have shaped the product in the first place, the actual, specific human need the thing is meeting. It can only tell you, more cheaply and at greater volume, how a simulated audience reacts to whatever you decided to test. That's a real capability. It is not a substitute for having found the right thing to test.
Where this leaves us
The most useful habit I took out of those years wasn't a formula, ours or Nielsen's. It was a reflex: whenever a number arrived looking clean, ask what its fiddle factor was, and who had set it. That reflex doesn't disappear because the number now comes with a methodology paper and a 90 percent reliability score attached. If anything, that's exactly the moment to ask it again.
So here's the question I'd put to anyone bringing synthetic consumer data into a real decision. Somewhere in your process, there used to be a person whose job was to distrust the purchase intent number before it reached the board. Does that person, or that step, still exist, now that the number arrives looking like science?
Sources: Maier, Aslak, Fiaschi, Rismal, Fletcher, Luhmann, Dow, Pappas and Wiecki, "LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings," arXiv:2510.08338 (October 2025); F. Thomas Juster, "Consumer Buying Intentions and Purchase Probability: An Experiment in Survey Design," Journal of the American Statistical Association (1966); Chandon, Morwitz and Reinartz, "Do Intentions Really Predict Behavior? Self-Generated Validity Effects in Survey Research," Journal of Marketing (2005); MarTech, "Synthetic research is a promise with a catch" (April 2026).
