“What giants?” asked Sancho Panza.
“Those you see over there,” replied his master, “with the long arms; sometimes they are almost two leagues long.”
“Look, your grace,” Sancho responded, “those things that appear over there aren’t giants but windmills, and what looks like their arms are the sails that are turned by the wind and make the grindstone move.”
“It seems clear to me,” replied Don Quixote, “that thou art not well-versed in the matter of adventures: these are giants; and if thou art afraid, move aside and start to pray whilst I enter with them in fierce and unequal combat.”
— Miguel de Cervantes, Don Quixote, Part I, Chapter VIII [1]
Introduction
The scientific method is built around a simple idea: formulate a hypothesis, test it against reality, and decide whether the evidence supports it.
In data science, we do this constantly. We run experiments, build proofs of concept, compare models, validate performance, and ask whether an idea works before investing further in it. We have become remarkably good at testing hypotheses before trusting them.
But what happens after they work?
A hypothesis that survives an experiment can become a model. A model that performs well can become a system. A successful system can become a process, and a process repeated for long enough can eventually become the way things are done. Somewhere along that path, an interesting reversal can occur: what once had to prove itself against reality can eventually become part of the lens through which we interpret it.
And reality does not stand still.
Populations change, and as a consequence, data-generating processes change with them. Technologies evolve, organizations adapt, and yet past success provides a powerful reason to keep trusting the assumptions that produced it.
This raises a question that goes beyond model monitoring or technical performance:
How often do we retest the assumptions behind something that still seems to work?
In this article, I want to explore that question at different levels, d to data acience methods, and consider a deceptively simple possibility:
What if we keep adapting our models and methods without adapting the way we understand the problem?
Kodak and Moody’s illustrate two different forms of the same underlying problem. Kodak represents the more familiar story: an organization struggles to move beyond the logic that made it successful. Moody’s presents a more subtle (and perhaps more dangerous) version. The organization did adapt. Models were revised, methodologies evolved, and new information was incorporated as markets changed. Yet those changes could remain bounded by the assumptions already embedded in the system [2]. Kodak asks what happens when we fail to change. Moody’s asks a harder question: what if we are changing all the time, but not changing what really matters?
When Success Becomes a Constraint
Everybody knows the cautionary tale of Kodak, the classic case study of a company that failed to adapt to change. But I want to use Kodak to make a point that goes beyond the usual story of technological disruption.
For decades, Kodak built an extraordinarily successful business around a particular understanding of photography: cameras generated demand for film, film generated recurring revenue, and processing completed the ecosystem. Digital photography did not simply introduce a new technology. It challenged the assumptions that made that system work [5].
Experience is valuable precisely because it allows us to recognize patterns and make decisions without rediscovering everything from scratch. But there is a paradox: the more successful a particular interpretation of the world becomes, the easier it is to stop seeing it as an interpretation at all. This can be reinforced by what behavioral theory calls status quo bias: once a particular way of working is established, we become disproportionately inclined to preserve it rather than reconsider the alternatives. Past success makes that tendency even easier to justify.
The same thing can happen in data science.
Our models do not begin with algorithms. They begin with decisions about how a problem should be represented. We decide what data matters, how it should be transformed, which relationships are worth modelling, what success looks like, and which assumptions are reasonable enough to proceed.
At first, these are choices. But when they work repeatedly, they become practices. Practices become processes, and processes eventually become methodology. What began as a hypothesis about how to solve a problem can quietly become the accepted way of solving it.
And this creates a more subtle kind of risk. The problem is no longer simply whether a model becomes outdated. It is whether we can keep updating the model while leaving the way we frame the problem largely untouched.
Models Capture Reality Through Assumptions
A model learns from the world it has seen. The difficult question is whether it remains useful when the world changes.
The question is not only how to predict Drug B, but why relationships learned from Drug A should still hold for it.
A similar image as above was shown to me at the end of an interview. With about five minutes left and almost no context, the interviewer asked me to explain what I was seeing. The actual slide was slightly different, but the problem was essentially the same.
Several questions immediately came to mind, but one stood out: if we learn from Drug A, what makes us confident that the relationships we learned will still hold for Drug B?
You have probably heard this many times: a predictive model is, by definition, a simplification. It learns relationships from observations generated under particular conditions. We choose variables, define outcomes, make assumptions, and reduce a complex reality into something we can model.
There is nothing inherently wrong with using historical drugs to predict the uptake of a new one [3, 4]. In fact, learning from related populations, domains, or tasks is a well-established idea in statistics and machine learning. But doing so requires assumptions about what is transferable. Are the relevant populations comparable? Are the mechanisms driving uptake sufficiently stable? Do the predictors have the same meaning? Has the market or data-generating process changed in a way that breaks the relationships learned historically?
If those assumptions are explicit and we have evidence that they reasonably hold, using Drug A to learn about Drug B may be entirely justified. The problem is assuming transportability rather than establishing it [4].
And that was precisely what made the interview question difficult. With the information available on the slide, I could understand the proposed modelling strategy, but I could not conclude that applying relationships learned from historical drugs to Drug B was necessarily the right approach. The most important information was not only the model or its historical performance. It was whether the assumptions that allowed us to move from A to B were defensible.
Once Drug B launches and its actual uptake becomes observable, those assumptions can be confronted with genuinely external evidence. Until then, historical performance tells us what worked in the world we observed, not automatically what will work in the world we are trying to predict.
A model learns from the world it has seen. The difficult question is what must remain true for it to work in the world it has not.
And this leads to a deeper problem. We know that moving a model into a new environment should force us to question what makes that transfer valid. But what if, instead, we keep modifying a model while the assumptions defining the problem become increasingly difficult to see, let alone challenge?
That is where Moody’s becomes particularly interesting.
Moody’s: The Illusion of Adaptation
Moody’s provides a much more concrete example of what happens when the world changes faster than the assumptions used to model it.
Moody’s is one of the major credit rating agencies. Part of its job is to assess how risky financial products are and translate that risk into ratings such as AAA. In the years leading up to the 2008 financial crisis, Moody’s rated securities backed by thousands of residential mortgages, meaning that part of its job was precisely to assess the credit risk embedded in those products.
The problem is quite familiar: given the characteristics of a pool of mortgages and the borrowers behind them, estimate how that pool will perform under different economic conditions.
To estimate their risk, Moody’s developed models that used information about individual borrowers and mortgages, together with historical data and simulations of economic conditions, to estimate how much money a pool of mortgages could lose.
So, the process was highly model-driven. Moody’s received a loan-level dataset containing information about the individual mortgages in a pool. Analysts fed this information into a proprietary model that estimated expected losses and how much protection would be required for the security to achieve a particular rating. Those outputs then informed analysts and rating committees before a final rating was published.
Then the mortgage market changed dramatically…
Subprime lending expanded, lending standards became much looser, increasingly complex securities appeared, and mortgages requiring little or no documentation became far more common. The population Moody’s was trying to model was no longer the same.
And Moody’s noticed.
It developed a more sophisticated model (called M3) that simulated mortgage performance under different economic conditions. As subprime lending exploded, it went further and developed M3 Subprime, specifically calibrated for this new type of mortgage [2].
But the assumptions underlying the model changed far less.
Moody’s continued to rely heavily on historical relationships to simulate future economic conditions. The macroeconomic simulation engine used for M3 was carried into M3 Subprime. The company continued to trust information supplied about the underlying loans even as lending standards and data quality deteriorated. And because there was limited historical evidence for this rapidly changing subprime market, some parameters had to rely heavily on expert judgement [2].
In summary, Moody’s improved the model, added complexity, recalibrated parameters, and created a version specifically for subprime mortgages. What it did not fully do was rebuild its representation of the problem around the possibility that the market itself had fundamentally changed.
When the housing market collapsed, the assumptions embedded in that representation no longer described the world the model was being asked to predict.
Methods Can Become Models Too
The same reasoning can be applied beyond individual models.
Data science itself has accumulated methods in response to real problems. The scientific method gives us a framework for formulating hypotheses, testing them against evidence, and revising them when they fail. Software engineering adds reproducibility, testing, and maintainability. DevOps and MLOps address deployment, monitoring, versioning, and continuous operation. Governance adds controls for increasingly consequential systems.
These practices did not appear arbitrarily. They are accumulated solutions to problems we encountered along the way. And that is precisely why they deserve the same scrutiny as models.
A decision that works becomes a practice. A practice becomes a process. The process acquires tools, roles, KPIs, documentation, and governance. Eventually, we may become very good at executing it without remembering which assumptions made it appropriate in the first place.
This does not mean constantly reinventing how we work. That would defeat the purpose of accumulated experience. It means recognizing that methods, like models, have conditions under which they make sense.
We routinely ask whether data has drifted, whether a model needs retraining, or whether its performance has deteriorated. Perhaps we should occasionally ask the equivalent question about our methods:
What would need to change in the world for this way of working to stop making sense?
Conclusion
We spend a great deal of effort validating ideas before we trust them. Perhaps we should spend more time revalidating the assumptions that made them successful in the first place.
Kodak reminds us that success can make the assumptions behind a way of working increasingly difficult to see. Analogue-based pre-launch forecasting makes the role of those assumptions explicit: transferring patterns from historical products is valid only insofar as the conditions that make them comparable still hold. Moody’s then reveals the more subtle problem: a system can continue to adapt while leaving precisely those underlying assumptions largely untouched.
A model can be monitored, retrained, and even improved while the way the problem itself is framed remains largely unchanged. Monitoring can tell us when performance deteriorates. Explanation and uncertainty can help us understand when a prediction deserves less confidence. But none of these, by themselves, tells us whether the assumptions defining the problem still make sense.
Perhaps one of the most important properties of a good model is not only knowing what it predicts, but knowing where it stops and where questioning should begin.
The lesson is not that experience, models, or established methods should be distrusted. It is almost the opposite. They are valuable because they encode what we have learned. But accumulated knowledge should remain open to evidence, perspectives, and questions that come from outside the frame in which it was built.
What we have learned should never become indistinguishable from what must be true.
Evolution depends on keeping our assumptions open to challenge, especially when success has made them difficult to see. That is part of good scientific practice, but it is rarely an easy one.
References
[1] Cervantes Saavedra, M. de. (1605). Don Quixote of La Mancha, Volume I, Part One, Chapter VIII. English translation in the Publiconsulting Media online edition. Publiconsulting Media. Chapter VIII — Don Quixote of La Mancha
[2] Omidvar, O., Safavi, M., & Glaser, V. L. (2023). Algorithmic routines and dynamic inertia: How organizations avoid adapting to changes in the environment. Journal of Management Studies, 60(2), 313-345.
[3] Robey, S. H., & David, F. S. (2017). Drug launch curves in the modern era. Nature Reviews Drug Discovery, 16, 13–14. DOI: 10.1038/nrd.2016.236.
[4] Guseo, R., Dalla Valle, A., Furlan, C., Guidolin, M., & Mortarino, C. (2017). Pre-launch forecasting of a pharmaceutical drug. International Journal of Pharmaceutical and Healthcare Marketing, 11(4), 412–438. DOI: 10.1108/IJPHM-07-2016-0036.
[5] Lucas Jr, H. C., & Goh, J. M. (2009). Disruptive technology: How Kodak missed the digital photography revolution. The Journal of Strategic Information Systems, 18(1), 46-55.
[6] Morrison, E. W. (2023). Employee voice and silence: Taking stock a decade later. Annual review of organizational psychology and organizational behavior, 10(1), 79-107.

