Model governance · Monte Carlo research · 8 min read

A model can run and still be wrong.

One of the most useful habits I developed in my independent study was learning not to treat a successful run as evidence that the model was correct. The harder question was: what kind of wrong am I looking at?

Complex models fail in several different ways. The code can be wrong. The data can be wrong. The economic assumption can be wrong. The documentation can be wrong even when the implementation is right. And sometimes nothing is technically broken—the output is simply telling you that one of your assumptions is much stronger than you realized.

Suspicious output is a diagnostic signal

During the Monte Carlo project, I repeatedly encountered outputs that looked implausible before I knew exactly why. Inflation could compound into a country state that felt crisis-like. A warehouse-age ramp could appear inconsistent with the written report. A rare event could show up often enough across 10,000 paths to make the whole distribution look distorted.

My first instinct early in the project was to fix the thing that looked wrong. Over time, that became a rule I actively resisted. Before changing a formula, distribution, or parameter, I tried to trace the result back through its lineage.

I started separating failure types

Failure typeQuestion I learned to ask
DataDid the input come from the intended source, period, unit, and reference class?
CalibrationDoes the fitted distribution reproduce the important shape and range of the underlying evidence?
Economic logicEven if the math is valid, does the mechanism make business sense?
ImplementationDoes the executable model actually do what the intended rule says?
DocumentationIs the written explanation describing the current executable model—or an older version?

The warehouse-ramp issue changed my debugging behavior

One review made the warehouse-age ramp look inconsistent. It would have been easy to “correct” the model immediately. Instead, I traced the implementation, the executive outputs, and the documentation separately. The executable model and results were internally consistent; the written description was stale.

That mattered because the correct fix was documentation, not economics. Changing the model would have introduced a new error while making the report look cleaner.

Other problems really were economic

That same discipline also made it easier to recognize when the model itself needed revision. Inflation/FX behavior and network-support logic were not just wording problems. Their interaction could produce economically questionable dynamics, so the mechanics were revisited and the validation suite expanded.

The distinction: a suspicious result is not evidence for a particular fix. It is evidence that the chain from source → assumption → transformation → implementation → output needs to be inspected.

Tests became model documentation

By the end of the project, validation was not an appendix I ran before submission. It was part of the architecture: 75 core checks, 22 inflation/FX checks, 60 financing-continuation invariants, and 12 scenario-reproduction checks.

Those tests did more than catch bugs. They forced the model to state what it believed. If a financing policy was supposed to defer a warehouse under a particular cash/debt constraint, an invariant could make that rule executable. If all twelve headline scenarios were supposed to reproduce the final paper, the reproduction suite made divergence visible.

What I carry into other work

I use the same mental model outside quantitative finance. In software, an analytics chart can be mathematically correct while the underlying product semantics are wrong. In operations, a dashboard can be accurate while the data-collection workflow makes the metric unreliable. In consulting, a slide can be well written while the recommendation depends on an unsupported premise.

The most useful debugging question is often not “where is the bug?” but “which layer is making the claim I no longer trust?”