Quantitative modeling · Monte Carlo research · 9 min read

Why dependence modeling became the hardest part of my Monte Carlo project.

Sampling one variable is easy. Sampling 23 country variables so their individual distributions and relationships remain economically coherent is a different problem.

Monte Carlo simulation sounds like a problem about random numbers. In my international-expansion research, the random-number generation became one of the easier parts.

Independent draws would create impossible countries

Country variables such as macroeconomic, institutional, cost, and market conditions are not independent. Sampling each factor separately would create combinations that look statistically convenient but economically incoherent.

The goal was not to claim a perfect structural model of every country. It was to preserve enough observed co-movement that simulated states did not routinely combine mutually inconsistent conditions.

First: let each variable keep the shape its evidence supported

I calibrated income-group-specific empirical marginals using adaptive/asymmetric kernel estimation where appropriate, partial pooling, numerical grids, and large diagnostic draw sets. That let bounded, skewed, or irregular variables keep their own shape instead of forcing every factor through a normal distribution.

Where subgroup samples were thin, partial pooling reduced overfitting. Where evidence was genuinely insufficient, I kept the structural assumption visible instead of converting it into fake empirical precision.

Then: preserve rank relationships across variables

I estimated pairwise Spearman dependence and converted it through the Gaussian-copula relationship. Rank dependence made sense because the variables had very different marginal shapes; I wanted the co-movement structure without requiring all of the raw variables to be jointly Gaussian.

But an assembled pairwise dependence matrix is not guaranteed to be positive semidefinite—the numerical property required to use it as a valid correlation structure.

Mathematical validity required a repair step

I used Higham’s nearest-correlation procedure to repair the matrix, then validated the joint system with 50,000 draws per income group after inverse-CDF transformation.

The repair itself was not the end. I compared simulated rank relationships back to the target structure because a matrix that is mathematically legal can still be a poor representation of the intended dependence after transformation.

The lesson: “correlated inputs” is not one checkbox. It is a chain of assumptions that has to remain mathematically valid after calibration, assembly, transformation, and simulation.

Dependence was only one layer of correlation

The country state also evolved over time, which introduced a separate question: how persistent should a condition be from one year to the next? Cross-sectional dependence answers what moves with what; temporal persistence answers how much of this year survives into next year. Those mechanisms had to stay conceptually separate.

Why I did not choose the most sophisticated model available

Costco’s company-level history was much shorter than the country panel. Ten annual observations were not enough evidence to justify fitting an elaborate multivariate company model, so I used a simpler central treatment plus robustness comparison instead.

The lesson: model complexity should be earned by evidence. A more sophisticated statistical object is not automatically a more defensible one.

Validation changed how I thought about the model

The final archived suite included core checks, inflation/FX checks, financing-continuation invariants, and scenario reproduction. The model became more trustworthy not because it was more complicated, but because it could explain and reproduce its own behavior.