What Agent-Based Models See That Regressions Miss
The Trap of Spurious Correlation
In 1974, Granger and Newbold demonstrated that two independent random walks, when regressed against each other, produce high $R^2$ values and significant $t$-statistics with probability approaching one as the sample size grows. The regression looks convincing. The relationship is entirely fabricated.
This is the canonical spurious regression problem. But there is a subtler and more insidious version: two phenomena that are genuinely causally linked at the micro level, yet appear unrelated — or even negatively correlated — when measured at the aggregate level and subjected to standard regression. The primitives of statistical inference are not equipped to see the mechanism. They can only see the covariance structure, and covariance is a marginal statistic that integrates out everything that matters.
Agent-based models operate differently. They do not estimate relationships between aggregate variables. They generate those relationships as emergent properties of explicitly modelled micro-level rules. What looks like noise at the macro level resolves into structure at the micro level — structure that is invisible to regression but transparent to simulation.
Why Standard Regression Fails Here
The classical linear model assumes that the conditional expectation $\mathbb{E}[Y \mid X]$ is a useful summary of the relationship between $X$ and $Y$. For many physical and economic systems, it is not — because the relationship is non-linear, path-dependent, or mediated by a distribution of micro states that the aggregate variables $X$ and $Y$ do not capture.
Consider a system of $N$ interacting agents with individual states $\omega_i \in \Omega$. The aggregate observable is a function $X = \Phi(\boldsymbol{\omega})$ where $\boldsymbol{\omega} = (\omega_1, \ldots, \omega_N)$. The macro-level distribution satisfies:
$(1)$
Equation (1) is a marginalisation over an astronomically large micro state space. Two different micro configurations can produce identical macro observables; two similar macro observables can be generated by micro dynamics that are causally opposite. Any regression of $X$ on $Y$ conflates all of this. The ABM preserves it.
The Formal Advantage
The ABM’s formal advantage over regression is not computational — it is epistemological. A regression asks: given the data, what function best describes the relationship between $X$ and $Y$? An ABM asks: given a model of micro behaviour, what aggregate relationship does the system generate?
The first question is answered by minimising a loss function over observed data. The second is answered by forward simulation. When the true data-generating process involves interaction effects, non-linearities, path dependence, or multiple equilibria, the first question has a misleading answer. The second question does not.
The practical upshot: whenever you observe a time series or cross-section and a standard regression produces insignificant, unstable, or sign-switching coefficients — do not conclude that there is no relationship. Conclude that the relationship, if it exists, may only be visible at the micro level.
Takeaway
Spurious regression is usually discussed as a statistical nuisance to be corrected with better tests. The deeper point is different: even well-specified regressions between non-integrated series can fail to detect genuine causal structure if that structure lives at the micro level and is invisible at the aggregate. Agent-based models are not a replacement for statistical inference — they are a different tool, operating on a different level of description, and asking a different question. The combination of both is what produces reliable insight.
Building a market simulation or calibration pipeline? I’d be glad to help.