The Lagging Truth

The Escalation Cost: Intensity, Duration, and the Growing Damage of Regime Change — the plain-English companion

A short explainer covering the same ground as the article below. The written companion carries the full detail.

Companion to the research paper of the same title. This is education, not investment advice. Nothing here tells you what to buy, sell, or predict. It explains what the paper found, how the checking worked, and what the results do and do not mean.

The claim in one sentence

When the world changes, your trailing measurements stay blind for a while — and the damage you take during that blind stretch equals the intensity of the change raised to the power of its duration. Both numbers can be computed in advance, from three quantities any organization already has.

It started as a banking problem

This result began as a special case. An earlier paper in this series examined a banking regulation — a rule that reads a twenty-year trailing average of credit (the average of the last twenty years, always trailing behind the present) and tells banks when to build rainy-day capital. Under realistic conditions, that rule amplified the very cycles it was built to damp. The surprise came afterward: the failure turned out not to be about banking at all.

It was an instance of something general. Whenever an organization measures a slow-moving variable with a trailing average, and then acts on that measurement in a way that feeds back into the variable itself, the loop can turn self-reinforcing instead of self-correcting. Banks do this. Factories do this. Retailers, pension funds, and government agencies do this. The banking paper had found one room on fire; this paper is about the building.

And it asks the question the banking paper couldn’t: not just whether a loop is safe, but what it costs when the world shifts underneath it. Picture an apartment building that swaps its old boiler for a new one overnight. Every thermostat in the building was tuned to the old boiler’s behavior. For a while, the readings still reflect a machine that no longer exists. Engineers call that a regime change — the rules governing a system changed, and everything calibrated to the old rules is now quietly wrong. This paper computes the bill for that quiet stretch.

The shower with the slow pipe

Start with the one idea everything else stands on: some numbers have momentum. Ask what tomorrow’s temperature will be, and “about the same as today” is a good guess — weather carries an echo of itself from day to day. Demand for a factory’s product works the same way: a busy month tends to be followed by another busy month. The strength of that echo is called persistence. A persistence near zero means each period starts fresh, like coin flips. A persistence near one means the past hangs on hard, like a heavy flywheel that keeps spinning.

Now put a decision-maker in the loop. A companion paper in this series, The Measurement Trap, worked out when such loops are safe, and its result is easiest to feel in a shower with a slow pipe. The water takes a while to respond to the knob. Turn the knob gently and you settle at a comfortable temperature. Turn it aggressively and you scald, overcorrect, freeze, and overcorrect again — oscillation. That paper proved there is a precise speed limit: the product of how strongly your measurement amplifies persistent movement and how hard you push back must stay under a fixed number, π²/2, about 4.93. Below the limit, corrections settle. Above it, correction becomes the oscillation.

For any such loop, mathematics offers a single verdict number — an amplification factor, written ρ (the technical name is the spectral radius). Each period, the gap between where the system is and where it should be gets multiplied by ρ. Below one, gaps shrink on their own; the system heals. Above one, gaps grow; the system feeds on itself. There is no middle.

Two sciences, each holding half the answer

Here is the gap this paper fills, and it is worth seeing plainly, because the paper’s whole contribution lives in it.

One field — control theory, the engineering of feedback — answers the question: is this loop stable at its current settings? It hands you the amplification factor and the speed limit. But it speaks about a world whose rules hold still.

A second field — adaptive filtering, the science of how measurements catch up — answers a different question: after the rules change, how long until your measurement reflects the new reality? It prices the lag. But it doesn’t tell you what that lag does to a system that is meanwhile acting on stale numbers.

Neither field alone prices the dangerous stretch in between: the rules have changed, your dashboard hasn’t noticed, and you are steering the new world with the old world’s map. The paper’s third theorem is the bridge — an exact identity welding the two fields into one computable quantity. That is why this paper sits underneath the whole series: every other paper in the program is a case of this one.

Damage compounds like interest

The central theorem prices the blind stretch. Call the amplification factor before the change ρ₁ and after the change ρ₂ — a regime change in the dangerous direction means the loop that was healing (below one) starts compounding (above one). Call the length of the blind stretch τ — the number of periods before your trailing measurement absorbs the new reality, which depends directly on how long your measurement window is.

The theorem: damage during the blind stretch is bounded by the intensity ratio raised to the duration.

Damage ≤ (ρ₂ ÷ ρ₁)^τ

That exponent is the entire point, and it is why the title says escalation. Damage during the blind period doesn’t add up — it compounds, the way unpaid interest compounds. Ten percent worse for ten blind periods is not “twice as bad as ten percent for five periods.” It is far worse, because each period’s damage multiplies the last. A modest change endured for a long blind stretch can out-damage a violent change caught quickly. Anyone who has watched a small unpaid balance snowball already understands the mathematics of this theorem.

And every ingredient is computable before the change happens, from three quantities an organization already tracks: how persistent its key variable is, how long its measurement window is, and how hard it pushes on the gap it measures.

How long should your rearview mirror be?

The window deserves its own moment, because it is the parameter organizations control most directly and think about least.

A measurement window is how much history you average over — how long your rearview mirror is. A store manager estimating demand from the last four weeks has a four-week window; a regulator reading a twenty-year trend has a thousand-week one. Short windows are jumpy: they chase every random wiggle, and a policy steering by them lurches. Long windows are smooth but sleepy: after the world changes, a long window keeps averaging in the dead regime for a long time. The blind stretch τ is the price of the long window.

The paper’s second theorem proves this tradeoff has a unique best answer — a single optimal window length, given by an exact formula, balancing estimation jitter against regime-change blindness. Not “shorter is better” or “longer is better,” but a computable sweet spot that moves in stated directions: more persistent variables and costlier regime changes push the best window shorter.

The theorem’s sharpest edge, though, is institutional, and the paper states it plainly: institutions rarely decide their measurement windows. They inherit them — from vendor software defaults, reporting calendars, and procedures nobody remembers choosing. Then a regime change arrives, and the institution discovers what the inheritance costs. The paper even notes the reverse calculation the identity licenses. An organization that doesn’t know its own effective window can infer it from its own crisis history — how long past deviations lasted, and how much they grew, brackets the window that produced them. The paper offers that as a principle the mathematics permits, not a validated method, and says so.

What happened when the theory met data

A theorem about damage should predict actual damage. The paper tests this the hard way — with the rules for judging every test written down and locked before the data was touched. Think of a pool player calling the pocket before the shot: pre-registration means the paper cannot quietly move the target after seeing where the ball went. Where a rule had to be amended, the amendment is dated and disclosed, and the amended rule was still locked before the relevant data run.

The 2008 crisis. Across seventeen U.S. sectors, the paper computed each sector’s predicted damage and compared it against how far each sector’s inventory-to-sales ratio — the classic gauge of supply-chain stress — actually swung. The comparison is a rank test: line the sectors up by predicted damage, line them up again by realized swing, and score how well the two orderings agree, from −1 (opposite) through 0 (unrelated) to +1 (identical). The score came out 0.3456, and the locked rule graded the episode a support. The paper also asked whether the full formula beat its own ingredients used alone. Its frozen rule left both follow-up comparisons unresolved — one sat at a probability of 0.0515 against a 0.05 line, a hair’s width from resolving, and the rule declined to round that up to a win.

The thirty-four-year test. The primary make-or-break test rolled the diagnostic forward through thirty-four years, always predicting with data available at the time and grading against what came after — never letting the model peek at the years it was being tested on. Across the nine sectors that genuinely oscillate around the stability boundary, pooled rank agreement was 0.1505, with a 0.0090 probability of arising by chance. Verdict: support. The paper immediately fences what this proves: the ordering of damage across sectors is validated; the compounding exponent itself is not, because testing the exponent would need variation these series cannot provide. The paper says so in the same breath as the result.

COVID, the planned null. Before checking, the paper registered COVID as a case where the theory should find nothing — and the reason is the theory’s scope in one story. The theorem prices one specific event: persistence rises, a decaying loop starts compounding, lag becomes expensive. COVID was not that event. Demand collapsed and rebounded through many channels at once, and in fourteen of seventeen sectors measured persistence fell. A framework that claimed damage rankings there would be detecting crises in general rather than the one mechanism it prices. The null arrived as registered — rank agreement 0.0760, chance probability 0.3933. The diagnostic declined to fire on the most famous supply-chain disruption in living memory, because that disruption is not the kind of event it prices.

The Beer Game. The Beer Game is a famous supply-chain simulation — a chain of players ordering from each other, in which small demand wiggles amplify into huge upstream swings. The paper ran it as a Monte Carlo: thousands of repetitions with fresh random demand each time, every competing strategy facing the identical demand sequence within each run, so differences between strategies are never differences in luck. Against a standard keep-the-shelf-topped-up ordering rule, a simple tool built on the diagnostic — damp your ordering only when measured persistence crosses 0.83 — cut mean cost from 122,962.38 to 122,729.71. Small in percentage (0.19 percent) but statistically solid (paired probability 0.0005), winning 72.2 percent of head-to-head runs; the full theorem-guided policy did better still, at 122,487.37. The paper brackets it: these are properties of this simulation’s own construction, and an earlier claim tying the baseline to an external industry figure was withdrawn rather than defended.

The map of where it helps — and where it hurts

The paper’s most useful pages for a practitioner may be its least flattering ones. Four simulation studies were built specifically to find the edges where acting on the diagnostic stops helping and starts hurting. They found four, and each is reported with every cell of the grid shown — including the unresolved cells — because a grid filtered to its favorable cells is a search dressed up as an experiment.

Value is conditional on breathing room. A previously reported result — that the tool turns beneficial in longer supply chains — survived re-testing only where capacity had headroom. Re-run at five times the original precision, the tight-capacity cells show the tool harmful at every chain length tested. More than half of the original grid had been unresolved noise read as signal.

Pricing advice is a cliff, and one-directional. Using the diagnostic to guide price raises pays off where capacity strain is moderate — and turns net-negative above a strain threshold, a cliff rather than a slope. Using it to recommend price cuts lost money in every environment tested: −581.36, −1,238.29, and −2,315.25 across the three demand regimes. The defensible reading, in the paper’s own words, is narrow: the calculation offers meaningful guidance on when not to cut prices, and no comparable license to cut them.

Permanent damage splits the verdict. Some markets punish mistakes permanently: customers who leave don’t come back. The economist’s word is hysteresis. There, the raise strategy survives when the regime has genuinely shifted, and fails in merely noisy regimes — the measurement’s own jitter switches the policy on and off, and each false alarm bleeds customers permanently.

And the deepest edge: some situations cannot be rescued by better measurement at all. When persistence doesn’t jump once but keeps drifting, the diagnostic-driven policy loses to the crudest alternative on the menu — a fixed damping setting, no measurement, no calculation, nothing. Here is what sharpens that finding: the paper also tested a version handed the true persistence value at every step, a perfect measurement no real firm could ever have. The perfect version loses too (paired contrasts −1.0688 and −1.5125 against the fixed rule). If noisy measurement were the problem, perfect measurement would fix it; it doesn’t. The limitation is in the recipe — a rule tuned to a snapshot of a moving target — and no precision reaches a parameter that will not hold still. The paper then names exactly which drift shapes were and were not tested, so the boundary of the claim is a list, not a shrug.

Three questions any organization can answer today

Strip the framework to essentials and it is an audit any organization can run on itself. Three questions: How persistent is the variable you steer by? How long is the window you steer with? How hard do you push on the gap you measure? Those three numbers determine the amplification factor; the amplification factor determines whether your deviations decay or compound; and the pair — intensity and blind duration — prices what the next regime change costs you while your measurements catch up.

Two properties make this an audit rather than an academic exercise. All three numbers are observable to the institution itself, from data it already has. And two of the three are policy choices: an organization that cannot change how persistent its environment is can still change how long it looks and how hard it reacts.

The framework then ranks the levers, and the ranking is the practical heart of the paper. Interventions that shorten the blind window — faster reporting, higher-frequency data, estimating current conditions instead of waiting for confirmed history — attack the exponent directly. They pay off exponentially, because the window’s length is the power everything else is raised to. Interventions that soften the push — damping, smoothing, rate limits — can pull a loop back inside the stability boundary, at the price of responsiveness, and the paper states the conditions under which that trade is worth making. Interventions that reduce persistence itself — simplifying the supply base, adding buffer capacity, pooling demand — are the most durable and the slowest to build, because they change the environment rather than the reaction to it.

Read as a checklist for a chief executive: know your three numbers; treat your measurement window as a decision, not an inheritance; and when investing in stability, remember that speed of information compounds while everything else merely adds.

This audit is becoming a tool. A public calculator is planned at LaggingTruth.com/diagnostic. You will paste in a stretch of your own history — monthly demand, orders, whatever series your decisions actually track. It will hand back your three numbers, your amplification factor against the 1.0 line, and your history replayed with the blind windows marked around your roughest stretches — where this framework would have warned you. It will describe the dynamics of the numbers you give it — education, not advice — and a signup form for launch notification sits with the first public bet, below.

The chips moment

The paper’s supply-chain evidence lands on a live policy question: the United States is currently spending tens of billions to rebuild domestic semiconductor manufacturing, and nations rebuilding industrial capacity are making decades-long choices about chain structure right now.

What the paper contributes is a map of where instability concentrates. Along the U.S. goods chain — retail sales to wholesalers to manufacturers’ shipments to manufacturers’ new orders — it measures how much the swings grow at each step. The chain amplifies 2.52 times end to end, and the growth is not evenly spread: the sharpest jump, 2.37 times, happens at the final step, where shipments become new factory orders. Exclude the COVID years and end-to-end amplification is 5.20 times; for durable goods, 7.48. The whiplash lives upstream, at exactly the tiers that on-shoring policy proposes to rebuild.

Semiconductors specifically: the paper had pre-registered a claim that CHIPS-dependent sectors would sit at the very peak of the instability ranking, and dropped it with the reason stated. They sit in the top cluster, but exact rank depends on the measurement specification — the fully written-out recipe for how the ranking is computed — and one such recipe floors eight sectors into an unrankable tie. What survives is blunter than a ranking. Under the primary specification, the semiconductor sector’s estimated amplification factor sits above the stability boundary at every level of capacity utilization measured — that is, however close to full tilt its factories were running, from slack conditions (1.0499) to near-full capacity (1.0780), with utilization currently reading 75.35. A hunt for a utilization threshold that flips the sector across the boundary found none — the paper reports that as unadjudicable rather than converting it into a convenient answer.

Why would chip-making be structurally prone to this? The complexity literature supplies a mechanism the paper endorses carefully: products with many interdependent components, chains with many tiers, and densely interconnected supply bases propagate a disturbance through more paths and hold it longer. In this paper’s terms, complexity is a mechanism for persistence, and persistence is the input to the damage bound. That is why the sectors this framework ranks as unstable and the sectors that literature calls complex are substantially the same list, arrived at from different directions. The paper offers that as a coherence check, not a tested causal claim, and no experiment in it estimates complexity’s effect.

Put the pieces together and the policy reading writes itself, within the paper’s stated limits. A country laying down new manufacturing capacity is choosing, for decades, the three parameters this framework audits. How many tiers will its chains have? How much complexity will each product carry? And how fast will information move through the system it is building? That last parameter is the cheapest to decide well right now. The framework’s ranking says the information choice compounds.

What might shortening the blind window look like in physical terms? One everyday illustration: a chain whose suppliers sit a truck ride away can often run on a faster rhythm than one strung across an ocean — orders confirmed against tomorrow’s delivery rather than next month’s container, production adjusted weekly rather than quarterly. Nearer, simpler supply bases also tend toward fewer tiers, which speaks to the persistence lever as well. To be plain about the boundary: the paper never tests geography, and distance is no guarantee — a far supplier with live data can outrun a near one with paperwork. The illustration is of the levers the framework prices, not a measured result. The paper also flags, explicitly as exploration, a financing corollary. If an ecosystem’s instability is structural, stabilizing it means investing in the tiers that carry the persistence, not only the visible final-assembly stage — and the composition of credit, not merely its quantity, determines whether that investment happens.

Two public bets

Two dated, falsifiable predictions ride with the paper, registered publicly so a miss cannot be quietly redefined.

First, a standing claim any firm can test on itself: compute your amplification factor from your three numbers — demand persistence, measurement window, reaction strength. Above 1.0, your response to the next demand shock amplifies; below, it decays.

Second, a sector-level bet extracted mechanically from the committed analysis: the seventeen-sector panel splits into nine boundary-crossing sectors and eight never-crossing ones, both lists registered verbatim. At the next trigger event, the flagged nine should swing harder than the quiet eight under the registered test. The paper registers an awkwardness rather than hiding it. Under its committed classification, the CHIPS-dependent computers sector lands in the never-crossing class while wholesale machinery lands in the flagged one — an earlier informal sketch that said otherwise is superseded, in writing.

What the checking caught in this paper

Every paper in this series runs through the same general machinery — hash-pinned inputs, a machine-checked ledger of numbers, a verification program, an adversarial review. That machinery is described once in the series’ shared verification note, which follows every companion on LaggingTruth.com. What belongs here is what the process caught and changed in this paper specifically.

Three failures of the checking itself are disclosed in the record. The manuscript reached mid-draft with no Methods section while every automated gate stayed green — the gate compared the manuscript against an outline that shared the same hole, and two artifacts that agree can agree on an omission. One experiment’s pre-registered decision rule was mis-calibrated: measured after the fact, it could not have fired at the effect actually present in the data. The rule’s measured insensitivity is disclosed in the paper, alongside a correctly targeted secondary analysis whose reading was committed before its script existed. And during the review round, a commit went through while the verification gate was failing, because the procedure read the wrong program’s success signal and reported a pass — the gate was right, the reading of it was wrong. Each episode produced a specific new check that now runs on every build.

The adversarial review — a memory-isolated session working from a curated package, recomputing the full ledger before reading — confirmed eighteen findings. Sixteen were fixed and two rebutted in writing, one cure arriving as a pre-registered amendment whose unresolved result the paper reports exactly as the frozen rule dictates. A boundary defect in the primary test’s final evaluation point was found by the author side running the code, invisible from every output the reviewer could reach, and disclosed unprompted. A second review round over the corrected package recomputed the enlarged ledger with zero mismatches and a byte-identical rebuild.

Beyond the review: two pre-registered cross-domain extensions — sovereign ratings and unemployment insurance — were withdrawn when their preconditions failed, and are reported as characterizations, not findings. The capacity-threshold test is reported as unadjudicable. And the paper’s own monitoring dashboard, run against history, confirms regime shifts two to five months after they begin — the paper’s thesis applied to its own apparatus, and reported as such.

What this cannot do

  • It does not detect crises in general. COVID is the registered counterexample: a multi-channel shock where persistence fell is outside the mechanism, and the diagnostic correctly found nothing there.
  • The compounding exponent is not empirically validated. The data validates the intensity ordering; the exponent is proved, not measured, and the paper says which is which.
  • The Beer Game savings are model-bound. They are properties of that simulation’s construction, with no claim of transfer to any real firm’s systems.
  • Price guidance is one-directional and conditional. When not to cut prices, under moderate capacity strain — no license to raise on the wrong side of the cliff, and none to cut anywhere tested.
  • Under drifting persistence, the recipe fails — even with perfect measurement. A fixed rule beat both the measured policy and one handed the true parameter; only one drift shape was tested, and the untested shapes are named.
  • The complexity link is coherence, not cause. No experiment here estimates complexity’s effect on persistence; the financing nexus is flagged as exploratory with no ledgered quantity.
  • Its own monitor is reactive. Two to five months behind the onsets it was checked against — a dashboard, not an early-warning system.
  • The predictions may take years to grade. Both wait on a trigger event; by rule, no trigger means untestable and carried forward, never a pass.

The takeaway

Two mature sciences each held half of a question every organization faces: one could say whether a feedback loop is stable where it sits, the other how fast a measurement catches up once the world moves. This paper welds them and prices what lives in between — the blind stretch where damage compounds like unpaid interest, as intensity raised to duration. The theory survived its called shots: it ordered the 2008 damage across seventeen sectors, held up through thirty-four years of out-of-sample rolling, and correctly found nothing in COVID, the crisis it was scoped not to price. Its edges are mapped and named — a cliff in the pricing advice, a one-way license, and a drift regime where no measurement, however perfect, rescues a rule tuned to a world that won’t hold still. What remains is an audit of three numbers any organization already has, and a ranking of levers in which faster information pays off exponentially. And — at the moment a nation is pouring concrete for the next fifty years of its manufacturing base — a map showing that the whiplash concentrates in exactly the tiers being rebuilt. Two public bets now wait to grade it.


This companion is licensed CC BY-NC 4.0. The research paper it accompanies is licensed CC BY-NC-ND 4.0, and the analysis and verification code is MIT-licensed. Education, not advice: nothing in this document is financial advice, an investment recommendation, or a forecast.