Skip to content

The dissipation meter — currents and entropy production

What it reads. Whether the strategic dynamics are in thermodynamic equilibrium or merely stationary. The Glauber-logit chain (a random player revises to a logit choice over pure payoffs) has a stationary distribution π over joint profiles; the meters are the probability current \(J^*(a,a') = \pi(a)w(a{\to}a') - \pi(a')w(a'{\to}a)\) and the entropy production rate

\[\sigma_{\mathrm{EP}} = \tfrac12 \sum_{a,a'} \big[\pi(a)w(a{\to}a') - \pi(a')w(a'{\to}a)\big]\,\log\frac{\pi(a)w(a{\to}a')}{\pi(a')w(a'{\to}a)} \;\ge\; 0 .\]

Why it means something (tier: exact — K3 + Schnakenberg network theory). In an exact potential game the logit revision probabilities are the heat-bath conditionals of \(\pi \propto e^{\lambda\Phi}\): the chain is reversible, \(J^* = 0\), \(\sigma_{\mathrm{EP}} = 0\) — strategic equilibrium is thermodynamic equilibrium. Off potentiality, detailed balance breaks: the system holds a non-equilibrium steady state that continuously circulates (Edgeworth-style cycling lives here) and dissipates.

Calibration state. Gate dynamics.exact:

Reading Requirement Artifact
π vs e^{λΦ}/Z, congestion n=2/3, λ ∈ {0.7, 1.2, 2} ≤ 10⁻¹⁰ gibbs_agreement.json
Potential games EPR and |J*| < 10⁻¹² equilibrium_reads_zero.json
RPS, matching pennies EPR > 10⁻³, circulation present ness_reads_positive.json
1,000-game α sweep at fixed λ = 1.2 ρ(EPR, α) > 0.9 and ρ(EPR, ℛ) > 0.8 chain_comovement.json

Measured, honestly split: marginally ρ(EPR, α) = 0.990 and ρ(EPR, ℛ) = 0.993 — but the second number is α-driven. Stratified by α, within-level ρ(EPR, ℛ) is +0.80..+0.88 up to α ≈ 0.65, degrades above, and reverses to −0.355 at α = 0.95 (per-level values in the artifact): among near-pure-harmonic games the reciprocity and dissipation meters decouple — the first realisation of C1's falsifier, and the programme's first genuine finding (F-0004). The mechanism is now measured (F-0007, artifact decoupling_mechanism.json): at α = 0.95, ℛ's within-level variation is denominator-driven (ρ(ℛ, 1/‖χ+χᵀ‖) = 0.993) — and the numerator alone also decouples from EPR (ρ = −0.37), refuting the natural 'use the numerator' correction. Interpretation (argued, not proved): the response layer is a local derivative at the QRE point — a marginal-strategy object — while EPR is a functional of the joint-profile stationary flux; nothing identifies one with the other except in potential games, where both vanish. The measurement is consistent with that reading; the tier of the correlations is derived, of the reading itself, argument. Scope table: ℛ answers is this system potential? (exact, any λ, any α); EPR answers how hard is it circulating?; use ℛ levels comparatively only for α ≲ 0.65 — measured at λ = 1.2 on 3×3 two-player families; the crossover's λ- and size-dependence is open (C1 remainder).

Artifacts regenerate via uv run python -m experiments.dynamics_calibration (make reproduce).

Limitations, stated once. This is the exact, generator-level meter: it needs the full profile space, so it runs on small games. Real-data dissipation goes through the trajectory estimators (KLD, TUR bound, NEEP — a later unit) which are validated against this meter on synthetic ground truth before touching anything empirical. The co-movement result is rank-order evidence at fixed λ, on 3×3 two-player families.

Frontier refinements (unit science.frontier, 2026-08-11)

The α = 0 criticality peak is a pure λ×payoff-scale fold (σ(λ, s·u) = σ(sλ, u) exactly; verified across scales 1/2/4). The supercritical frontier λ_c(α) is monotone descending — 7.8 at α = 0.55 to 3.0 at α = 0.80, with no median crossing below α ≈ 0.5 by λ = 15 at 40-game sampling. And F-0004's reversal decomposes (F-0010): the within-level coupling COLLAPSE at α = 0.95 is universal; the sign is λ-dependent across the three λ tested (+0.25 → +0.03 → −0.23; −0.26 at 4×4), each individual sign within ~2 null-SD — the initial criterion failed 2/4 and the revision is on the record. λ_c(α) is the verified-unique crossing of the fixed-set median-ρ curve (single-crossing checked per level); the 5-game phase map's α=0.5 onset vs this 40-game onset between 0.50–0.55 is recorded as sampling variability.

Driving the system: the housekeeping/excess split (unit thermo.protocols, F-0012)

The instantaneous entropy production of any distribution p splits exactly as σ_tot = σ_hk + σ_ex (Hatano–Sasa / Esposito–Van den Broeck): the housekeeping part is the fuel a NESS burns just to exist (identically zero on potential games — detailed balance), the excess part is relaxation, σ_ex = −d/dt D(p‖π). For stepwise λ-quenches the Hatano–Sasa IFT ⟨e^{−Y}⟩ = 1 holds for every game — including off detailed balance, where Jarzynski does not apply — — but note the honesty caveat (red-team O-1): the exact weighted-transfer computation returns 1 by algebraic telescoping for stepwise protocols, so it is an identity verification, not a bug-catcher; the correctness test is the independent sampled-trajectory CI, powered at α = 0.5 where ⟨Y⟩ ≈ 10⁻² (protocol_ift_checks.json; the α = 0.95 read is retained but low-power).

The pre-registered reading (predictions written in config before the run; the intended prior commit aborted on a hook failure — F-0012 records this plainly) found the driving-cost inversion (F-0012): across the α family, excess quench dissipation collapses (0.036 → 3×10⁻⁵ nats) while accumulated housekeeping grows to 12.4 nats, exactly linear in protocol duration (σ_hk depends only on the generator and its π, so constancy per hold is structural — the informative number is the burn rate and its growth with α). The collapse's MECHANISM is open: the first-pass 'λ-insensitive NESS' explanation was refuted by a red-team probe on an asymmetric mix whose NESS genuinely moves. Potential games pay per change and drive for free in the quasi-static fine-step limit (⟨Y⟩ ∝ 1/K, measured slopes ≈ −1); harmonic games pay rent per unit time whether driven or not. Artifacts: protocol_quench_scan.json, protocol_ift_checks.json.

Fast quenches (unit science.quench_regimes, F-0014). When relaxation is cut short, excess dissipation leaves the floor and rises toward the frozen divergence D(π_start‖π_end) — the τ→0 telescoping identity. The second EFE campaign found single-spectral-gap interpolation between the two exact limits adequate to ~0.1 dex (held-out median; 0.46 max — the multi-mode remainder is open). Scope guards (red-team round 2): the frozen limit is exact only AT τ = 0 and its approach is non-uniform — loop-like NESS paths (π_end ≈ π_start with excursions between, e.g. α = 0.95 with a long ramp) break the interpolation by ~1 dex, and that consumed-probe failure is reported alongside the held-out numbers, not averaged into a median. Run-1's verdict was winner_failed_validation — the machine stopped confidently after one probe and the pre-registered held-out guard refused it; min_probes in run_campaign is that lesson institutionalised as a stopping gate (validation stays with the held-out guard). Artifact: fast_quench_campaign.json.

The working quench model (unit science.quench_multimode, F-0015). The third campaign replaced the global crossover with a per-step, path-aware recursion: p_k = π_k + (p_{k−1} − π_k)e^{−g_k τ}, accumulating Y along the tracked distribution. It dominates the F-0014 formula on both median (0.073 vs 0.114 dex) and worst case (0.44 vs 1.11 — the loop-path failure is structural to endpoint comparison and vanishes when the path is followed). Honest negative in the same artifact: naively adding the second eigenmode makes things worse — the off-detailed-balance generator is non-normal, and the principled multi-mode treatment (bi-orthogonal) is parked until a use-case needs sub-0.4-dex crossover accuracy. Artifact: quench_multimode_campaign.json (full hypothesis × probe residual table).

Quench dissipation from data (unit thermo.hs_estimator, F-0016). The plug-in Ŷ estimator needs only observed state windows — and its development is a case study in diagnostics: the natural self-calibration (the IFT ⟨e^{−Ŷ}⟩ = 1) turned out measurably insufficient (a 45% bias hid behind an IFT of 1.01; exponential-average errors cancel where mean errors don't — both failed runs are in git). A per-window relaxation gate looked like the fix — until the second red-team round broke that too (safety factors are game-dependent, coverage fails across seeds, the gate collapses on single-trajectory data). The unit is OPEN and the module is banner-marked experimental: the honest state of the art is that plug-in Hatano–Sasa estimation from quench data is not yet certifiable, and the measured failure map (four independent modes, all in F-0016) is the deliverable. One robust qualitative consequence stands: on ramps into concentration (the F-0006 gap collapse, 0.88 → 0.15 across λ 0.5 → 5.5) hold times must vastly exceed anything practical — data-side quench estimation is hardest exactly where rationality concentrates. Artifact: hs_estimator_sweep.json (void for certification, kept for the record).

Can a month of data be quoted? No — and the reason is instructive (unit thermo.hs_estimator.smalln, F-0019). R8 asked whether the certified estimator's n ≥ 200 floor could come down to the ~30 trajectories a month of market data provides. It cannot: no interval construction passes the registered criteria below the existing floor. But the binding constraint is not the number — interval coverage holds at every n down to 20 (19–20/20, width ≈ 1.9× the bootstrap SE, for percentile, bootstrap-t and the heuristic t-widened alike). What fails is the instrument's own decision machinery: agreement with a reference gate decision runs 10/20 at n = 20–30 and only reaches 17/20 at n = 200, and the anomaly flag still flips under a physically-null permutation 3–6 times in 20 below n = 200. So the estimate converges roughly an order of magnitude faster than the verdict that certifies it — which is precisely why F-0017's monthly reads flickered. Two mandated diagnostics then sharpened this. Against the generator's exact settling status (no sampling noise) n = 200 scores 19/20 and passes — so the earlier 17/20 was the finite-sample comparator's own noise and the existing floor is solid, while small-n still fails (10/20 at n = 30) regardless of reference. And a within-group permutation that preserves the SE split's composition collapses the flag flips from 6 → 0 at n = 30 and 3 → 0 at n = 100: almost all of the small-n instability comes from the estimator's i::4 split reassigning trajectories, not from the physics. The refusal is therefore implementation-bound — a split-independent SE (jackknife or bootstrapped τ̂) may stabilise the gate at n ≈ 30–50 and bring monthly windows back within reach through better machinery rather than more data. One tempting shortcut is closed off explicitly: you may not borrow a full-window settling verdict and apply it to a sub-window. A joint verdict over 211 day-pairs does not imply that January settled — a slow subset can hide inside a passing aggregate — so a month may be quoted only as an estimate with its interval and an explicit "gate stability unvalidated at this n" warning, never with a verdict. Artifact: smalln_certification.json. The gate's order-dependence is fixed; the small-n floor is not (unit thermo.hs_estimator.gate_se, F-0020). R8 left a sharp lead: the small-n instability looked like an artifact of one implementation choice — a 4-way i::4 trajectory split used to estimate the relaxation time's standard error — because a permutation preserving that split's composition made the flag flips vanish. R9 replaced it with three order-invariant estimators (leave-one-out jackknife, an analytic delta method, and trajectory bootstrap) and the lead was correct: flag flips go to zero at every n (0/20 against the incumbent's 7/6/0 at n = 30/50/100), and on the real CAISO panel the flip rate under physically-null day-order shuffles falls from 0.214 to 0.050, with one month collapsing 17/20 → 0/20. The SE method even moves admission: a month refused under the incumbent is admitted under all three alternatives, with the estimate provably identical. And yet the floor does not move. Agreement with the exact settling status at n = 30 fails for every candidate — the best, trajectory bootstrap, reaches 12/20 against an 18 bar while passing all four other criteria. The reason is worth stating precisely, because it retracts half of what R8 hoped: the relaxation-time estimate has an across-seed SD of ~35–40% of itself at n = 30, so an SE that is accurate (bootstrap sits 0.18–0.29 from an independently measured oracle, against the incumbent's 0.44–0.49) must report that variance and therefore must refuse holds that are genuinely settled. Flag stability was implementation-bound and is solved; gate accuracy at small n is a property of the relaxation-time estimator itself, and the next lever is a lower-variance τ̂ (multi-lag fitting, or pooling across windows), not a better error bar. Two by-products: the analytic delta method is dead on arrival (its dτ/dρ gradient explodes at the ρ clip floor, and it overstates the SE ~3× at n = 30 by dropping a π̂ term that partly cancels match-rate noise), and a second, independent order-dependence turned up outside R9's scope — the CI/IFT bootstrap draws resample indices from a fixed seed, so permuting trajectories re-rolls the interval while leaving the estimate exact to six decimals. That one cannot be fixed by reseeding: the anomaly flag is a hard boolean thresholded on a Monte-Carlo interval, and when the bound sits at the threshold the noise has to be reported rather than hidden. Artifacts: gate_se_read.json, gate_se_realdata.json.