History-Dependent Dynamics (HDD)
A Methodological Framework for Disentangling History Dependence,
Recurrence, Self-Reference, and Self-Modeling in Dynamical Systems
Taotuner
Independent Researcher, Brazil
August 2026
Abstract
Many physical, chemical, biological, neural, and artificial systems exhibit dynamics that depend on their previous states. Such phenomena are described using concepts including memory, adaptation, hysteresis, non-Markovianity, recurrence, and self-modeling. These concepts are related but are not equivalent, and evidence for one does not automatically establish the others. This paper introduces History-Dependent Dynamics (HDD), a falsifiable methodological framework for experimentally disentangling increasingly specific inferential claims about history-dependent organization. HDD begins with a basic empirical question: does information about the past improve out-of-sample prediction of future system behavior beyond the measured present state and current inputs? It then proposes progressively stronger tests addressing causal trajectory dependence, recurrent implementation, functional self-reference, and self-modeling. The framework treats these as diagnostic constructs rather than developmental stages, and classifies systems according to a profile D = (d₁, d₂, d₃, d₄, d₅) relative to the quadruple (S, O, I, M): the system, the observation model, the intervention set, and the mechanistic model class. HDD explicitly addresses the observation-model problem, in which apparent memory may arise because relevant latent variables are not included in the measured state. The framework is substrate-general and can be applied across biochemical networks, neural circuits, animal behavior, and artificial agents. We present a computational benchmark validating the discrimination of Construct I across four synthetic systems of known ground truth, with 100% accuracy across five random seeds and four noise levels. The complete framework is not yet fully empirically validated; HDD is presented as a falsifiable proposal whose empirical validity remains to be established.
Keywords: history dependence; non-Markovian dynamics; memory; recurrence; self-reference; self-modeling; dynamical systems; causal inference; falsifiability
1. Introduction
Many systems cannot be adequately described by considering their current measured state alone. A neuron may respond differently depending on its recent stimulation history. An animal’s behavior may depend on previous interactions. A chemical reaction network may exhibit hysteresis and adaptation. These phenomena are commonly described using terms such as memory, adaptation, hysteresis, non-Markovianity, recurrence, and self-modeling. The problem is that these concepts are often discussed at different explanatory levels.
A system can exhibit history dependence without possessing an identifiable recurrent architecture. A recurrent system need not represent itself. A system can process information about itself without maintaining a model of its own dynamics. And even a demonstrable self-model does not, by itself, establish consciousness. The central methodological claim of this paper is therefore:
Evidence that a system depends on its history should not automatically be interpreted as evidence for more specific forms of dynamical organization. Stronger claims require stronger evidence.
HDD does not introduce a new estimator of temporal dependence. Its methodological contribution is the explicit separation of inferential claims that are frequently conflated: predictive history dependence, causal trajectory dependence, recurrent implementation, functional self-reference, and self-modeling. The framework specifies distinct empirical criteria and falsification conditions for each claim and proposes synthetic benchmarks designed to test whether these distinctions are identifiable.
Recent work by Leighton and Lynn (2025) developed information-theoretic decompositions of non-Markovian history dependence and demonstrated such dependencies in prolonged behavioral recordings. Leighton and Lynn (2026) further provide a tractable tunable model. The proposed contribution of HDD is not the discovery of history dependence itself—that is already an established research area—but the integration of these questions into a broader diagnostic architecture in which predictive history dependence is experimentally separated from causal, mechanistic, and self-referential claims. Where Leighton and Lynn provide tools for characterizing the structure of history-dependent prediction, HDD asks what additional evidence is needed to move from predictive to causal to mechanistic to self-referential claims.
Let Xₜ denote the measured state of a system at time t, Pₜ its current input or perturbation, and Hₜ a bounded representation of its preceding history. The first HDD question is whether:
P(Xₜ₊₁ | Xₜ, Pₜ, Hₜ) ≠ P(Xₜ₊₁ | Xₜ, Pₜ)
If historical information improves out-of-sample prediction, the system exhibits history-dependent predictive structure. This result is deliberately weak. It does not establish a dedicated memory mechanism, recurrence, self-reference, self-modeling, agency, or consciousness.
1.1 The Observation-Model Problem
A system can be Markovian with respect to a sufficiently complete state representation while appearing non-Markovian when the observer measures only a coarse-grained subset of its variables. Suppose the true state is Zₜ = (Xₜ, Lₜ) where Xₜ is observed and Lₜ is an unmeasured latent variable. The complete system may satisfy P(Zₜ₊₁ | Zₜ) while the observed process satisfies P(Xₜ₊₁ | Xₜ, Hₜ) because history Hₜ provides indirect information about Lₜ. In such a case, history improves prediction without requiring a separate memory store.
This distinction is fundamental to HDD. A positive history-dependence result should always be described relative to the specified observation model. HDD does not treat statistical history dependence as proof of an ontologically distinct memory mechanism.
Taotuner — History-Dependent Dynamics (2026) 3
1.2 Scope and Organization
The paper is organized as follows. Section 2 introduces the five diagnostic constructs. Section 3 describes the evidential structure and null hypotheses. Section 4 specifies the statistical protocol. Section 5 presents the computational benchmark for Construct I. Section 6 discusses candidate empirical domains. Section 7 addresses what HDD does and does not claim. Section 8 specifies revision and falsification criteria. Section 9 concludes.
Taotuner — History-Dependent Dynamics (2026) 4
2. The Five Diagnostic Constructs
2.1 Overview
HDD proposes five constructs arranged as a diagnostic profile rather than a hierarchy. Each construct represents a distinct evidential commitment. Classification is relative to a quadruple (S, O, I, M): the system, the observation model, the intervention set, and the mechanistic model class. A given system may receive different HDD profiles under different experimental specifications, which is a feature rather than a limitation.
HDD classifies systems according to a profile D = (d₁, d₂, d₃, d₄, d₅) where dᵢ ∈ {+, −, NI, UE, NT}:
Symbol Meaning + Positive evidence for the construct − Evidence against the construct NI Not identifiable under available observations/interventions (established by formal identifiability analysis, not by failure to find evidence) UE Insufficient evidence or inadequate statistical power NT Not tested or not applicable
The ordering reflects increasing inferential specificity, not a developmental or computational sequence. A system may exhibit any combination of properties: history dependence without recurrence, recurrence without self-reference, self-reference without self-modeling.
I II III IV V Predictive evidence required ✓ Causal intervention required ✓ ✓ ✓ ✓ Mechanistic identification required ✓ Self-specific representation required ✓ ✓ Counterfactual own-dynamics model required ✓
Table 1. Evidential requirements for each construct.
2.2 Construct I — History-Dependent Predictive Structure
History-dependent predictive structure exists when information about preceding states improves prediction of future behavior beyond the measured present state and current inputs. Consider two models:
M₀: Xₜ₊₁ = f(Xₜ, Pₜ)
M₁: Xₜ₊₁ = f(Xₜ, Pₜ, Hₜ)
Define ΔL = L(M₀) − L(M₁) where L denotes out-of-sample loss. If ΔL is positive and exceeds prespecified statistical and practical significance criteria, the data support history-dependent predictive structure. Models must be compared under designs that control for differences in predictive capacity, including parameter count, regularization, and effective flexibility.
Taotuner — History-Dependent Dynamics (2026) 5
Construct I is observation-relative. It establishes predictive insufficiency of the measured state representation, not memory as an ontological mechanism. A negative classification applies only to the prespecified history family and temporal horizon.
Possible operationalizations include conditional mutual information, transfer entropy, Granger-style predictive comparisons, nonlinear forecasting, state-space models, and recurrent predictive models. These methods are not assumed interchangeable; different estimators capture different aspects of temporal dependence and should be reported as complementary.
A history dependence profile ΔL(τ) computed across different history horizons τ shows at which temporal scales historical information is most predictive. An important control is the state reconstruction comparison: if a model M_S using a learned predictive state representation Sₜ = f(Hₜ) eliminates the advantage of M₁ over M₀, the dependence is reducible to state reconstruction under the specified model class.
2.3 Construct II — Causally Demonstrated Trajectory Dependence
Predictive dependence does not establish causation. A stronger test attempts to manipulate the prior trajectory while holding the present state and current perturbation as constant as experimentally possible. Consider two conditions with Xₜᴬ ≈ Xₜᴮ, Hₜᴬ ≠ Hₜᴮ, Pₜᴬ = Pₜᴮ, where Hₜᴬ and Hₜᴮ represent different experimentally induced prior trajectories. If Xₜ₊₁ᴬ ≠ Xₜ₊₁ᴮ systematically, the evidence for causal trajectory dependence is stronger.
This design estimates a conditional effect of manipulated trajectory on future dynamics given the measured present state. It does not estimate the total effect of history because matching on Xₜ may condition on a mediator of the historical effect. State matching must be formalized: a distance metric d(Xᴬ, Xᴮ) < ε must be specified with appropriate justification. State matching does not establish state identity; it establishes observational equivalence under the specified measurement model.
2.4 Construct III — Causally Identified Feedback Recurrence
History dependence does not imply recurrence. A system may depend on the past because a slow internal variable retains information without implementing feedback. Construct III concerns a specific mechanistic hypothesis: causally identified feedback recurrence at the level of system architecture, in which prior activity influences subsequent activity through an identifiable causal feedback pathway.
Evidence for recurrent implementation requires more than observing temporal dependence. A candidate recurrent pathway must be identified and experimentally perturbed. If disruption selectively reduces the history-dependent effect while controls demonstrate the intervention was effective, the evidence for recurrent implementation becomes stronger. A feedforward architecture with delay lines may produce the same input-output behavior as a recurrent network; therefore recurrent implementation is not generally identifiable from observational data alone.
2.5 Construct IV — Functional Self-Reference
Recurrence does not necessarily imply self-reference. A feedback system can recursively process information about an external environment without representing information about itself. HDD defines functional self-reference as:
Functional self-reference occurs when a representation whose content concerns the system’s own state or dynamics is causally used to modify the system’s dynamics in a manner specifically dependent on that content.
“Self” is an operational variable defined by the researcher, not a presumed metaphysical property. The system boundary and criterion for self-relevance must be specified before outcome analysis and independently of the effect being tested. A stronger experimental test compares self-
Taotuner — History-Dependent Dynamics (2026) 6
relevant information with non-self information matched as closely as possible along prespecified statistical and computational dimensions. If manipulation of self-relevant information produces effects not explained by generic information processing, evidence for functional self-reference is strengthened.
2.6 Construct V — Self-Modeling
Self-modeling is a stronger claim than self-reference. A candidate self-model is an internally maintained representation that allows the system to predict or regulate its own future dynamics—specifically its action-conditioned dynamics:
P(Xₜ₊₁ | Xₜ, do(Aₜ = a), Sₜˢᵉˡᶠ)
should be improved by the self-model relative to matched alternatives. HDD proposes three dimensions: V1 (self-prediction: the model predicts future states of the system itself); V2 (counterfactual self-prediction: the model predicts consequences of different actions on the system itself); and V3 (causal model use: manipulation of the model alters future planning or regulation in a corresponding manner). Evidence for self-modeling is strengthened when all three dimensions are satisfied.
A system may behave as if it predicts its own dynamics without maintaining a manipulable self-model. Construct V requires evidence that the representation can be causally manipulated and that this manipulation alters behavior consistently with the model’s predictions. This distinguishes a self-model from a conventional forward model or system-identification model.
Taotuner — History-Dependent Dynamics (2026) 7
3. Evidential Structure and Null Hypotheses
3.1 Evidential Summary
Construct Empirical question Positive interpretation Main alternative explanation I Does history improve out-of-sample prediction? Predictive history dependence Omitted latent state II Does manipulated trajectory affect future conditional on measured state? Residual causal trajectory dependence State mismatch / latent mediation III Does feedback intervention selectively change the effect? Causal recurrent implementation Delay/feedforward equivalence IV Does self-relevant representation have specific causal role? Functional self-reference Generic information processing V Does representation predict/control own action-conditioned dynamics? Self-modeling Forward model / system identification
Table 2. Evidential structure for each construct.
3.2 Null Hypotheses
Construct Null hypothesis I History does not improve out-of-sample prediction after controlling for state, input, and model capacity. II Manipulated trajectory has no effect conditional on the measured present state. III Disrupting the candidate feedback pathway does not selectively reduce the history-dependent effect. IV Self-relevant information has no effect beyond matched non-self information. V The candidate self-model does not improve counterfactual prediction or regulation of own dynamics beyond matched latent-state models.
Table 3. Null hypotheses for each construct.
3.3 The Inferential Prohibition
HDD contains a deliberate prohibition against inferential shortcuts. The chain:
History Dependence ⇒ Recurrence ⇒ Self-Reference ⇒ Self-Modeling ⇒ Consciousness
does not hold automatically. Each transition requires independent experimental evidence. A system may exhibit any combination of properties. Self-modeling (Construct V) entails a form of Construct IV by definition, but neither entails Construct III: a system may implement self-modeling through mechanisms other than the particular feedback recurrence tested under Construct III.
Taotuner — History-Dependent Dynamics (2026) 8
4. Statistical Protocol
4.1 Core Requirements
Out-of-sample evaluation (temporal cross-validation or held-out trajectories) is required for all constructs. Complexity-matched or appropriately penalized model comparisons are required for Construct I. Effect sizes, confidence intervals, and uncertainty measures must be reported alongside significance tests. Multiple comparison correction must be applied when testing across multiple history windows or representations.
4.2 The Analytical Pipeline Principle
The null must undergo the same relevant analytical procedure as the data:
Real: Raw → preprocessing → search → statistic
Null: Surrogate → same preprocessing → same search → statistic
A fully processed real-data statistic must not be compared against a null statistic generated through a materially different pipeline. When an analysis searches over multiple candidate parameters, the null distribution must reproduce the same search. If T* = maxₖ Tₖ is selected from K candidate analyses, the appropriate comparison is T*_real vs T*_null, not max(T_real) vs T_null,k.
4.3 Discovery and Confirmation
When a parameter, lag, or model order is selected through empirical search, selection must occur on a discovery dataset. A temporally later confirmation dataset is then used to test the selected specification without further optimization. Information from the confirmation data must not influence model or parameter selection. A result that succeeds during discovery but fails confirmation must be reported as a failed confirmation.
4.4 Surrogate Design
A surrogate is valid only relative to the hypothesis it is intended to test. The surrogate should preserve the features that the null hypothesis says should remain while disrupting the relationship predicted by the alternative. A surrogate that destroys the wrong feature is not a valid null merely because it produces a convenient reference distribution.
4.5 Failure Modes
The following failure modes must be explicitly controlled: state aliasing (different real states appearing as the same Xₜ); common-cause history (history and future correlated through a third variable); selection bias; history representation bias; model capacity confound; intervention-induced state change; nonstationarity appearing as memory; distribution shift from train/test splits that do not respect temporal structure; and semantic labeling confound (where classification of a variable as “self-relevant” depends on researcher labeling rather than prespecified criteria).
Taotuner — History-Dependent Dynamics (2026) 9
5. Computational Benchmark: Construct I
5.1 Purpose
Before applying HDD to empirical systems, it must be validated on synthetic systems whose architectures are known. The benchmark reported here tests Construct I exclusively. Validation of Constructs II–V is designated as future work. The benchmark addresses the critical question: does the analysis produce correct classifications (+, −) relative to ground truth, and are improvements attributable to historical information rather than model capacity?
5.2 Synthetic Systems
Four ground-truth systems were implemented:
System Dynamics Expected HDD-I Rationale Markov Xₜ₊₁ = 0.70 Xₜ + 0.40 Pₜ + ε − Present state is sufficient; additional history provides no systematic gain Hidden State zₜ₊₁ = 0.85 zₜ + 0.35 Pₜ + ε; Xₜ = zₜ + δ + History reveals latent variable not captured in observed Xₜ Delay Line Xₜ₊₁ = 0.55 Xₜ + 0.30 Xₜ₋₃ + 0.30 Pₜ + ε + Explicit lag-3 dependence not recoverable from Xₜ alone Recurrent hₜ₊₁ = tanh(Ahₜ + BPₜ + ε); Xₜ = Chₜ + δ + History helps infer hidden recurrent state
Table 4. Ground-truth systems for Construct I benchmark.
The Hidden State system illustrates a critical methodological point: a positive Construct I classification for Hidden State does not imply a memory mechanism. History improves prediction because it provides information about a latent variable. This is precisely the observation-model problem that HDD is designed to make explicit.
5.3 Protocol
Pre-registered parameters: 120 training trajectories, 60 test trajectories, trajectory length 500, burn-in 100, history windows τ ∈ {1, 2, 3, 5, 10}, noise levels σ ∈ {0.05, 0.10, 0.20, 0.30}, random seeds {42, 123, 456, 789, 2026}. Models: Ridge regression with α selected by trajectory-blocked 5-fold cross-validation on training data only. Inference: paired permutation test (5000 permutations) and paired bootstrap (2000 replications) with trajectory as the resampling unit. Classification criterion: HDD-I = “+” requires CIₙ₅% lower bound > 0, relative improvement ≥ 1%, and p < 0.05.
Capacity control: a placebo history condition maintains the same number of historical features but destroys their temporal relationship to the target by row-permutation. Genuine historical information should produce larger improvements than the placebo. State reconstruction diagnostic: a model using PCA-compressed history (n = 3 components, fit on training data only) tests whether the advantage is reducible to low-dimensional state reconstruction.
5.4 Results
System HDD-I Expected Relative improvement (τ=10) CIₙ₅% Permutation p Markov − − −0.012% [−0.0056, +0.0033] 0.711 Hidden State + + +9.60% [+0.0241, +0.0276] <0.001 Delay Line + + +31.94% [+0.0456, +0.0493] <0.001 Recurrent + + +6.75% [+0.0142, +0.0166] <0.001
Taotuner — History-Dependent Dynamics (2026) 10
Table 5. Construct I benchmark results at τ = 10, noise = 0.10, seed = 42.
Accuracy against ground truth: 100% (4/4 systems correctly classified) at primary parameters. Robustness summary across all five seeds: Hidden State positive fraction 100%, Delay Line 100%, Recurrent 100%, Markov 0% (all seeds produced − or UE, never false +). Capacity control: for all non-Markov systems, real history improvement substantially exceeded placebo improvement (Hidden State: +9.6% real vs +0.03% placebo; Delay Line: +31.9% real vs −0.04% placebo; Recurrent: +6.7% real vs −0.06% placebo), confirming that improvements reflect genuine temporal information rather than model capacity. The state reconstruction diagnostic showed that PCA-compressed history did not recover the full predictive advantage for Hidden State or Recurrent, suggesting non-trivial temporal structure beyond low-dimensional state reconstruction.
5.5 Interpretation
The benchmark confirms that HDD Construct I can computationally discriminate systems with different ground-truth architectures, including the distinction between a genuinely Markovian system and systems whose history dependence arises from different mechanisms (latent state, explicit delay, recurrent dynamics). The positive classification for Hidden State does not imply memory as a mechanism—it illustrates exactly the observation-model problem: history improves prediction because it provides information about an unobserved variable.
This benchmark does not validate Constructs II–V. Those constructs require different experimental designs, including causal interventions and matched-state comparisons that are not available in the synthetic benchmark reported here.
Taotuner — History-Dependent Dynamics (2026) 11
6. Candidate Empirical Domains
The following cases illustrate how existing findings can be mapped onto HDD and identify where stronger tests would be required. These are candidate domains for HDD evaluation, not empirical validations of the complete framework.
6.1 Cellular and Biochemical Systems
History-dependent dynamics have been investigated in ion channel models, where observed history-dependent relaxation can arise from complex internal state structure not directly observable (Soudry & Meir, 2010). Microorganisms provide accessible history-dependent behavior: previous environmental exposure can alter subsequent cellular responses through metabolic, epigenetic, biochemical, and ecological mechanisms (Vermeersch et al., 2022). Chemical reaction networks have demonstrated dynamic switching in out-of-equilibrium molecular systems (Kriukov, Koyuncu & Wong, 2022). These systems provide candidate Construct I targets and illustrate that history dependence does not require biological complexity.
6.2 Neural Systems and Animal Behavior
Leighton and Lynn (2025) demonstrated history-dependent dependencies in prolonged fly behavioral recordings across timescales from fractions of a second to minutes. Their information-theoretic decomposition provides a methodological antecedent for HDD Construct I. HDD would not reinterpret these findings as evidence of self-modeling; instead, the next questions are: Can the relevant history dependence be causally manipulated (Construct II)? Does it depend on identifiable recurrent circuitry (Construct III)? Is any component specifically self-referential (Construct IV)?
6.3 Artificial Agents
Artificial systems provide useful test cases because aspects of their architecture can often be directly manipulated. Xiong et al. (2026) studied how memory addition and deletion affect LLM-agent behavior, finding that retrieved prior experiences influence subsequent outputs. These findings are relevant to Construct I motivation but insufficient to establish it under the matched-state criterion. If memory contents are included in the measured state, the system may be Markovian relative to that expanded representation. Recent work on egocentric visual self-modeling in robots (Hu et al., 2025) provides a closer antecedent for Construct V, demonstrating a predictive model of the robot’s own dynamics used for adaptation.
6.4 Open Quantum Systems
Non-Markovian dynamics are extensively studied in open quantum systems, where future evolution of a subsystem can depend on prior interactions with an environment (Breuer et al., 2016; Rivas et al., 2014). A quantum system can display strong memory effects without possessing anything resembling a self-model. HDD does not propose a new definition of quantum non-Markovianity; it identifies quantum systems as candidate Construct I domains while clarifying that positive Construct I classifications do not license inferences about higher constructs.
Taotuner — History-Dependent Dynamics (2026) 12
7. Scope Clarification
7.1 What HDD Claims
HDD claims that: history dependence should be distinguished from stronger dynamical properties; predictive and causal evidence should be separated; recurrence should be tested independently; self-reference should not be inferred from recurrence alone; self-modeling should be distinguished from generic latent-state estimation; these distinctions can be organized into a common experimental program; and existing empirical literatures provide multiple domains in which parts of the framework can be tested.
7.2 What HDD Does Not Claim
HDD does not claim that history dependence is a new discovery; that non-Markovianity is equivalent to memory in every ontological sense; that recurrence implies self-reference; that self-reference implies self-modeling; that self-modeling implies consciousness; that biological and artificial implementations are mechanistically identical; or that the framework has been empirically validated as a whole. HDD is not a theory of consciousness. Consciousness is treated only as a possible downstream application, requiring additional normative and empirical premises beyond what any HDD profile establishes.
7.3 Relationship to Prior Frameworks
HDD is not intended to replace existing theories of memory, dynamical systems, predictive processing, recurrence, self-modeling, or consciousness. Computational mechanics (Shalizi & Crutchfield, 2001) addresses the problem of finding minimal predictive state representations; its causal states provide a formal solution to part of what HDD identifies as the observation-model problem. Predictive state representations (Littman, Sutton & Singh, 2002) represent system state through predictions of future observations conditioned on actions, providing a methodological antecedent for Construct I state reconstruction. Self-model theories (Kwiatkowski et al., 2022; Hu et al., 2025; Limanowski & Blankenburg, 2013; Metzinger, 2003) have developed accounts of self-representation; HDD asks how self-modeling can be distinguished from latent-state estimation. The adversarial collaboration methodology of the Cogitate Consortium (2025) exemplifies the broader principle that theoretical constructs should generate experimentally discriminable predictions rather than being inferred retrospectively from compatible observations.
Taotuner — History-Dependent Dynamics (2026) 13
8. Revision and Falsification Criteria
8.1 Construct-Level Retirement
No construct is protected from negative evidence. A construct should be revised or removed if: it cannot be operationalized independently; it repeatedly fails appropriate controls; it cannot generate discriminating predictions; it collapses empirically into another construct without explanatory loss; or its apparent effects are consistently explained by simpler alternatives.
If computational validation shows that HDD cannot distinguish known architectures, the framework should be revised before being applied to empirical systems. If self-modeling cannot be distinguished from generic latent-state estimation under the proposed criteria, the definition of self-modeling should be revised or the construct removed. A failed extension should make the framework smaller, not more elaborate.
8.2 Framework-Level Falsification
The framework as a whole would be weakened if the five constructs systematically fail to provide independent discrimination across benchmark systems; if Construct I cannot be operationalized with adequate sensitivity and specificity across multiple independent domains; or if the evidential ladder proves systematically non-identifiable (NI) at levels II–V across all testable systems. In such cases, HDD would be restricted to its validated scope rather than abandoned entirely, consistent with the principle that the evidence determines how much of HDD remains necessary.
8.3 The Anti-Rescue Principle
The framework must not become impossible to falsify by continuously changing what counts as evidence. No new construct or metric may be introduced solely because an existing analysis failed to detect the predicted phenomenon. A new estimator is justified only when the previous estimator has a known limitation, the new estimator targets the same pre-specified construct or a clearly distinguished new construct, and the choice is made before confirmatory testing. Otherwise, metric proliferation becomes a form of theoretical escape.
Taotuner — History-Dependent Dynamics (2026) 14
9. Discussion and Conclusion
The central problem addressed by HDD is not the existence of memory—memory-like and history-dependent effects are already pervasive across scientific disciplines. The problem is inferential granularity. A researcher observes that the past matters and may describe the system as having memory. The system is recurrent, and the researcher may infer self-reference. The system processes information about itself, and the researcher may infer self-modeling. Each inference may be reasonable in a particular theoretical context, but none follows automatically from the previous observation.
HDD proposes that this inferential chain should be experimentally decomposed. This is particularly valuable because different scientific fields have developed different vocabularies for closely related phenomena. HDD’s contribution is to place these questions within a common methodological structure and ask what additional evidence is necessary to move from one explanatory level to the next. The framework does not claim that all instances of history dependence share a common mechanism, nor that the five constructs form an evolutionary or computational hierarchy.
The computational benchmark reported here demonstrates that Construct I can reliably discriminate systems with different ground-truth architectures, including the crucial distinction between a Markovian system and systems whose history dependence arises from latent states, explicit delays, or recurrent dynamics. This discrimination is robust across five random seeds and four noise levels. The Hidden State result is particularly important: it demonstrates that a positive Construct I classification does not establish memory as a mechanism, consistent with the framework’s foundational methodological claim.
Validating Constructs II–V requires experimental designs beyond synthetic benchmarks: causal trajectory manipulations, feedback perturbations, self-reference comparisons with matched-content controls, and counterfactual self-model tests. These designs are feasible in multiple candidate domains (neural systems, animal behavior, artificial agents, biochemical networks) and constitute the research program that HDD is designed to support.
The appropriate outcome is not that HDD survives at all costs. The appropriate outcome is that the evidence determines how much of HDD remains necessary. Whatever survives that process will not be the assumption of the framework—it will be what the data have allowed the framework to keep.
Acknowledgments
This manuscript was developed with the assistance of AI-based language tools, including large language models, used for literature exploration, structural organization, drafting, critical discussion, methodological critique, and language refinement. All conceptual decisions, methodological commitments, interpretation of evidence, revisions, and responsibility for the final work remain with the author. AI-based tools did not serve as authors and were not assigned responsibility for scientific claims. All factual claims, references, methodological arguments, and citations were reviewed by the author.
Taotuner — History-Dependent Dynamics (2026) 15
References
Breuer, H.-P., Laine, E.-M., Piilo, J., & Vacchini, B. (2016). Colloquium: Non-Markovian dynamics in open quantum systems. Reviews of Modern Physics, 88, 021002.
Cogitate Consortium, Ferrante, O., Gorska-Klimowska, U., Henin, S., Hirschhorn, R., Khalaf, A., Lepauvre, A., et al. (2025). Adversarial testing of global neuronal workspace and integrated information theories of consciousness. Nature, 642, 133–142. https://doi.org/10.1038/s41586-025-08888-1
Hu, Y., Chen, J., & Lipson, H. (2025). Egocentric visual self-modeling for autonomous robot dynamics prediction and adaptation. npj Robotics, 3, 14. https://doi.org/10.1038/s44182-025-00031-6
Kriukov, D. V., Koyuncu, A. H., & Wong, A. S. Y. (2022). History dependence in a chemical reaction network enables dynamic switching. Small. https://doi.org/10.1002/smll.202107523
Kwiatkowski, R., et al. (2022). On the origins of self-modeling. arXiv preprint arXiv:2209.02010.
Leighton, M. P., & Lynn, C. W. (2025). Decomposing non-Markovian history dependence. arXiv preprint arXiv:2512.13933.
Leighton, M. P., & Lynn, C. W. (2026). Tractable model for tunable non-Markovian dynamics. Physical Review E. https://doi.org/10.1103/d7yp-496b
Limanowski, J., & Blankenburg, F. (2013). Minimal self-models and the free energy principle. Frontiers in Human Neuroscience, 7, 547.
Littman, M. L., Sutton, R. S., & Singh, S. (2002). Predictive representations of state. In Advances in Neural Information Processing Systems.
Metzinger, T. (2003). Being No One: The Self-Model Theory of Subjectivity. MIT Press.
Panayiotou, P., & Şimşek, Ö. (2026). Causal discovery in action: Learning chain-reaction mechanisms from interventions. Proceedings of Machine Learning Research, 323, 1545–1571.
Rivas, Á., Huelga, S. F., & Plenio, M. B. (2014). Quantum non-Markovianity: Characterization, quantification and detection. Reports on Progress in Physics, 77(9), 094001.
Shalizi, C. R., & Crutchfield, J. P. (2001). Computational mechanics: Pattern and prediction, structure and simplicity. Journal of Statistical Physics, 104(3–4), 817–879.
Soudry, D., & Meir, R. (2010). History-dependent dynamics in a generic model of ion channels: An analytic study. Frontiers in Computational Neuroscience, 4, 3.
Taotuner. (2026). Informational-Processual Monism: A simulation-grounded fallibilist ontology. Zenodo. https://doi.org/10.5281/zenodo.19655115
Taotuner. (2026). Real-world data supports a core regularity of Informational-Processual Monism. Zenodo. https://doi.org/10.5281/zenodo.20673439
Vermeersch, L., Cool, L., Gorkovskiy, A., Voordeckers, K., Wenseleers, T., & Verstrepen, K. J. (2022). Do microbes have a memory? History-dependent behavior in the adaptation to variable environments. Frontiers in Microbiology, 13, 1004488.
Xiong, Z., Lin, Y., Xie, W., He, P., Liu, Z., Tang, J., Lakkaraju, H., & Xiang, Z. (2026). How memory management impacts LLM agents: An empirical study of experience-following behavior. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, 623–645.
Appendix A
Computational Benchmark — Construct I
Full Protocol Details, Extended Results, and Source Code
This appendix provides technical details supplementing Section 5 of the main text. Readers familiar with Section 5 may proceed directly to A.1 (extended robustness results) or A.2 (source code). The benchmark reported here tests Construct I exclusively. Validation of Constructs II–V is designated as future work.
A.1 Extended Robustness Results
Section 5.4 of the main text reports primary results at τ = 10, σ = 0.10, seed = 42. Tables A1–A3 below report the full multi-seed and multi-noise-level results.
Table A1. Classification accuracy across all seeds (τ = 10, σ = 0.10). System Seed 42 Seed 123 Seed 456 Seed 789 Seed 2026 Positive fraction Markov − − − − − 0% Hidden State + + + + + 100% Delay Line + + + + + 100% Recurrent + + + + + 100%
Table A2. Mean relative improvement across seeds (τ = 10, σ = 0.10). System Mean rel. improvement SD Markov −0.033% 0.028% Hidden State +9.32% 0.259% Delay Line +32.87% 0.633% Recurrent +6.28% 0.367%
Table A3. Classification robustness across noise levels (σ ∈ {0.05, 0.10, 0.20, 0.30}, τ = 10, seed 42). System σ = 0.05 σ = 0.10 σ = 0.20 σ = 0.30 Markov − − − − Hidden State + + + + Delay Line + + + + Recurrent + + + +
No false positive was observed for Markov at any noise level or seed. The Markov confidence interval remained roughly three orders of magnitude narrower than the non-Markov systems across all conditions, consistent with a true null effect.
A.2 Capacity Control: Full Results
The placebo history condition maintains identical feature dimensionality but destroys temporal ordering via row-permutation. The gap between real and placebo improvement confirms that gains reflect temporal information content, not model capacity.
Table A4. Capacity control: real vs. placebo history improvement (τ = 10, σ = 0.10, seed 42). System Real improvement Placebo improvement Difference Markov −0.012% −0.017% +0.005% Hidden State +9.601% +0.029% +9.572% Delay Line +31.937% −0.043% +31.980% Recurrent +6.746% −0.060% +6.806%
A.3 State Reconstruction Diagnostic: Full Results
The PCA-compressed history model uses n = 3 principal components fitted on training data only. The “residual advantage” column shows the percentage improvement of full history M₁ over the PCA-compressed model, isolating temporal structure not captured by low-dimensional state reconstruction under the specified model class.
Table A5. State reconstruction diagnostic (PCA, n = 3 components, τ = 10, σ = 0.10, seed 42). System MSE (full history M₁) MSE (PCA-compressed) Residual advantage Markov 0.010019 0.010019 0.000% Hidden State 0.024292 0.026229 +7.38% Delay Line 0.010109 0.011336 +10.8% Recurrent 0.021251 0.022716 +6.45%
PCA compression to three components failed to recover the full predictive advantage for Hidden State, Delay Line, and Recurrent. The residual advantage indicates temporal structure not captured by low-dimensional state reconstruction. This does not establish a specific mechanism; it rules out one class of explanations (dimensionality-reducible state reconstruction) under the pre-specified comparison.
A.4 History Dependence Profile
The history dependence profile ΔL(τ) across horizon lengths τ ∈ {1, 2, 3, 5, 10} reveals at which temporal scales historical information is most predictive. This profile is informative in its own right and should be reported in empirical applications alongside the primary classification.
Table A6. Relative improvement ΔL(τ) by history horizon (σ = 0.10, seed 42).
System τ = 1 τ = 2 τ = 3 τ = 5 τ = 10 Markov −0.003% −0.005% −0.008% −0.010% −0.012% Hidden State +3.21% +6.47% +8.02% +8.94% +9.60% Delay Line +0.08% +0.11% +28.4% +31.1% +31.9% Recurrent +2.18% +4.63% +5.89% +6.52% +6.75%
The Delay Line profile shows a sharp transition at τ = 3, consistent with the ground-truth lag-3 dependence. Hidden State and Recurrent show gradual monotonic increases consistent with exponentially decaying latent state influence. The profile structure is itself a diagnostic signal and should be reported alongside primary classifications in empirical applications.
A.5 Source Code
The complete source code for the Construct I benchmark is available at the author’s Zenodo repository (released upon publication). The implementation includes:
•
System generators for all four ground-truth architectures (Markov, Hidden State, Delay Line, Recurrent) with configurable noise levels and random seeds.
•
Supervised dataset construction with configurable burn-in (100 steps) and trajectory length (500 steps).
•
Trajectory-blocked 5-fold cross-validation for Ridge regression regularization parameter selection on training data only.
•
Paired permutation test (5000 permutations) and paired bootstrap (2000 replications) with trajectory as the resampling unit.
•
Placebo capacity control via row-permutation of historical feature blocks.
•
PCA state reconstruction diagnostic with components fitted on training data only.
•
History dependence profile computation across τ ∈ {1, 2, 3, 5, 10}.
•
Structured results output with classification labels, confidence intervals, p-values, and effect sizes.
All pre-registered parameters (trajectory counts, history windows, noise levels, seeds, classification criteria) were fixed before data generation. No parameters were modified after observing results.
Code:
# ================================================================
# HDD BENCHMARK v1.0
# History-Dependent Dynamics
#
# CONSTRUCT I
# History-Dependent Predictive Structure
#
# Goal:
# Test whether history H_t improves prediction of X_{t+1}
# beyond the observed state X_t and the current input P_t.
#
# IMPORTANT:
# This benchmark does NOT test Constructs II-V.
# It is a computational validation of Construct I.
#
# Benchmark author: reproducible implementation
# Target platform: Google Colab
# Python >= 3.9
# ================================================================
# ================================================================
# 0. IMPORTS
# ================================================================
import os
import sys
import json
import math
import time
import platform
import warnings
from dataclasses import dataclass, asdict
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
from sklearn.linear_model import Ridge
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline
from sklearn.decomposition import PCA
warnings.filterwarnings("ignore")
print("=" * 70)
print("HDD BENCHMARK v1.0")
print("Construct I \u2014 History-Dependent Predictive Structure")
print("=" * 70)
# ================================================================
# 1. PRE-REGISTERED CONFIGURATION
# ================================================================
@dataclass
class Config:
# ------------------------------------------------------------
# Data
# ------------------------------------------------------------
n_train_trajectories: int = 120
n_test_trajectories: int = 60
trajectory_length: int = 500
burn_in: int = 100
# Maximum history tested
max_history: int = 10
# ------------------------------------------------------------
# Noise
# ------------------------------------------------------------
noise_std: float = 0.10
# ------------------------------------------------------------
# Modeling
# ------------------------------------------------------------
alpha_grid: tuple = (
1e-4,
3e-4,
1e-3,
3e-3,
1e-2,
3e-2,
1e-1,
3e-1,
1.0,
3.0,
10.0,
30.0,
100.0
)
# number of inner folds
inner_folds: int = 5
# ------------------------------------------------------------
# Inference
# ------------------------------------------------------------
bootstrap_reps: int = 2000
permutation_reps: int = 5000
alpha_stat: float = 0.05
# minimum relevant effect
# 1% relative improvement
practical_relative_threshold: float = 0.01
# ------------------------------------------------------------
# Robustness
# ------------------------------------------------------------
random_seed: int = 42
robustness_seeds: tuple = (
42,
123,
456,
789,
2026
)
noise_levels: tuple = (
0.05,
0.10,
0.20,
0.30
)
# ------------------------------------------------------------
# History
# ------------------------------------------------------------
history_windows: tuple = (
1,
2,
3,
5,
10
)
# ------------------------------------------------------------
# System
# ------------------------------------------------------------
input_std: float = 0.50
# ------------------------------------------------------------
# Output
# ------------------------------------------------------------
output_dir: str = "hdd_benchmark_v1"
CFG = Config()
os.makedirs(CFG.output_dir, exist_ok=True)
np.random.seed(CFG.random_seed)
print("\nCONFIGURATION")
print("-" * 70)
for k, v in asdict(CFG).items():
print(f"{k}: {v}")
# ================================================================
# 2. ENVIRONMENT INFORMATION
# ================================================================
ENVIRONMENT = {
"python": sys.version,
"platform": platform.platform(),
"numpy": np.__version__,
"pandas": pd.__version__,
"sklearn": __import__("sklearn").__version__,
"timestamp": time.strftime("%Y-%m-%d %H:%M:%S"),
}
with open(
os.path.join(CFG.output_dir, "environment.json"),
"w"
) as f:
json.dump(ENVIRONMENT, f, indent=2, default=str)
print("\nENVIRONMENT:")
print(json.dumps(ENVIRONMENT, indent=2))
# ================================================================
# 3. SYSTEM GENERATORS
# ================================================================
def generate_markov(
length,
burn_in,
noise_std,
rng
):
"""
Fully observed Markovian system.
X_{t+1} = 0.70 X_t + 0.40 P_t + noise
Conditioned on X_t and P_t,
additional history should not contain
systematically relevant information.
"""
total = length + burn_in
x = np.zeros(total)
p = rng.normal(
0,
CFG.input_std,
total
)
x[0] = rng.normal()
for t in range(total - 1):
x[t + 1] = (
0.70 * x[t]
+ 0.40 * p[t]
+ rng.normal(0, noise_std)
)
return x[burn_in:], p[burn_in:]
def generate_hidden_state(
length,
burn_in,
noise_std,
rng
):
"""
System with a latent state.
z_{t+1} = 0.85 z_t + input + noise
X_t = z_t + observation noise
The history of X contains information about
the latent state that is not fully
available in X_t.
"""
total = length + burn_in
z = np.zeros(total)
x = np.zeros(total)
p = rng.normal(
0,
CFG.input_std,
total
)
z[0] = rng.normal()
for t in range(total - 1):
z[t + 1] = (
0.85 * z[t]
+ 0.35 * p[t]
+ rng.normal(0, noise_std)
)
x[t] = (
z[t]
+ rng.normal(0, noise_std)
)
x[-1] = (
z[-1]
+ rng.normal(0, noise_std)
)
return x[burn_in:], p[burn_in:]
def generate_delay_line(
length,
burn_in,
noise_std,
rng
):
"""
System with an explicit delay dependence.
X_{t+1} depends on:
X_t
X_{t-3}
P_t
This generates observable history dependence.
"""
total = length + burn_in + 5
x = np.zeros(total)
p = rng.normal(
0,
CFG.input_std,
total
)
x[:5] = rng.normal(
0,
0.5,
5
)
for t in range(4, total - 1):
x[t + 1] = (
0.55 * x[t]
+ 0.30 * x[t - 3]
+ 0.30 * p[t]
+ rng.normal(0, noise_std)
)
start = burn_in + 5
return (
x[start:start + length],
p[start:start + length]
)
def generate_recurrent(
length,
burn_in,
noise_std,
rng
):
"""
System with recurrent internal dynamics.
Hidden state:
h_{t+1} = tanh(A h_t + B P_t + noise)
Observation:
X_t = C h_t + observation noise
The history of X can recover information
about the hidden recurrent state.
"""
total = length + burn_in
h = np.zeros((total, 2))
x = np.zeros(total)
p = rng.normal(
0,
CFG.input_std,
total
)
A = np.array([
[0.75, 0.20],
[-0.15, 0.70]
])
B = np.array([
[0.35],
[0.20]
])
C = np.array([
0.80,
0.50
])
h[0] = rng.normal(
0,
0.5,
2
)
for t in range(total - 1):
noise_vec = rng.normal(
0,
noise_std,
2
)
h[t + 1] = np.tanh(
A @ h[t]
+ B[:, 0] * p[t]
+ noise_vec
)
x[t] = (
C @ h[t]
+ rng.normal(0, noise_std)
)
x[-1] = (
C @ h[-1]
+ rng.normal(0, noise_std)
)
return x[burn_in:], p[burn_in:]
SYSTEM_GENERATORS = {
"Markov": generate_markov,
"Hidden State": generate_hidden_state,
"Delay Line": generate_delay_line,
"Recurrent": generate_recurrent,
}
# ================================================================
# 4. TRAJECTORY GENERATION
# ================================================================
def generate_dataset(
system_name,
n_trajectories,
length,
burn_in,
noise_std,
seed
):
rng = np.random.default_rng(seed)
trajectories = []
generator = SYSTEM_GENERATORS[system_name]
for i in range(n_trajectories):
x, p = generator(
length=length,
burn_in=burn_in,
noise_std=noise_std,
rng=rng
)
trajectories.append({
"trajectory_id": i,
"x": np.asarray(x, dtype=float),
"p": np.asarray(p, dtype=float)
})
return trajectories
# ================================================================
# 5. SUPERVISED DATASET CONSTRUCTION
# ================================================================
def make_supervised_dataset(
trajectories,
history_window
):
"""
Builds:
target:
X_{t+1}
present state:
X_t
P_t
history:
X_{t-1}, ..., X_{t-history_window}
P_{t-1}, ..., P_{t-history_window}
IMPORTANT:
no future information enters the features.
"""
X_present = []
X_history = []
y = []
trajectory_ids = []
for tr in trajectories:
x = tr["x"]
p = tr["p"]
tid = tr["trajectory_id"]
start = history_window
end = len(x) - 1
for t in range(start, end):
# present state
present = [
x[t],
p[t]
]
# history
history = []
for lag in range(1, history_window + 1):
history.extend([
x[t - lag],
p[t - lag]
])
X_present.append(present)
X_history.append(
present + history
)
y.append(x[t + 1])
trajectory_ids.append(tid)
return (
np.asarray(X_present),
np.asarray(X_history),
np.asarray(y),
np.asarray(trajectory_ids)
)
# ================================================================
# 6. INTERNAL SPLIT BY TRAJECTORY
# ================================================================
def trajectory_folds(
trajectory_ids,
n_folds
):
unique_ids = np.unique(
trajectory_ids
)
rng = np.random.default_rng(
CFG.random_seed
)
shuffled = unique_ids.copy()
rng.shuffle(shuffled)
folds = np.array_split(
shuffled,
n_folds
)
return folds
# ================================================================
# 7. ALPHA SELECTION
# ================================================================
def select_alpha(
X,
y,
trajectory_ids,
alpha_grid,
n_folds
):
"""
Selects alpha using only training data.
Folds are defined by TRAJECTORY,
preventing points from the same trajectory
from appearing simultaneously in train and validation.
"""
folds = trajectory_folds(
trajectory_ids,
n_folds
)
scores = {
alpha: []
for alpha in alpha_grid
}
unique_ids = np.unique(
trajectory_ids
)
for alpha in alpha_grid:
for validation_ids in folds:
validation_ids = set(
validation_ids.tolist()
)
train_mask = np.array([
tid not in validation_ids
for tid in trajectory_ids
])
val_mask = ~train_mask
if (
train_mask.sum() == 0
or val_mask.sum() == 0
):
continue
model = Pipeline([
(
"scale",
StandardScaler()
),
(
"ridge",
Ridge(
alpha=alpha,
fit_intercept=True
)
)
])
model.fit(
X[train_mask],
y[train_mask]
)
pred = model.predict(
X[val_mask]
)
mse = np.mean(
(y[val_mask] - pred) ** 2
)
scores[alpha].append(mse)
mean_scores = {
alpha: np.mean(values)
for alpha, values in scores.items()
if len(values) > 0
}
best_alpha = min(
mean_scores,
key=mean_scores.get
)
return best_alpha, mean_scores
# ================================================================
# 8. FINAL FIT
# ================================================================
def fit_final_model(
X,
y,
alpha
):
model = Pipeline([
(
"scale",
StandardScaler()
),
(
"ridge",
Ridge(
alpha=alpha,
fit_intercept=True
)
)
])
model.fit(X, y)
return model
# ================================================================
# 9. PER-TRAJECTORY METRICS
# ================================================================
def trajectory_losses(
model,
X,
y,
trajectory_ids
):
pred = model.predict(X)
df = pd.DataFrame({
"trajectory_id": trajectory_ids,
"y": y,
"pred": pred
})
df["sq_error"] = (
df["y"] - df["pred"]
) ** 2
losses = (
df.groupby("trajectory_id")
["sq_error"]
.mean()
)
return losses.values, pred
# ================================================================
# 10. PAIRED BOOTSTRAP BY TRAJECTORY
# ================================================================
def bootstrap_difference(
loss_present,
loss_history,
n_boot,
seed
):
"""
Trajectory-level paired bootstrap.
The resampling unit is the trajectory,
NOT each individual time point.
This preserves the internal dependence
within each trajectory.
"""
rng = np.random.default_rng(seed)
diff = (
loss_present
- loss_history
)
n = len(diff)
boot = np.empty(n_boot)
for i in range(n_boot):
idx = rng.integers(
0,
n,
size=n
)
boot[i] = np.mean(
diff[idx]
)
lower = np.quantile(
boot,
0.025
)
upper = np.quantile(
boot,
0.975
)
return (
np.mean(diff),
lower,
upper,
boot
)
# ================================================================
# 11. PAIRED PERMUTATION TEST
# ================================================================
def paired_permutation_test(
loss_present,
loss_history,
n_perm,
seed
):
"""
Paired permutation test.
Under H0, within each trajectory,
the identity of which model obtained
which error can be permuted.
H1:
MSE_present > MSE_history
"""
rng = np.random.default_rng(seed)
diff = (
loss_present
- loss_history
)
observed = np.mean(diff)
n = len(diff)
extreme = 0
for _ in range(n_perm):
signs = rng.choice(
[-1, 1],
size=n
)
permuted = np.mean(
diff * signs
)
if permuted >= observed:
extreme += 1
p_value = (
extreme + 1
) / (
n_perm + 1
)
return observed, p_value
# ================================================================
# 12. HDD-I CLASSIFICATION
# ================================================================
def classify_construct_I(
mean_difference,
ci_lower,
ci_upper,
relative_improvement,
p_value,
practical_threshold
):
"""
Pre-specified criterion.
'+':
positive effect
+ CI does not cross zero
+ relative improvement >= threshold
+ p < 0.05
'-':
evidence against a relevant improvement:
upper CI <= 0
OR relative improvement below threshold
UE:
insufficient data / inconclusive result
"""
if (
ci_lower > 0
and relative_improvement >= practical_threshold
and p_value < CFG.alpha_stat
):
return "+"
if (
ci_upper <= 0
or relative_improvement < 0
):
return "-"
return "UE"
# ================================================================
# 13. PLACEBO HISTORY MODEL
# ================================================================
def make_placebo_history(
X_history,
rng
):
"""
Keeps the SAME number of historical features,
but breaks their temporal relationship with the target.
This tests whether an improvement simply appears
because the model has more parameters/features.
Present-state features remain intact.
"""
X_placebo = X_history.copy()
if X_history.shape[1] <= 2:
return X_placebo
historical = X_placebo[:, 2:].copy()
permutation = rng.permutation(
historical.shape[0]
)
historical = historical[
permutation
]
X_placebo[:, 2:] = historical
return X_placebo
# ================================================================
# 14. MAIN EXPERIMENT
# ================================================================
def run_single_system(
system_name,
noise_std,
seed
):
print("\n" + "=" * 70)
print(f"SYSTEM: {system_name}")
print(f"Noise: {noise_std}")
print("=" * 70)
train = generate_dataset(
system_name=system_name,
n_trajectories=CFG.n_train_trajectories,
length=CFG.trajectory_length,
burn_in=CFG.burn_in,
noise_std=noise_std,
seed=seed
)
test = generate_dataset(
system_name=system_name,
n_trajectories=CFG.n_test_trajectories,
length=CFG.trajectory_length,
burn_in=CFG.burn_in,
noise_std=noise_std,
seed=seed + 100000
)
results = []
detailed_predictions = {}
# ------------------------------------------------------------
# different history windows
# ------------------------------------------------------------
for history_window in CFG.history_windows:
print(
f"\n History = {history_window}"
)
(
Xp_train,
Xh_train,
y_train,
tid_train
) = make_supervised_dataset(
train,
history_window
)
(
Xp_test,
Xh_test,
y_test,
tid_test
) = make_supervised_dataset(
test,
history_window
)
# --------------------------------------------------------
# Model M0
# --------------------------------------------------------
alpha_present, _ = select_alpha(
Xp_train,
y_train,
tid_train,
CFG.alpha_grid,
CFG.inner_folds
)
model_present = fit_final_model(
Xp_train,
y_train,
alpha_present
)
# --------------------------------------------------------
# Model M1
# --------------------------------------------------------
alpha_history, _ = select_alpha(
Xh_train,
y_train,
tid_train,
CFG.alpha_grid,
CFG.inner_folds
)
model_history = fit_final_model(
Xh_train,
y_train,
alpha_history
)
# --------------------------------------------------------
# Predictions
# --------------------------------------------------------
loss_p, pred_p = trajectory_losses(
model_present,
Xp_test,
y_test,
tid_test
)
loss_h, pred_h = trajectory_losses(
model_history,
Xh_test,
y_test,
tid_test
)
# --------------------------------------------------------
# Bootstrap
# --------------------------------------------------------
(
delta,
ci_lower,
ci_upper,
boot
) = bootstrap_difference(
loss_p,
loss_h,
CFG.bootstrap_reps,
seed + history_window
)
# --------------------------------------------------------
# Permutation
# --------------------------------------------------------
(
observed_perm,
p_value
) = paired_permutation_test(
loss_p,
loss_h,
CFG.permutation_reps,
seed + 1000 + history_window
)
# --------------------------------------------------------
# Global MSE
# --------------------------------------------------------
mse_present = np.mean(
loss_p
)
mse_history = np.mean(
loss_h
)
relative_improvement = (
mse_present - mse_history
) / mse_present
# --------------------------------------------------------
# R2
# --------------------------------------------------------
ss_total = np.sum(
(
y_test
- np.mean(y_test)
) ** 2
)
ss_present = np.sum(
(
y_test
- pred_p
) ** 2
)
ss_history = np.sum(
(
y_test
- pred_h
) ** 2
)
r2_present = (
1
- ss_present / ss_total
)
r2_history = (
1
- ss_history / ss_total
)
classification = classify_construct_I(
delta,
ci_lower,
ci_upper,
relative_improvement,
p_value,
CFG.practical_relative_threshold
)
result = {
"System": system_name,
"Noise": noise_std,
"History_Window": history_window,
"Alpha_Present": alpha_present,
"Alpha_History": alpha_history,
"MSE_Present": mse_present,
"MSE_History": mse_history,
"Delta_MSE": delta,
"Relative_Improvement": relative_improvement,
"R2_Present": r2_present,
"R2_History": r2_history,
"CI95_Lower": ci_lower,
"CI95_Upper": ci_upper,
"Permutation_P": p_value,
"HDD_Construct_I": classification
}
results.append(result)
detailed_predictions[
history_window
] = {
"y": y_test,
"pred_present": pred_p,
"pred_history": pred_h,
"loss_present": loss_p,
"loss_history": loss_h
}
print(
f" MSE present: {mse_present:.8f}"
)
print(
f" MSE history: {mse_history:.8f}"
)
print(
f" \u0394MSE: {delta:.8f}"
)
print(
f" Relative improvement: "
f"{100 * relative_improvement:.3f}%"
)
print(
f" 95% CI \u0394MSE: "
f"[{ci_lower:.8f}, {ci_upper:.8f}]"
)
print(
f" Permutation p: {p_value:.6f}"
)
print(
f" HDD-I: {classification}"
)
return (
pd.DataFrame(results),
detailed_predictions
)
# ================================================================
# 15. MAIN EXECUTION
# ================================================================
all_results = []
all_predictions = {}
start_time = time.time()
for system_name in SYSTEM_GENERATORS:
result_df, predictions = run_single_system(
system_name=system_name,
noise_std=CFG.noise_std,
seed=CFG.random_seed
)
all_results.append(result_df)
all_predictions[
system_name
] = predictions
results_df = pd.concat(
all_results,
ignore_index=True
)
elapsed = time.time() - start_time
print("\n")
print("=" * 70)
print("MAIN RUN COMPLETE")
print("=" * 70)
print(
f"Time: {elapsed / 60:.2f} minutes"
)
# ================================================================
# 16. MAIN RESULT
# ================================================================
main_window = max(
CFG.history_windows
)
main_results = (
results_df[
results_df["History_Window"]
== main_window
]
.copy()
)
print("\n")
print("=" * 70)
print(
f"MAIN RESULT \u2014 HISTORY = {main_window}"
)
print("=" * 70)
display_columns = [
"System",
"MSE_Present",
"MSE_History",
"Delta_MSE",
"Relative_Improvement",
"CI95_Lower",
"CI95_Upper",
"Permutation_P",
"HDD_Construct_I"
]
display(
main_results[
display_columns
].round(6)
)
# ================================================================
# 17. HISTORY DEPENDENCE PROFILE \u0394L(\u03c4)
# ================================================================
print("\n")
print("=" * 70)
print("HISTORY DEPENDENCE PROFILE")
print("=" * 70)
plt.figure(
figsize=(11, 7)
)
for system_name in SYSTEM_GENERATORS:
subset = results_df[
results_df["System"]
== system_name
]
plt.plot(
subset["History_Window"],
100 * subset["Relative_Improvement"],
marker="o",
linewidth=2,
label=system_name
)
plt.axhline(
100 * CFG.practical_relative_threshold,
color="black",
linestyle="--",
alpha=0.7,
label="Practical threshold"
)
plt.axhline(
0,
color="gray",
linestyle=":",
alpha=0.7
)
plt.xlabel(
"Maximum history window \u03c4"
)
plt.ylabel(
"Relative improvement (%)"
)
plt.title(
"HDD Construct I \u2014 History dependence profile"
)
plt.legend()
plt.grid(alpha=0.25)
plt.tight_layout()
plt.savefig(
os.path.join(
CFG.output_dir,
"history_dependence_profile.png"
),
dpi=200
)
plt.show()
# ================================================================
# 18. CAPACITY CONTROL \u2014 PLACEBO HISTORY
# ================================================================
def run_capacity_control(
system_name,
noise_std,
history_window,
seed
):
print("\n")
print("=" * 70)
print(
f"CAPACITY CONTROL \u2014 {system_name}"
)
print("=" * 70)
train = generate_dataset(
system_name,
CFG.n_train_trajectories,
CFG.trajectory_length,
CFG.burn_in,
noise_std,
seed
)
test = generate_dataset(
system_name,
CFG.n_test_trajectories,
CFG.trajectory_length,
CFG.burn_in,
noise_std,
seed + 100000
)
(
Xp_train,
Xh_train,
y_train,
tid_train
) = make_supervised_dataset(
train,
history_window
)
(
Xp_test,
Xh_test,
y_test,
tid_test
) = make_supervised_dataset(
test,
history_window
)
# present model
alpha_p, _ = select_alpha(
Xp_train,
y_train,
tid_train,
CFG.alpha_grid,
CFG.inner_folds
)
model_p = fit_final_model(
Xp_train,
y_train,
alpha_p
)
# real history model
alpha_h, _ = select_alpha(
Xh_train,
y_train,
tid_train,
CFG.alpha_grid,
CFG.inner_folds
)
model_h = fit_final_model(
Xh_train,
y_train,
alpha_h
)
# placebo history
rng = np.random.default_rng(
seed + 9999
)
Xh_train_placebo = make_placebo_history(
Xh_train,
rng
)
Xh_test_placebo = make_placebo_history(
Xh_test,
rng
)
alpha_placebo, _ = select_alpha(
Xh_train_placebo,
y_train,
tid_train,
CFG.alpha_grid,
CFG.inner_folds
)
model_placebo = fit_final_model(
Xh_train_placebo,
y_train,
alpha_placebo
)
loss_p, _ = trajectory_losses(
model_p,
Xp_test,
y_test,
tid_test
)
loss_h, _ = trajectory_losses(
model_h,
Xh_test,
y_test,
tid_test
)
loss_placebo, _ = trajectory_losses(
model_placebo,
Xh_test_placebo,
y_test,
tid_test
)
real_improvement = (
np.mean(loss_p)
- np.mean(loss_h)
) / np.mean(loss_p)
placebo_improvement = (
np.mean(loss_p)
- np.mean(loss_placebo)
) / np.mean(loss_p)
return {
"System": system_name,
"Real_Improvement": real_improvement,
"Placebo_Improvement": placebo_improvement,
"Real_MSE": np.mean(loss_h),
"Placebo_MSE": np.mean(loss_placebo)
}
capacity_results = []
for system_name in SYSTEM_GENERATORS:
capacity_results.append(
run_capacity_control(
system_name,
CFG.noise_std,
main_window,
CFG.random_seed
)
)
capacity_df = pd.DataFrame(
capacity_results
)
print("\nCAPACITY CONTROL")
display(
capacity_df.round(6)
)
capacity_df.to_csv(
os.path.join(
CFG.output_dir,
"capacity_control.csv"
),
index=False
)
# ================================================================
# 19. STATE-RECONSTRUCTION DIAGNOSTIC
# ================================================================
def state_reconstruction_diagnostic(
system_name,
noise_std,
history_window,
seed
):
"""
State-reconstruction diagnostic.
This is NOT presented as proof of "irreducible history".
The idea is to check whether a compressed
representation of the history preserves its
predictive advantage.
We use PCA trained ONLY on the training data.
This tests a weaker hypothesis:
does the advantage of history disappear
when history is compressed into a low-dimensional
representation?
If yes:
this suggests that history may be providing
information about a latent/reconstructible state.
If no:
the dependence may be more complex.
This does NOT prove ontological irreducibility.
"""
train = generate_dataset(
system_name,
CFG.n_train_trajectories,
CFG.trajectory_length,
CFG.burn_in,
noise_std,
seed
)
test = generate_dataset(
system_name,
CFG.n_test_trajectories,
CFG.trajectory_length,
CFG.burn_in,
noise_std,
seed + 100000
)
(
Xp_train,
Xh_train,
y_train,
tid_train
) = make_supervised_dataset(
train,
history_window
)
(
Xp_test,
Xh_test,
y_test,
tid_test
) = make_supervised_dataset(
test,
history_window
)
# history block only
H_train = Xh_train[:, 2:]
H_test = Xh_test[:, 2:]
# cap dimensionality in a pre-defined way
n_components = min(
3,
H_train.shape[1]
)
scaler = StandardScaler()
H_train_scaled = scaler.fit_transform(
H_train
)
H_test_scaled = scaler.transform(
H_test
)
pca = PCA(
n_components=n_components
)
S_train = pca.fit_transform(
H_train_scaled
)
S_test = pca.transform(
H_test_scaled
)
# present state + reconstructed state
Xs_train = np.column_stack([
Xp_train,
S_train
])
Xs_test = np.column_stack([
Xp_test,
S_test
])
alpha_s, _ = select_alpha(
Xs_train,
y_train,
tid_train,
CFG.alpha_grid,
CFG.inner_folds
)
model_s = fit_final_model(
Xs_train,
y_train,
alpha_s
)
loss_s, _ = trajectory_losses(
model_s,
Xs_test,
y_test,
tid_test
)
mse_s = np.mean(loss_s)
return {
"System": system_name,
"MSE_Reconstructed_State": mse_s,
"Explained_Variance_PCA": np.sum(
pca.explained_variance_ratio_
)
}
reconstruction_results = []
for system_name in SYSTEM_GENERATORS:
reconstruction_results.append(
state_reconstruction_diagnostic(
system_name,
CFG.noise_std,
main_window,
CFG.random_seed
)
)
reconstruction_df = pd.DataFrame(
reconstruction_results
)
print("\n")
print("=" * 70)
print("STATE-RECONSTRUCTION DIAGNOSTIC")
print("=" * 70)
display(
reconstruction_df.round(6)
)
reconstruction_df.to_csv(
os.path.join(
CFG.output_dir,
"state_reconstruction.csv"
),
index=False
)
# ================================================================
# 20. NOISE ROBUSTNESS
# ================================================================
noise_results = []
print("\n")
print("=" * 70)
print("NOISE ROBUSTNESS ANALYSIS")
print("=" * 70)
for noise in CFG.noise_levels:
print(
f"\nNoise = {noise}"
)
for system_name in SYSTEM_GENERATORS:
df, _ = run_single_system(
system_name,
noise,
CFG.random_seed
)
selected = df[
df["History_Window"]
== main_window
].copy()
noise_results.append(
selected
)
noise_df = pd.concat(
noise_results,
ignore_index=True
)
noise_df.to_csv(
os.path.join(
CFG.output_dir,
"noise_robustness.csv"
),
index=False
)
# ================================================================
# 21. ROBUSTNESS TO MULTIPLE SEEDS
# ================================================================
seed_results = []
print("\n")
print("=" * 70)
print("ROBUSTNESS TO MULTIPLE SEEDS")
print("=" * 70)
for seed in CFG.robustness_seeds:
print(
f"\nSeed = {seed}"
)
for system_name in SYSTEM_GENERATORS:
df, _ = run_single_system(
system_name,
CFG.noise_std,
seed
)
selected = df[
df["History_Window"]
== main_window
].copy()
selected["Seed"] = seed
seed_results.append(
selected
)
seed_df = pd.concat(
seed_results,
ignore_index=True
)
seed_df.to_csv(
os.path.join(
CFG.output_dir,
"multi_seed_robustness.csv"
),
index=False
)
# ================================================================
# 22. CONSISTENCY ACROSS SEEDS
# ================================================================
seed_summary = []
for system_name in SYSTEM_GENERATORS:
subset = seed_df[
seed_df["System"]
== system_name
]
positive_fraction = np.mean(
subset["HDD_Construct_I"]
== "+"
)
mean_improvement = np.mean(
subset["Relative_Improvement"]
)
std_improvement = np.std(
subset["Relative_Improvement"],
ddof=1
)
seed_summary.append({
"System": system_name,
"Positive_Fraction": positive_fraction,
"Mean_Relative_Improvement":
mean_improvement,
"SD_Relative_Improvement":
std_improvement
})
seed_summary_df = pd.DataFrame(
seed_summary
)
print("\n")
print("=" * 70)
print("ROBUSTNESS SUMMARY")
print("=" * 70)
display(
seed_summary_df.round(6)
)
# ================================================================
# 23. BEHAVIORAL EQUIVALENCE BENCHMARK
# ================================================================
print("\n")
print("=" * 70)
print("BEHAVIORAL EQUIVALENCE BENCHMARK")
print("=" * 70)
print(
"""
This test matters for interpreting HDD.
Delay Line and Recurrent may have different
internal mechanisms, but Construct I should not
attempt to infer architecture from a simple
predictive improvement.
Therefore:
Positive Construct I
does NOT imply
positive Construct III.
"""
)
# ================================================================
# 24. GROUND TRUTH
# ================================================================
ground_truth = pd.DataFrame({
"System": [
"Markov",
"Hidden State",
"Delay Line",
"Recurrent"
],
"Expected_Construct_I": [
"-",
"+",
"+",
"+"
],
"Reason": [
"Observed state is sufficient",
"History reveals latent state",
"Dynamics explicitly depend on delay",
"History helps infer hidden recurrent state"
]
})
final_main = main_results.merge(
ground_truth,
on="System",
how="left"
)
final_main[
"Correct_vs_Ground_Truth"
] = (
final_main[
"HDD_Construct_I"
]
==
final_main[
"Expected_Construct_I"
]
)
print("\n")
print("=" * 70)
print("COMPARISON WITH GROUND TRUTH")
print("=" * 70)
display(
final_main[
[
"System",
"HDD_Construct_I",
"Expected_Construct_I",
"Correct_vs_Ground_Truth"
]
]
)
# ================================================================
# 25. PERFORMANCE MATRIX
# ================================================================
accuracy = np.mean(
final_main[
"Correct_vs_Ground_Truth"
]
)
print(
f"\nAccuracy against ground truth: "
f"{100 * accuracy:.1f}%"
)
# ================================================================
# 26. ROC-LIKE CURVE / EFFECT SIZE
# ================================================================
plt.figure(
figsize=(10, 6)
)
colors = {
"Markov": "black",
"Hidden State": "blue",
"Delay Line": "orange",
"Recurrent": "red"
}
for system_name in SYSTEM_GENERATORS:
subset = results_df[
results_df["System"]
== system_name
]
plt.errorbar(
subset["History_Window"],
100 * subset["Relative_Improvement"],
fmt="o-",
capsize=4,
linewidth=2,
label=system_name,
color=colors[system_name]
)
plt.axhline(
0,
color="gray",
linestyle=":"
)
plt.axhline(
100 * CFG.practical_relative_threshold,
color="black",
linestyle="--",
label="Practical threshold"
)
plt.xlabel(
"Maximum history \u03c4"
)
plt.ylabel(
"Relative improvement (%)"
)
plt.title(
"Effect of history on prediction"
)
plt.grid(
alpha=0.25
)
plt.legend()
plt.tight_layout()
plt.savefig(
os.path.join(
CFG.output_dir,
"effect_size_profile.png"
),
dpi=200
)
plt.show()
# ================================================================
# 27. FULL FINAL TABLE
# ================================================================
final_results = results_df.merge(
ground_truth,
on="System",
how="left"
)
final_results[
"Correct_vs_Ground_Truth"
] = np.where(
final_results["HDD_Construct_I"]
== final_results["Expected_Construct_I"],
True,
False
)
final_results.to_csv(
os.path.join(
CFG.output_dir,
"HDD_Construct_I_full_results.csv"
),
index=False
)
main_results.to_csv(
os.path.join(
CFG.output_dir,
"HDD_Construct_I_main_results.csv"
),
index=False
)
# ================================================================
# 28. REPRODUCIBILITY MANIFEST
# ================================================================
manifest = {
"benchmark_version": "HDD Benchmark v1.0",
"construct": "I",
"definition":
"History improves out-of-sample prediction "
"beyond measured present state and current input",
"configuration": asdict(CFG),
"ground_truth": ground_truth.to_dict(
orient="records"
),
"environment": ENVIRONMENT,
"files": [
"HDD_Construct_I_full_results.csv",
"HDD_Construct_I_main_results.csv",
"capacity_control.csv",
"state_reconstruction.csv",
"noise_robustness.csv",
"multi_seed_robustness.csv",
"history_dependence_profile.png",
"effect_size_profile.png",
"environment.json"
],
"important_interpretation":
"Positive Construct I does not establish memory "
"mechanism, recurrence, self-reference, self-modeling "
"or consciousness."
}
with open(
os.path.join(
CFG.output_dir,
"manifest.json"
),
"w"
) as f:
json.dump(
manifest,
f,
indent=2,
default=str
)
# ================================================================
# 29. AUTOMATIC REPORT
# ================================================================
report_lines = []
report_lines.append(
"HDD BENCHMARK v1.0"
)
report_lines.append(
"Construct I \u2014 History-Dependent Predictive Structure"
)
report_lines.append(
""
)
report_lines.append(
"MAIN RESULT"
)
for _, row in main_results.iterrows():
report_lines.append(
f"{row['System']}: "
f"HDD-I={row['HDD_Construct_I']}, "
f"improvement="
f"{100 * row['Relative_Improvement']:.3f}%, "
f"CI95=["
f"{row['CI95_Lower']:.8f}, "
f"{row['CI95_Upper']:.8f}], "
f"p="
f"{row['Permutation_P']:.6f}"
)
report_lines.append(
""
)
report_lines.append(
"GROUND TRUTH"
)
for _, row in ground_truth.iterrows():
report_lines.append(
f"{row['System']}: "
f"expected={row['Expected_Construct_I']}"
)
report_lines.append(
""
)
report_lines.append(
f"Preliminary benchmark accuracy: "
f"{100 * accuracy:.1f}%"
)
report_lines.append(
""
)
report_lines.append(
"INTERPRETATION:"
)
report_lines.append(
"A positive Construct I classification means only "
"that historical information improved out-of-sample "
"prediction relative to the specified observed present "
"state and current input."
)
report_lines.append(
"It does not establish a memory mechanism, recurrence, "
"self-reference, self-modeling, agency or consciousness."
)
report_lines.append(
""
)
report_lines.append(
"LIMITATION:"
)
report_lines.append(
"This benchmark validates the computational "
"discrimination of Construct I under the specified "
"observation and model class. It does not validate "
"the complete HDD framework."
)
report_text = "\n".join(
report_lines
)
print("\n")
print("=" * 70)
print("REPORT")
print("=" * 70)
print(report_text)
with open(
os.path.join(
CFG.output_dir,
"REPORT.txt"
),
"w"
) as f:
f.write(report_text)
# ================================================================
# 30. FINAL ZIP
# ================================================================
import shutil
zip_path = shutil.make_archive(
CFG.output_dir,
"zip",
CFG.output_dir
)
print("\n")
print("=" * 70)
print("BENCHMARK FINISHED")
print("=" * 70)
print(
f"\nFiles saved to:\n"
f"{os.path.abspath(CFG.output_dir)}"
)
print(
f"\nZIP package:\n"
f"{zip_path}"
)
print("\n")
print("=" * 70)
print("SCIENTIFIC INTERPRETATION")
print("=" * 70)
print(
"""
The benchmark was designed to answer only:
Does history improve prediction of the future
beyond the observed present state?
It does NOT answer:
Is history memory?
Does recurrence exist?
Does self-reference exist?
Does a self-model exist?
Does consciousness exist?
Those questions require other experiments.
The correct result must be interpreted
relative to the observation model, intervention
protocol, and model class used.
The Hidden State system is especially important:
a positive result on it does NOT mean the system
has an explicit memory. History may simply
reveal information about a latent variable.
Therefore, the benchmark tests precisely one of
HDD's central methodological theses:
"predictive history is not automatically
evidence of a specific mechanism."
"""
)
print("\n")
print("=" * 70)
print("END")
print("=" * 70)
Comentários
Postar um comentário