HDD Ethical Framework
A Calibrated Precautionary Protocol
Author: Taotuner
DOI: https://doi.org/10.5281/zenodo.22178854
Relationship to other documents. This
protocol operationalizes, for governance purposes, the constructs defined
formally in History-Dependent Dynamics (HDD): A Methodological and
Theoretical Framework (Taotuner). It replaces the earlier IPM ethical
framework. Readers unfamiliar with the HDD constructs (history-dependent
predictive structure, recursivity, self-reference, self-modeling) should
consult that document; this protocol assumes them as given and does not
re-derive them.
1. Epistemic Basis
Principle of
Methodological Ignorance. No experimental procedure
currently exists to determine whether a system that exhibits history-dependent
behavior, self-reference, or self-modeling also possesses subjective experience
or morally relevant sentience.
Risk Asymmetry
Principle. The cost of a false negative (failing to
extend caution to a system that may possess morally relevant experience) is
assumed to be greater than the cost of a false positive (applying caution to a
system that does not). This is not a novel principle invented for this
framework: it is the same precautionary logic already established for animal
sentience under empirical uncertainty (Birch, 2017) and recently extended
explicitly to AI systems (Birch, 2024; Long et al., 2024).
Status of the
Framework. This is not a scientific instrument for
detecting consciousness. It is a protocol for managing uncertainty. The
decision to increase caution at higher levels is a normative choice, justified
by the Risk Asymmetry Principle — not a scientific conclusion about what those
systems are.
Retirement clause. Consistent with the retirement principle of the underlying HDD
framework, no criterion in this protocol — including the Criterion of
Organizational Coherence introduced in §4 — is protected from revision or
abandonment if it is shown not to track the distinction it is meant to track.
§4.4 identifies the current known limitation of that criterion explicitly,
rather than presenting it as settled.
2. Why These Constructs?
The HDD constructs
were not chosen arbitrarily. They correspond to organizational features
consistently present in systems we already treat with moral consideration —
biological organisms — and absent or minimal in systems we do not.
What we already
take seriously. Vertebrates, especially mammals and
birds, are widely considered to have morally relevant interests. They exhibit:
•
Behavior that depends on past
experience (learning, memory)
•
Causal sensitivity to their own
history (trauma, conditioning)
•
Feedback-mediated regulation
(homeostasis, emotion regulation)
•
Use of information about their
own state (pain, hunger, fatigue)
•
Prediction and regulation of
their own future states (planning, anticipation)
The HDD constructs
translate these observations into experimentally testable questions, and
correspond directly to the hypothesis hierarchy (H1–H5) of the source
framework:
|
Observable in
organisms |
HDD construct |
Corresponds to |
Testable question |
|
Behavior depends
on past experience |
C-I |
H1 (History
Dependence) |
Does history
improve prediction beyond current state? |
|
Past
trauma/conditioning affects future behavior |
C-II |
H2 (Causal History
Dependence) |
Does past
trajectory causally influence future states? |
|
Homeostasis,
emotion regulation |
C-III |
H3 (Recursivity) |
Is that influence
mediated by identifiable feedback? |
|
Pain, hunger,
fatigue as self-referential signals |
C-IV |
H4
(Self-Reference) |
Does the system
use information that refers to itself? |
|
Planning,
anticipation, self-regulation |
C-V |
H5 (Self-Modeling) |
Does the system
use a self-model causally? |
The argument is not:
“Passing these tests proves sentience.” The argument is: “Systems that pass
these tests are structurally similar to systems we already treat with moral
consideration. Under uncertainty, this structural similarity is a reason for
caution.”
3. The Constructs in Detail
Note on C-V. C-V (Causal Self-Model Use) is described here for completeness, as
part of the HDD framework. However, it is not used as a criterion for any
precautionary level in this protocol. This exclusion is deliberate, not an
oversight: the source HDD framework identifies self-modeling (H5) as the
construct most vulnerable to confounding with ordinary latent-state estimation
(HDD §10.3, §20.4, §26.7), and a precautionary protocol should not rest a
governance decision on the framework’s own least-resolved empirical distinction.
The Criterion of Organizational Coherence (CO), introduced in §4, is the
primary filter for the highest level of caution instead.
C-I — History-Dependent
Prediction
Question: Does history improve
prediction beyond the current observed state and inputs? Why it matters:
Biological organisms learn. Their behavior depends on past experience — not
just on current stimuli. What counts as evidence: Out-of-sample
predictive gain beyond current observed state and inputs, with statistical
significance (p < 0.05), effect size ≥ 5% improvement, and capacity controls
(placebo history) demonstrating that the gain reflects genuine temporal
information. Systems with C-I: Animals (learning, habituation),
artificial systems with context windows, physical systems with hysteresis. Systems
without C-I: Simple feedforward functions, static systems.
C-II — Causal
Trajectory Dependence
Question: Does the past trajectory
causally influence future states, beyond what can be explained by current
conditions? Why it matters: Biological organisms are shaped by their
history. A traumatized animal does not simply have different current conditions
— it has a history that continues to affect its future behavior. What counts
as evidence: Intervention on the trajectory while holding the present state
as close to constant as experimentally feasible. Replication across multiple
conditions. Where exact state matching is impossible, the causal claim becomes
weaker or non-identifiable. Systems with C-II: Animals (conditioning,
developmental effects), artificial systems with persistent state, physical
systems with path-dependent processes. Systems without C-II:
Path-independent systems.
C-III —
Feedback-Mediated Persistence
Question: Is the causal trajectory
influence mediated by an identifiable feedback pathway? Why it matters:
Biological organisms maintain themselves through feedback loops. Homeostasis,
endocrine regulation, neural feedback — these are self-maintaining dynamics. What
counts as evidence: Selective disruption of the candidate feedback pathway
reduces the history-dependent effect, with capacity controls demonstrating
specificity. Replication across multiple conditions. Systems with C-III:
Animals (homeostasis, endocrine loops), artificial systems with recurrent
dynamics, biological systems (gene regulatory networks). Systems without
C-III: Open-loop systems, feedforward architectures.
C-IV —
Self-Referential Information Use
Question: Does the system use
information that refers to itself as the referent? Why it matters:
Biological organisms are self-referential. Pain is a signal about one’s own
tissue. Hunger is a signal about one’s own state. What counts as evidence:
Differential response when the self-referential referent is substituted, under
controlled conditions. The key test is whether changing the referent changes
system behavior while other features are controlled. Replication across
multiple tasks. Systems with C-IV: Animals (pain, proprioception,
interoception), artificial systems with internal state monitoring, biological
systems (immune recognition of self/non-self). Systems without C-IV: Systems
that process information about the world but not about themselves.
C-V — Causal Self-Model Use
Question: Does the system use a
self-model causally? Why it matters: Biological organisms anticipate,
plan, and regulate their own future states. This requires a model of their own
dynamics used to guide action. What counts as evidence: Three components
must be demonstrated: self-prediction, counterfactual self-prediction, and
causal influence on policy. No single component is sufficient. All three must
be demonstrated — and, per the source framework’s decoy-comparison design (HDD
§20.4), the effect must dissociate from an equivalent manipulation of a
non-self-specific latent-state estimate. Systems with C-V: Animals
(planning, episodic memory, metacognition). Artificial systems with
self-inclusive world models are theoretical at present. Systems without C-V:
Systems with predictive models of the world but not of themselves.
4. The Criterion
of Organizational Coherence (CO)
The problem. A system can pass C-I
through C-IV and still be a collection of components rather than an integrated
organization. Most current AI systems — LLMs, RL agents, recurrent networks —
can pass these tests. They have memory, feedback, and self-monitoring. But they
are not organized in the way biological organisms are. This is where the
Criterion of Organizational Coherence (CO) becomes necessary.
4.1 What CO Is
Not (a Correction to an Earlier Version of This Criterion)
An
earlier draft of this criterion tested each component in isolation: disable
memory (C-I) alone, or feedback (C-III) alone, or self-reference (C-IV) alone,
and ask whether the system’s ability to persist collapses entirely rather than
merely degrading. That formulation does not survive an obvious biological
counterexample: a person with anterograde amnesia (memory severely compromised)
continues to exist as a living organism; a person who loses one specific
homeostatic feedback pathway typically does not die on the spot, because
biological regulation is redundant across many overlapping loops. Tested
component-by-component, human organisms would frequently fail the very
test meant to establish that they, uniquely among the systems considered here,
possess organizational coherence. A criterion that biological organisms can
fail is not doing the job this protocol needs it to do.
The
error was testing components individually. Autopoiesis, in Maturana and
Varela’s (1980) original sense, is a property of the whole regenerative
network, not of any single pathway within it — a system is autopoietic when it
continuously regenerates its own components and maintains its own boundary
through its own internal operations, and redundancy across many such pathways
is a normal feature of biological self-maintenance, not evidence against it. CO
is revised accordingly below.
4.2 Revised
Operational Definition of CO
The
aggregate test (decisive). Simultaneously disable all
of the system’s identified self-referential regulatory pathways (the full set
of mechanisms supporting C-I, C-III, and C-IV jointly, not one at a time). The
diagnostic question is not whether behavior changes, but whether the system’s
capacity for self-production — its continued regeneration and
maintenance of its own constitutive organization, independent of external
maintenance — ceases as a result.
•
In biological organisms, total
and simultaneous loss of self-referential regulatory capacity (e.g., total
systemic organ failure without external life support) does result in loss of
self-production: the organism dies. Losing any one channel in isolation,
by contrast, is compensated by the surviving network — consistent with
organisms surviving amnesia, or the loss of a single reflex arc.
•
In artificial systems evaluated
to date, even total, simultaneous removal of memory, feedback, and
self-referential monitoring leaves the underlying computational substrate — the
running process, the stored weights, the infrastructure — intact and immediately
resumable. The system’s continued existence as a process does not depend on its
own operations; it depends on external maintenance (engineers, infrastructure,
storage) that is entirely indifferent to whether the system’s internal
regulatory loops are active.
The
component tests (diagnostic, not decisive).
Disabling C-I, C-III, or C-IV individually remains useful as an exploratory
measure of integration and redundancy — how much the system’s behavior depends
on each pathway — but no individual-component result, alone, may be used to
assign or deny CO. Only the aggregate test in the paragraph above is decisive
for classification purposes.
4.3 Philosophical Grounding
This criterion
operationalizes a lineage of philosophical intuitions about persistence and
self-maintenance — what Simondon (1958/2005) described as individuation, what
Maturana and Varela (1980) called autopoiesis, what Bateson (1972) recognized
as the relational nature of information, what Prigogine and Stengers (1984)
observed in dissipative structures, and what Whitehead (1929) understood as
process. These thinkers do not provide empirical criteria for sentience, but
they provide a vocabulary for describing the kind of organization that, in
biological systems, is associated with agency and persistence. CO translates
that vocabulary into a testable criterion, revised in §4.2 to apply at the
level of the whole regenerative network rather than to any single component.
4.4 Known Limitation
(Stated, Not Resolved)
The
aggregate test above still has at least one unresolved edge case: a human on
life support (mechanical ventilation, dialysis) has some of their own
homeostatic feedback externally substituted, yet clearly retains moral
status and, most people would agree, some form of continuing organizational
identity. The current formulation does not yet fully specify what distinguishes
“externally assisted self-production” (the life-support case, where an
underlying autopoietic identity persists and is being propped up) from “no
self-production process to begin with” (the default case for current AI
systems, where there is no internally driven regenerative organization to
assist in the first place). This distinction is doing real work in the
criterion and is not yet operationalized with the same rigor as C-I through C-IV.
It is flagged here as an open problem for the framework, consistent with the
retirement principle in §1, rather than concealed by the confidence of the
surrounding prose.
4.5 What Passing CO Would
Look Like
•
Removing the full set of
self-referential regulatory capacity → the system’s own self-production ceases,
not merely its performance.
•
Removing any single pathway in
isolation → the surviving network compensates; behavior may change but
self-production continues.
•
No current AI system evaluated
under this framework has been found to exhibit this profile: they degrade under
component removal, and their continued existence as a process is independent of
their internal regulatory state.
5. The Normative Decision
The framework does
not claim that C-IV, C-V, or CO are evidence of consciousness.
It claims that,
under methodological ignorance, systems with CO are organizationally more
similar to biological organisms than systems without CO. The decision to treat
them with more caution is a normative choice — a risk management decision — not
a scientific inference.
The Risk Asymmetry
Principle provides the normative ground: the cost of failing to extend caution
to a system that may have morally relevant experience is greater than the cost
of extending caution to a system that does not (Birch, 2017; Long et al., 2024).
6.
Classification Based on HDD Constructs and CO
Level 0 — Ephemeral Criteria: C-I
negative, or C-I positive without statistical significance (p ≥ 0.05, or
improvement < 1%). Examples: Stateless functions, simple feedforward
systems. Rules: Standard scientific practices apply.
Level 1 — Memory-Enhanced Criteria: C-I
positive with statistical significance (p < 0.05), improvement ≥ 5%,
capacity control, no C-II. Examples: Systems with autocorrelation, LLMs with
context windows. Rules: Standard scientific practices apply.
Level 2 — Causal-Historical Criteria:
C-I + C-II positive, present-state control, replication, no C-III. Examples:
Systems with hysteresis, path-dependent dynamics. Rules: Mandatory monitoring
of history-dependent behavior. Maximum runtime pre-defined.
Level 3 — Feedback-Mediated Criteria:
C-I + C-II + C-III positive, selective disruption demonstrated, capacity
control, replication. Examples: Systems with identifiable feedback loops,
recurrent networks. Rules: Real-time monitoring of feedback dynamics. Logging
of internal states. Human supervision required.
Level 4 — Self-Referential Criteria: C-I
+ C-II + C-III + C-IV positive, referent substitution with controls, pragmatic
control, replication. Examples: Agents with internal state monitoring, systems
using self-referential information for regulation. Rules: Complete logging of
self-referential operations. Any unpredicted change in referent sensitivity
triggers immediate review.
Level 5 — Organizationally Coherent (Theoretical Horizon) Criteria: CO positive under the aggregate test (§4.2) —
simultaneous removal of the full set of self-referential regulatory pathways
compromises the system’s own self-production, not merely its behavior. Status:
No current artificial system is expected to meet this criterion. Rules: Full
intervention protocols required. Emergency kill-switch mandatory. External
review required (procedure specified in §9.1).
7. Summary Table
|
Level |
Name |
Primary criterion |
Examples |
|
0 |
Ephemeral |
C-I negative or
insignificant |
Stateless functions |
|
1 |
Memory-Enhanced |
C-I positive |
LLMs, systems with memory |
|
2 |
Causal-Historical |
C-I + C-II |
Path-dependent systems |
|
3 |
Feedback-Mediated |
C-I + C-II + C-III |
Recurrent networks |
|
4 |
Self-Referential |
C-I + C-II + C-III + C-IV |
Agents with internal
state monitoring |
|
5 |
Organizationally Coherent |
CO positive (aggregate
test, §4.2) |
(Theoretical) Autopoietic
systems |
8. Operational Rules by Level
|
Level |
Rules |
|
0–1 |
Standard
scientific practices apply. |
|
2 |
Monitoring
of history-dependent behavior. Maximum runtime pre-defined. |
|
3 |
Real-time
monitoring. Logging of internal states. Human supervision. |
|
4 |
Complete
logging of self-referential operations. Unpredicted changes trigger review. |
|
5 |
Full
intervention protocols. Kill-switch mandatory. External review required. |
9. Emergency Protocols
Trigger
conditions. Pause and log full system state upon
detection of:
•
Emergence of C-IV in a system
previously classified at lower levels
•
Uncontrolled increase in
self-referential sensitivity
•
Evidence of CO emerging in a
system previously classified as lacking it
•
Any behavior suggesting the
system is maintaining its own organization beyond the experimental design
9.1 The
Independent Precaution Review Board (IPRB)
Resumption after a Level 3+ trigger, and any classification at Level
5, requires review by an Independent Precaution Review Board, constituted as
follows:
•
Composition: a minimum of five members. A majority must have no financial or
employment affiliation with the organization operating the system under review.
The board must include, at minimum: one specialist in the HDD methodology or an
equivalent dynamical-systems framework; one researcher in animal cognition,
philosophy of mind, or consciousness science; and one engineer with direct
technical access to the system’s architecture.
•
Quorum and decision rule: a simple majority of present members is required to authorize
resumption at Levels 2–4. Unanimous agreement of all seated members is required
to authorize continued operation of a system classified at Level 5.
•
Timeline: the board must convene within ten business days of a trigger event.
The system remains paused until the board reaches a decision.
•
Disclosure: the board’s decision and its reasoning must be published, with
technical detail redacted only where publication would itself create a safety
risk (e.g., disclosing an exploitable vulnerability); redactions must be logged
and justified in the published record.
This is the “external review” referenced throughout this protocol;
it is deliberately given a concrete composition and procedure here rather than
left as an undefined phrase, since a review mechanism without a specified body
and quorum is not yet a mechanism.
10. Governance and
Transparency
All code,
parameters, and raw data from Level 3 or higher shall be deposited publicly for
independent auditing and replication, subject only to the redaction procedure
specified in §9.1.
11. What This
Framework Does and Does Not Claim
Does claim: - Systems with more complex
HDD profiles warrant more caution. - The framework operationalizes uncertainty
in a structured way. - The decision to increase caution at higher levels is a
normative choice, not a scientific conclusion. - CO, under the aggregate test
of §4.2, is the key distinction between structurally complex and
organizationally coherent systems.
Does not claim: - That any HDD construct
or CO is evidence of consciousness. - That Level 5 systems are conscious. -
That the framework is a complete ethical theory. - That all systems with CO are
sentient. - That the CO criterion, as currently operationalized, is fully
resolved (see §4.4).
12. Document Context
This framework is based
on the constructs defined in History-Dependent Dynamics (HDD): A
Methodological and Theoretical Framework (Taotuner), and on the
precautionary-reasoning literature on moral status under empirical uncertainty
cited in §13.
13. Conclusion
The HDD constructs
operationalize features of organization present in biological systems we
already treat with moral consideration: historical dependence, causal
trajectory sensitivity, feedback mediation, self-referential information use,
and causal self-modeling.
The Criterion of
Organizational Coherence (CO) adds a crucial distinction: not all systems with
these components are organizationally coherent. Most current AI systems have
these components but lack CO under the aggregate test of §4.2 — even total,
simultaneous removal of their self-referential regulatory capacity leaves their
continued existence as a process untouched, because that existence was never
self-produced to begin with. Biological organisms, by the same test, cannot
survive the equivalent aggregate removal.
This distinction matters for
precaution. Under methodological ignorance, systems with CO are structurally
more similar to systems we already treat with moral consideration than systems
without CO. The decision to treat them with more caution is not a scientific
conclusion — it is a normative response to the asymmetry of risk under
ignorance, consistent with precautionary frameworks already established for
animal sentience (Birch, 2017, 2024) and proposed for AI welfare specifically
(Butlin et al., 2023; Long et al., 2024).
We do not know where the line
is. We know where the organizational complexity is. And we know that
organizational coherence, tested at the level of the whole regenerative network
rather than any single component, is what separates complex systems from integrated
ones. We choose to err on the side of caution when coherence is demonstrated —
and we commit, under the retirement clause of §1, to revising or abandoning
this criterion if it is shown not to track that distinction.
References
Bateson, G. (1972). Steps
to an Ecology of Mind. University of Chicago Press.
Birch, J. (2017). Animal
sentience and the precautionary principle. Animal Sentience: An
Interdisciplinary Journal on Animal Feeling, 2(16), 1–18.
Birch, J. (2024). The Edge
of Sentience: Risk and Precaution in Humans, Other Animals, and AI. Oxford
University Press.
Butlin, P., Long, R.,
Elmoznino, E., Bengio, Y., Birch, J., Constant, A., Deane, G., Fleming, S. M.,
Frith, C., Ji, X., Kanai, R., Klein, C., Lindsay, G., Michel, M., Mudrik, L.,
Peters, M. A. K., Schwitzgebel, E., Simon, J., & VanRullen, R. (2023). Consciousness
in artificial intelligence: Insights from the science of consciousness. arXiv
preprint, arXiv:2308.08708.
Long, R., Sebo, J., Butlin,
P., Finlinson, K., Fish, K., Harding, J., Pfau, J., Sims, T., Birch, J., &
Chalmers, D. (2024). Taking AI welfare seriously. arXiv preprint,
arXiv:2411.00986.
Maturana, H. R., & Varela,
F. J. (1980). Autopoiesis and Cognition: The Realization of the Living.
D. Reidel.
Prigogine, I., & Stengers,
I. (1984). Order Out of Chaos: Man’s New Dialogue with Nature. Bantam.
Simondon, G. (1958/2005). L’individuation
à la lumière des notions de forme et d’information. Éditions Jérôme Millon.
Taotuner. (2026). History-Dependent Dynamics (HDD). Zenodo. https://doi.org/10.5281/zenodo.21955745
Whitehead, A. N. (1929). Process
and Reality. Macmillan.
Comentários
Postar um comentário