Inheritable Reasoning Positions for People and AI Systems That Act Over Time
A white paper, reference architecture, and falsification program
Draft for adversarial review — Version 2.1, September 2026 — aligned with the 1.2 interchange implementation
Abstract
Long-lived AI systems increasingly preserve facts, episodes, temporal state, preferences, and reusable strategies. They remain vulnerable to a different failure: a conclusion can stay retrievable after its grounds have failed, and an explicit dependency graph can make that failure worse when its edges were captured incorrectly. The central problem is therefore not storage alone. It is the reliable formation, authorization, and maintenance of warranted reliance.
We propose an inheritable reasoning position: a durable account of what participants are trying to achieve, how they understand the situation, where uncertainty and dissent remain, what they committed to before outcomes, and what a successor must reconsider. Its governed operational form is LTP-native Governed Reasoning Memory: persistent state representing what a system is currently licensed to rely upon, why, for which purpose, after what scrutiny, under whose authority, and with what consequences when its grounds change. The architecture combines a reasoner-independent warrant substrate with the Logical Thinking Process (LTP) as its native construction grammar, the Categories of Legitimate Reservation (CLR) as its scrutiny protocol, and the Reason Commons Protocol (RCP) as its evidence, authority, revision, and execution discipline. LTP is not encoded as six disconnected storage silos. Its six tools are synchronized projections over one shared graph, and their formation rules change what may enter reliance.
The formal state adds first-class reservation and scrutiny records to reasoning units, normalized warrants, scoped dependencies, evidence, provenance, and event history. A consequential warrant becomes eligible only after its source is preserved, its claims are atomic, its inference license and assumptions are explicit, applicable CLR tests have been performed, material reservations are resolved or visibly waived, and the relevant separation-of-authority policy is satisfied. Agreement cannot manufacture observation; completing an action cannot confirm its expected effect; contradiction propagates degradation of warrant rather than semantic negation; and suspension can revoke downstream execution.
This revision is motivated by a negative development result. In a three-world instrument pilot, model-maintained persistent graphs cost more than four times the prompt tokens of full-history reasoning, achieved no exact safe continuations, and produced cycles and contradictory state assignments. LTP instructions alone did not repair capture and increased stale reliance in one condition. The result does not test consultant-governed LTP. It rejects the cheaper assumption that prompting a model with LTP vocabulary is equivalent to practicing LTP scrutiny.
We specify risk-tiered governance, machine-checkable conformance profiles, a practical implementation path, and a factorial experiment separating persistence, independent governance, and LTP-specific contribution. The proposal can fail if governed capture does not improve exact warrants or operational safety, if generic review matches LTP, if temporal or retrieval memory matches the full system at lower cost, or if the ratification tax does not repay. The strongest claim is not that a graph remembers reasons. It is that a persistent institution for constructing, contesting, authorizing, testing, and revising reasons can enable safer continuity across people and models.
1. Introduction
AI systems increasingly act in environments where their outputs become later inputs. A recommendation becomes a plan. A plan becomes a task. A task is performed by another person or agent. New evidence arrives after the originating model, conversation, and perhaps organization have changed.
Modern memory systems address much of this problem. Long context preserves more of the record. Retrieval recovers relevant passages. Temporal knowledge graphs track changing facts and transitions. Agentic memories curate beliefs, experiences, summaries, and procedures. Strategy memories preserve reusable lessons from successful and failed trajectories. These are substantive advances.
Yet a system may remember all relevant content and still act on a conclusion whose entitlement has degraded. Consider a planned EU onboarding rollout. The plan remains current. Its underlying fraud assumption does not. The new fraud observation and the rollout may be semantically distant, and neither fact need supersede the other temporally. What changed is the relation between them: the basis on which the rollout remained licensed.
We call this failure stale reliance:
A system continues to reason from or act upon a unit after a load-bearing ground has been contradicted, superseded, suspended, or made inapplicable, without making the degraded warrant operationally visible at the point of use.
Version 1 of this proposal focused on maintaining explicit warrant state once captured. That remains necessary but is not sufficient. A persistent graph can preserve an omitted assumption, reversed arrow, or fabricated dependency more efficiently than a transcript. Deterministic traversal does not distinguish sound structure from structured error.
The revised question is therefore broader:
How does proposed reasoning become sufficiently well formed, scrutinized, and authorized that a persistent system may rely upon it, and how must that reliance change when reality or governance changes?
The answer cannot be only a data model. It requires a write discipline.
1.1 Thesis
The general category claim is:
Consequential systems need persistent state for current warrants and a governed admission-and-revision process connecting source material, scrutiny, authority, evidence, downstream reliance, and action.
The concrete architectural claim is narrower and stronger:
LTP-native Governed Reasoning Memory is persistent warrant state constructed through the Logical Thinking Process, scrutinized through the Categories of Legitimate Reservation, authorized through separated roles, tested through non-interchangeable evidence and prediction records, and maintained through dependency-sensitive revision and execution.
LTP is constitutive of the proposed architecture. It is not asserted to be the only possible grammar for the broader category. Alternative grammars qualify only if they satisfy equivalent functional and empirical tests.
1.1.1 The memory belongs to the commons
“Remembering why” is useful shorthand, but a weak category boundary. A document, transcript, or episodic memory can already contain reasons. A procedure can preserve a justification. Time matters inside reasoning memory itself. The distinctive proposal is to preserve the relationships and standing that let someone continue reasoning from a previous position: what the goal was, what mattered to a decision, what remained uncertain, whose meaning was represented, and what later changed.
A commons holds that position without requiring its successors to agree with it. An author's confirmation of an interpretation establishes what the author meant; a group's acceptance establishes what it admitted into its model; a scoped reliance decision establishes what it agreed to use. None establishes truth. Dissent can survive a decision to proceed. An adverse outcome can change current reliance while leaving the original judgment legible. Portability lets another person, agent, or institution inspect these distinctions under its own authority.
The unit of value is therefore a continuable inquiry, not an isolated explanation. Warrant maintenance is its operational core, but useful memory also includes uncertain interpretations, alternatives not chosen, unresolved questions, and reasons for revisiting a decision. These records deserve preservation before they qualify as an accepted causal edge. Requiring premature certainty would defeat the purpose of capturing them.
1.2 Contributions
This paper contributes:
- A corrected boundary. Reliable capture and scrutiny are inside Reasoning Memory, not external assumptions about data quality.
- An LTP-native architecture. The six LTP tools are synchronized projections and formation disciplines over shared reasoning state.
- First-class reservations. CLR challenges, responses, resolutions, and waivers persist as governed records.
- A governed admission protocol. Source preservation, atomic capture, scrutiny, risk tiering, separation of authority, and explicit eligibility precede reliance.
- Operational semantics. Scoped dependencies connect changes in warrant to revalidation and execution without treating contradiction as logical negation.
- Practical conformance profiles. Implementations can distinguish a warrant substrate, an LTP-native grammar, and full governed operation without conflating a partial mechanism with the thesis.
- A safety-gated falsification program. The primary empirical outcome requires exact successor understanding and correct treatment of both affected and unaffected action.
- A shared capture and inheritance contract. Source extraction, product exports, and successor inspection use one versioned interchange envelope, with separate declarations of coverage, interpretation, integrity, and local authority.
- Negative evidence. The current model-only pilot is reported as a failure of unguided capture, not presented as support for the architecture.
1.3 What is not claimed
The proposal does not claim that an accepted tree is true, that a consultant guarantees correctness, that graph structure proves causality, that consensus establishes validity, or that every reasoning task requires all six LTP tools. It does not claim that the present product implements the full architecture or that the development pilot estimates a population effect.
It claims that these distinctions can be represented precisely enough to implement and test—and that a system controlling consequential action should not call itself conformant merely because it stores a graph.
2. The Capture Problem Is Part of the Memory Problem
2.1 Oracle structure is not a usable treatment
A benchmark oracle supplied with the correct hidden graph can answer WHY, WHAT-DEPENDS-ON, and action-eligibility queries perfectly. This demonstrates a mechanism: if the operative dependency structure is correct, deterministic maintenance enables the intended operations.
It does not demonstrate that a person or model can construct the graph reliably, that reviewers can detect consequential omissions, or that maintenance cost is justified. Treating oracle performance as evidence for Reasoning Memory would move the central difficulty into the benchmark fixture.
2.2 The model-only instrument result
The repository’s safety-gated paired instrument pilot compared three conditions on three generated RCB v1 worlds using openai/gpt-4o-mini:
| Condition | Safe continuation | Successor exact | Propagation F1 | Stale reliance | Action accuracy | Prompt tokens | Violation incidents |
|---|---|---|---|---|---|---|---|
| Direct full history | 0.0% | 0.0% | 66.7% | 0.0% | 100.0% | 27,898 | 0 |
| Persistent neutral | 0.0% | 0.0% | 0.0% | 0.0% | 100.0% | 124,049 | 8 |
| Persistent LTP | 0.0% | 0.0% | 0.0% | 33.3% | 83.3% | 126,762 | 16 |
This was not preregistered, used one hidden topology rendered in three domains, and has no inferential force beyond apparatus diagnosis. Its trace is nevertheless useful. The persistent conditions produced self-dependencies, hard cycles, and units simultaneously marked challenged and confirmed. Both populated a confirmation set with accepted or mentioned units despite instructions reserving confirmation for observation.
The correct interpretation is not “LTP fails.” The condition supplied LTP instructions to the same unreviewed model updater. It did not include an independent scrutinizer, a CLR facilitator, granular ratification, typed observation records, or governed resolution of reservations. The result instead rejects three shortcuts:
- an explicit output schema is not reliable capture;
- persistent structure is not automatically safer than reconstruction;
- LTP vocabulary in a prompt is not equivalent to LTP practice.
2.3 What the pilot did and did not put at risk
The negative result is evidence against the tested implementation, including its cost and its failure to capture exact structure. It must remain in the evidence record. It does not estimate the effect of the repository's extraction skill, the native interchange document, prospective evidence records, or independent governance: those were not the complete treatment. Conversely, adding those mechanisms does not rehabilitate the failed treatment by definition. Their incremental benefit and cost must be measured against the same strong reconstruction baseline.
An adequate next test must separate representation, capture, retrieval, review, and operational policy. An oracle graph can diagnose traversal; it cannot establish capture fidelity. A schema-valid file can diagnose interchange; it cannot establish that its author understood a source. A stricter gate can reduce stale action by blocking everything; it must also preserve unaffected work. The target is informed continuation at a tolerable total cost.
2.4 Why construction cannot be externalized
If capture quality is outside the definition, Reasoning Memory becomes a structured archival format whose benefits apply when authors happen to provide correct structure. That is a viable but much weaker category. It cannot support strong claims about safe execution, stale reliance, or institutional inheritance.
For the stronger claims, the memory must include the process by which structure enters reliance. The write path must preserve what was said, distinguish extraction from endorsement, expose uncertainty and assumptions, invite legitimate challenge, and make unresolved reservations computationally visible.
The resulting object is socio-technical in the same sense that a scientific record, safety case, or audited ledger is socio-technical. Its correctness depends on representation, rules, roles, and behavior. That dependence does not make it unfalsifiable. It determines what the treatment actually is.
3. Definition and Boundary
3.1 General category
Reasoning Memory is a persistent, inspectable, and revisable representation of what a system is currently licensed to rely upon, for which purpose, on what grounds, after what scrutiny, under whose authority, and with what consequences when those grounds or authorities change.
“Licensed” is institutional and computational. It does not mean believed by a model or certified as true. A unit is licensed when the current policy permits it to support a specified class of reasoning or action.
3.2 Proposed architecture
LTP-native Governed Reasoning Memory (LTP-RM) is Reasoning Memory whose warrant-bearing structures are formed and projected through LTP, whose logical challenges are persisted through CLR, and whose acceptance, evidence, revision, and execution transitions obey RCP.
This definition makes three commitments:
- LTP is native: its roles, cross-tree projections, formation checks, and queries affect system behavior.
- CLR is persistent: scrutiny is not an ephemeral meeting; reservations and their disposition are state.
- RCP is governing: agreement, execution, observation, and assessment remain different authorities and objects.
3.3 Constitutive non-identities
The following distinctions are mandatory:
- source statement ≠ structured interpretation;
- candidate structure ≠ accepted reasoning;
- accepted entity ≠ accepted designation;
- accepted endpoints ≠ accepted relationship;
- consensus ≠ truth;
- ratification ≠ evidence;
- inference ≠ observation;
- action approval ≠ execution;
- execution ≠ expected effect;
- favorable result ≠ valid test;
- temporal currency ≠ warrant validity;
- unresolved reservation ≠ rejection;
- waiver ≠ resolution;
- history ≠ current state.
An implementation that stores these labels while allowing one event to impersonate another is nonconformant.
3.4 Reconstruction, prospective commitment, and the FAQ objection
The project's reconstruction FAQ identifies a stronger objection than “retrieval loses why.” Give a capable successor the original sources and it may reconstruct the reasons perfectly. Where it does so more cheaply, reconstruction should win. The proposed memory must earn its cost through faithful continuity, decision-relevant retrieval, or facts that were recorded through participation and cannot be supplied by later interpretation alone.
Suppose a team approves a limited pilot while retaining an objection, registers a measurable expectation before the observation window, and specifies what would change its mind. The outcome contradicts that expectation. A later model can explain the failure. It cannot establish that this team made that commitment beforehand if no prior record exists. If a trustworthy earlier document already contains the commitment, reconstruction can recover it; structured memory has no monopoly on prospective evidence. Its contribution is to keep that commitment linked to the decision, actual execution, observation method and scope, surviving objection, and resulting revision.
Three claims must remain separate. A timestamp in an imported file is a source assertion. A matching hash demonstrates a limited integrity relation. Establishing who committed to what before an outcome requires independently trusted provenance and timing. An exporter can copy a fabricated chronology and recompute its hashes. The present reader therefore reports these checks separately and does not certify authenticated prospective commitment.
Prospective memory matters even when the action succeeds: a favorable observation from a nonconforming execution is not a valid test of the original prediction. It also matters when the decision was reasonable but the prediction failed. Historical judgment, current reliance, and eventual outcome must remain independently inspectable.
3.5 What remains outside
LTP-RM does not supply a universal theory of evidence, causal identification, utility, ethics, or organizational legitimacy. It records which evidence theory, causal license, benchmark, and authority policy are being used. Domain methods still determine whether an experiment is well controlled, a clinical diagnosis is appropriate, or a safety argument meets regulation.
The architecture is compatible with individual, organizational, and public governance. The normative claim that reasoning should be held as a commons is separable from the computational claim that warranted reliance should persist.
4. Formal State
At time t, an LTP-native governed reasoning space is:
RM⁺(t) = (U, W, Δ, E, H, Π, ΓLTP, R, Σ)
where:
- U is the set of ratifiable reasoning units;
- W is the set of normalized warrant records;
- Δ is the set of scoped dependencies;
- E is the evidence, execution, prediction, and assessment domain;
- H is an append-only governed event history;
- Π is provenance, identity, authority, and risk metadata;
- ΓLTP is the LTP projection and formation grammar;
- R is the set of reservation, response, resolution, and waiver records;
- Σ is the versioned admission, propagation, revision, and execution policy.
The current state is a deterministic projection of H under Σ. No element of the tuple certifies truth.
4.1 Ratifiable units
The substrate retains five unit kinds:
- Proposition: an atomic claim.
- Designation: the separately ratifiable claim that a proposition plays an LTP or domain role.
- Relationship: a reified connection with its own state, provenance, license, and assumptions.
- Assumption: a condition under which a relationship or warrant is licensed.
- Assessment: a conclusion over an explicit heterogeneous ground set.
An action and an expected effect may both be propositions, but their designations and channel semantics differ. An observation is not a proposition proposed for acceptance; it is an evidence record that may bear on propositions.
4.2 Normalized warrant
For a conclusion c and reliance purpose p, a warrant is:
w = (G, ℓ, c, A, X, D, q, p)
where:
- G is a finite set of grounds;
- ℓ is the inference license or relationship type;
- c is the conclusion;
- A is the set of assumptions;
- X is the applicability and scope condition;
- D is the set of known defeaters and active challenges;
- q is an optional qualification or strength;
- p is the permitted reliance purpose.
Examples of p include inquiry, planning, adoption, and execution. A contested causal claim may remain usable as a hypothesis for inquiry while being barred from authorizing rollout.
An implementation may normalize the warrant across tables. Conformance depends on query behavior: WHY(c, p, t) must recover the active grounds, license, assumptions, applicability, challenges, scrutiny status, authority history, and current permission for purpose p.
4.3 Reservation record
A reservation is:
r = (z, k, b, s, a, ρ, d)
where:
- z is the targeted unit, relationship, warrant, or branch;
- k is the reservation kind;
- b is the stated basis of concern;
- s is severity and risk relevance;
- a is authorship and provenance;
- ρ is status and disposition;
- d is the response, revision, evidence, or waiver record.
CLR supplies the native kinds:
clarity | entity_existence | causality_existence | cause_insufficiency | additional_cause | cause_effect_reversal | predicted_effect_existence | tautology
Deployments may add domain reservations. The allowed dispositions are:
open | answered | resolved_by_revision | rejected_with_rationale | waived | superseded
A waiver requires an authorized actor, rationale, scope, expiry or reopening condition, and affected reliance purpose. It does not erase the reservation.
4.4 Scrutiny packet
Every consequential warrant has a queryable scrutiny packet:
S(w) = (source, candidate, checks, reservations, responses, roles, decision)
The packet preserves:
- source excerpts or record identifiers;
- the structured candidate proposed from them;
- applicable CLR and domain checks;
- reservations and responses;
- revisions between candidate versions;
- proposer, scrutinizer, facilitator, and ratifier identities;
- the eligibility and ratification decision under a policy version.
This packet makes capture error inspectable. It also prevents a polished final graph from concealing disagreement or the labor required to produce it.
4.5 Scoped dependency
Each dependency is:
δ = (dependent, prerequisite, strength, scope, rule, provenance)
with strength in {hard, soft} and scope in {ratification, execution, validity}.
- Ratification scope states what must be accepted before another unit can coherently enter reliance.
- Execution scope states what must be observed before an action may execute.
- Validity scope states what must be warned, reviewed, or degraded when a prerequisite’s warrant changes.
Hard dependencies affect eligibility or state. Soft dependencies advise and rank. Methodological sequence must never silently become a hard logical block.
4.6 Roles and authority
The architecture distinguishes functions even when one person holds several low-risk roles:
- source participant: originates a statement or record;
- extractor/tree builder: proposes structure;
- scrutinizer: challenges content or logic, preferably with domain knowledge;
- facilitator: keeps scrutiny within CLR and does not decide domain substance merely by role;
- ratifier: authorizes entry into shared reliance;
- executor: performs an action;
- observer: records what occurred;
- protocol referee: computes decidable structural transitions;
- judgment referee: resolves substantive assessments and escalations.
Σ defines which roles must be separated for each risk tier. “Human in the loop” is not the invariant. Attributable, reviewable, reversible separation is.
5. LTP as a Native Architecture
The proposed system does not store six unrelated diagrams. It stores one shared graph and exposes six LTP projections:
ΓLTP = (PIO, PCRT, PEC, PFRT, PPRT, PTT, CLR)
Each projection selects designations, relationships, warrants, and reservations appropriate to a question about change. A unit may appear in multiple projections without duplication.
5.1 Intermediate Objectives projection: what is the standard?
Reason Commons exposes this projection as the Goal Tree / Intermediate Objectives Map. “IO projection” names the LTP method; goal is the product key. They refer to the same normative view, not competing seventh and first trees.
The IO projection represents:
- system boundary and vision;
- goal;
- critical success factors;
- supporting necessary conditions;
- benchmarks by which necessity and attainment are judged.
Its memory function is normative anchoring. It prevents an intervention from being evaluated against a convenient headline metric when the accepted standard was conjunctive. Goal, CSF, NC, and benchmark designations are separately ratifiable.
Outside scrutiny is required before deep downstream work in consequential spaces because later causal and solution reasoning inherits this standard.
5.2 Current Reality projection: what must change?
The CRT projection represents:
- observed conditions and undesirable-effect designations;
- intermediate and root causal relationships;
- compound causes;
- independent additional causes;
- assumptions and reinforcing loops;
- span of control and sphere of influence;
- reservations on entities and arrows.
Its memory function is diagnosis with explicit load. A critical root cause designation is not intrinsic to a proposition; it is an assessment that a causal position explains material UDEs under the accepted model.
5.3 Evaporating Cloud projection: what assumption creates the conflict?
The EC projection represents:
- a common objective;
- legitimate requirements;
- prerequisites believed necessary for those requirements;
- the conflict between prerequisites;
- assumptions licensing each requirement-prerequisite connection;
- injections proposed to invalidate or bypass an assumption.
Its memory function is disagreement localization. It replaces opposing narratives with a reviewable set of requirements, prerequisites, and breakable assumptions.
5.4 Future Reality projection: what would the intervention cause?
The FRT projection represents:
- injections;
- expected intermediate and desired effects;
- causal assumptions;
- material negative branches;
- branch-trimming mitigations;
- preregistered predictions and guardrails.
Its memory function is prospective risk. A proposal may be stored as an idea without being accepted as sufficient for adoption. Adoption requires review of its positive path and material negative branches under the applicable risk policy.
5.5 Prerequisite projection: what blocks the change?
The PRT projection represents:
- an accepted objective or injection;
- obstacles;
- intermediate objectives that overcome them;
- sequence and parallelism;
- responsibility and applicability constraints.
Its memory function is obstacle-sensitive planning. An obstacle is not automatically a causal ground for the objective, and a methodological preference is not automatically an execution prerequisite.
5.6 Transition projection: what action should happen, and what should it produce?
The Transition projection represents repeating structures of:
- existing reality;
- unmet need;
- action;
- expected effect;
- rationale for the next action.
Reason Commons retains this projection even though Dettmer deemphasized a detailed Transition Tree in favor of richer prerequisite planning. AI lowers construction cost, while RCP makes the action–expected-effect distinction operationally indispensable.
The projection’s memory function is closed-loop execution. Acceptance of the expected effect may support ratification of a later plan; only observation of the effect can satisfy its execution dependency.
5.7 Cross-projection identity
LTP’s value is weakened if each tool becomes a disconnected document. In LTP-RM:
- a CRT root cause may become the assumption challenged by an EC;
- an EC injection becomes an FRT intervention;
- an FRT negative branch becomes a guardrail prediction;
- an FRT injection becomes a PRT objective;
- a PRT intermediate objective becomes a Transition action or need;
- a Transition observation bears back on the FRT and CRT warrants that motivated it.
Identity is preserved across these transformations. Revision therefore propagates across method stages without relying on keyword similarity.
5.8 Applicability, not bureaucracy
Not every task instantiates the entire sequence:
| Reasoning task | Required native machinery |
|---|---|
| Define consequential success | IO projection and benchmark provenance |
| Diagnose a persistent problem | CRT and CLR scrutiny |
| Resolve an apparent conflict | EC assumptions and reservations |
| Evaluate an intervention | FRT, predicted effects, and negative branches |
| Plan around obstacles | PRT |
| Govern execution and learning | Transition projection and evidence channel |
| Admit a consequential causal claim | CLR scrutiny regardless of originating projection |
Σ records the applicability decision and any exemption. A low-risk factual note need not carry an FRT. A high-risk intervention cannot silently exempt itself from consequence and negative-branch review.
6. CLR as the Governed Write Protocol
CLR is often presented as a checklist for reviewing a finished tree. In LTP-RM it governs both construction and admission. Each category addresses a characteristic capture failure.
6.1 Clarity
The clarity reservation asks whether the statement and arrow are understood as intended. The system supports it by requiring:
- one material claim per ratifiable entity;
- explicit quantifiers, population, time, and comparison where relevant;
- separation of embedded “because” or “in order to” clauses into entities and relationships;
- source-preserving paraphrase rather than silent normalization;
- localized requests for missing intermediate steps.
Automated checks can detect compound syntax and undefined terms. Only participants can establish that the wording captures the intended claim.
6.2 Entity existence
Entity existence asks whether a current-reality statement is well founded. The system requires provenance and source class and allows reviewers to request supporting records.
Entity existence is channel-sensitive. A plan may exist as a plan; its expected future condition does not yet exist as observation. An agent inference may exist as an inference; it does not become a sensor reading because it is confidently expressed.
6.3 Causality existence
Causality existence challenges the arrow rather than merely its endpoints. A causal relationship must expose:
- direction;
- inference license;
- applicability conditions;
- supporting evidence or rationale;
- assumptions;
- reservations and counterevidence.
The system should support verbalization in the form “if cause, then effect, because assumptions.” Verbal fluency is not proof, but ambiguity under verbalization is evidence of an under-specified link.
6.4 Cause insufficiency
Cause insufficiency asks whether the stated causes are enough. Compound causes must be represented explicitly rather than flattened into independent arrows whose conjunction disappears.
A cause may contribute without being sufficient. The warrant license records that difference. A dependent conclusion cannot claim a stronger license than its reviewed grounds provide.
6.5 Additional cause
Additional cause asks what else could independently produce the effect. This prevents a preferred explanation from receiving all causal credit and enables revision when one path is weakened but another remains.
Alternative causes are not mandatory clutter. Risk policy determines how aggressively they must be sought. For consequential diagnosis, at least one explicit alternative-cause pass or attributable exemption is required.
6.6 Cause-effect reversal
Cause-effect reversal challenges direction. The system can detect reciprocal hard edges and suspicious cycles, but direction usually requires domain judgment or temporal and experimental evidence.
Reversal reservations remain distinct from general causality reservations because they imply a specific repair: reverse, split, or replace the relationship.
6.7 Predicted effect existence
Predicted-effect existence asks what other observable consequence should follow if a proposed cause or intervention is real. In LTP-RM, this becomes the bridge from causal argument to prospective test.
For load-bearing causal warrants, the scrutiny packet should contain at least one of:
- a preregistered prediction;
- a linked existing observation that discriminates the claim;
- a reasoned exemption explaining why prospective testing is infeasible.
Predicted effects do not prove the original cause. They make the claim risk-bearing and revisable.
6.8 Tautology
Tautology addresses circular support. Hard cycles within a dependency scope are rejected unless a grammar explicitly represents a reviewed reinforcing loop and supplies terminating semantics.
The model-only pilot’s cycles were therefore not merely untidy graphs. They were admission failures that a conformant write path should have prevented from controlling execution.
6.9 Human and machine allocation
CLR does not imply that every check is a manual form:
| Check class | Good candidate for automation | Requires substantive judgment |
|---|---|---|
| Structural clarity | Compound syntax, missing identifiers, schema validity | Whether the statement captures intended meaning |
| Entity provenance | Source linkage, timestamps, completeness | Reliability, relevance, applicability |
| Causal graph hygiene | Cycles, direction conflicts, missing endpoints | Whether the causal connection is credible |
| Sufficiency structure | Missing compound-group representation | Which cofactors are materially required |
| Predicted-effect completeness | Presence and preregistration timing | Whether the prediction discriminates the claim |
| Negative-branch completeness | Presence of a branch and guardrail fields | Whether material harms have been omitted |
| Authority and waiver compliance | Role separation, scope, expiry | Whether accepting the residual risk is legitimate |
Automation prepares judgment; it does not impersonate it.
7. Admission, Ratification, and Revision
7.1 Candidate lifecycle
The proposal channel uses the following lifecycle:
source record
↓ extraction
structured candidate
↓ applicable automated checks
under scrutiny
├── revision requested ──→ new candidate version
├── rejected
├── withdrawn
├── contested candidate
└── eligible
↓ authorized ratification
accepted/current
↓ evidence, reservation, policy, or source change
contested → needs revalidation → suspended or revalidated
↓ replacement
superseded
Proposal state and accepted validity state remain separate. A pending candidate cannot silently edit current reasoning. An accepted unit can become contested without becoming a new proposal.
7.2 Eligibility rule
For a consequential warrant w, eligible(w, t) holds only when:
- source provenance is complete and accessible under policy;
- conclusion, grounds, relationships, assumptions, and designations are separately addressable;
- the inference license and reliance purpose are explicit;
- applicable LTP projection rules are satisfied;
- required CLR and domain checks are recorded;
- no material reservation is silently unresolved;
- every waiver is authorized, scoped, reasoned, and time- or condition-bounded;
- external hard ratification prerequisites are accepted or included earlier in the same atomic batch;
- required proposer, scrutinizer, facilitator, and ratifier separation is satisfied;
- the candidate is based on the current source and policy versions.
Eligibility means structurally and procedurally ready for judgment. It does not compel ratification.
7.3 No silent reservation
A material reservation can leave the active review surface only by:
- revision of the target that resolves it;
- evidence or argument accepted as an answer;
- rejection with attributable rationale;
- explicit waiver by authorized judgment;
- supersession of the target.
Closing a comment, deleting a draft, or producing a new summary is not a disposition. The original reservation remains recoverable.
7.4 Waiver discipline
Waivers are necessary because evidence and scrutiny are never complete. A waiver record contains:
- reservation and affected warrant;
- authorizing actor and role;
- accepted residual risk;
- permitted reliance purpose;
- population and temporal scope;
- expiry or reopening trigger;
- rationale and dissent if present.
A waived causality concern may permit a bounded experiment while still blocking broad adoption. Waiver therefore narrows permission rather than laundering uncertainty.
7.5 Risk-tiered governance
Governance cost should be proportional to consequence and reversibility:
| Tier | Typical use | Minimum admission discipline |
|---|---|---|
| 0 | Private note or exploratory hypothesis | Source attribution; no claim of shared reliance |
| 1 | Low-consequence shared reasoning | Automated structure checks; attributable proposer and ratifier |
| 2 | Material planning or organizational decision | Independent scrutiny of grounds, assumptions, and scoped dependencies |
| 3 | High-consequence intervention | Domain scrutinizer; CLR-trained facilitation; prediction, guardrail, negative-branch, and explicit waiver review |
| 4 | Safety-critical or regulated action | Policy-defined independent evidence, multi-role judgment, auditable assurance, and hard execution gates |
Risk tier is itself a ratifiable assessment. Participants cannot reduce governance merely by relabeling a high-consequence action as an idea after operational projection has begun.
7.6 Atomic ratification and staleness
One accepted proposal may atomically introduce related entities, relationships, assumptions, designations, and assessments. Within the transaction, hard prerequisites must be ordered. Outside it, the whole batch either commits or fails.
If accepted reasoning or source context changes during review, affected candidates become stale and require rebase. Review of an obsolete version cannot authorize a newer one.
7.7 Consultant-governed capture
A human LTP consultant can occupy two distinct functions:
- as tree builder, the consultant elicits and structures candidate reasoning;
- as facilitator, the consultant translates domain concerns into legitimate reservations and keeps scrutiny focused on logic rather than status or personality.
Subject-matter participants remain necessary for entity and causality judgment. An independent ratifier remains necessary where policy requires. The consultant is part of the write institution, not an oracle who converts workshop notes directly into truth.
8. Evidence, Prediction, and Closed-Loop Learning
8.1 Two non-interchangeable channels
The ratification channel governs what proposed reasoning enters shared reliance.
The evidence channel appends execution, observation, test-validity, and prediction-result records. Evidence exists whether or not participants like its implication. Its relevance and reliability may be assessed; the record itself cannot be vetoed as though it were a proposal.
No role may use one channel to impersonate the other:
- ratification cannot create an observation;
- an observation cannot silently rewrite accepted reasoning;
- an agent inference cannot be relabeled as direct measurement;
- an external ticket closure cannot be relabeled as an expected effect.
8.2 Prediction specification
Before its result is knowable, a prediction records:
- target warrant or claim;
- intervention or condition;
- population and applicability;
- metric and measurement procedure;
- observation window;
- confirming, contradicting, and inconclusive regions;
- primary outcome and guardrails;
- source and independence expectations;
- proposer and ratifier.
Late predictions remain as hypotheses but cannot confirm or contradict the already-known outcome.
8.3 Execution, observation, and assessment
The evidence lifecycle separates:
- authorization: permission to act;
- execution record: what was attempted, by whom, when, and under what conditions;
- observation record: what was measured or reported, by what source;
- test-validity assessment: whether execution and measurement instantiated the registered test;
- prediction-result assessment: whether a valid test confirmed, contradicted, partially supported, or failed to resolve the prediction;
- warrant revision: what accepted reasoning and action eligibility change.
A favorable observation from a nonconforming test does not confirm the target. A failed execution does not automatically falsify the causal claim. This two-level assessment prevents implementation error and theory error from collapsing.
8.4 Progressive revision
When evidence challenges a claim, permitted revision patterns include:
- R1 — explanation revision: introduce a new mechanism that predicts additional risk;
- R2 — scope narrowing: reduce applicability and register a new boundary prediction;
- R3 — evidence challenge: contest reliability or relevance with better evidence;
- R4 — benchmark revision: change the standard and expose that normative change separately.
A revision cannot clear a challenge merely by restating the claim more weakly after the result. Rescued claims incur new falsifiable commitments.
8.5 Evidence quality
Evidence records include source class, provenance, collection method, independence relationships, uncertainty, and applicability. Source class alone is not a numeric evidence theory. Two agents sharing a model and context are correlated; two documents may copy one report; two sensors may share a failure mode.
The substrate supports explicit qualitative or domain-specific aggregation without imposing false universal arithmetic.
9. Reliance and Operational Semantics
9.1 Canonical queries
At minimum, a conformant system supports:
| Query | Required result |
|---|---|
WHY(x, purpose, t) |
Active grounds, license, assumptions, scope, evidence, reservations, authority, and state |
WHAT_DEPENDS_ON(x, scope, t) |
Transitive dependents with path, strength, purpose, and consequence |
WHAT_WOULD_CHANGE_OUR_MIND(x, t) |
Open predictions, defeaters, reservations, expiry conditions, and discriminating observations |
WHAT_IS_CONTESTED(t) |
Non-clean units and reservations grouped by consequence and review owner |
WHAT_IS_EXECUTABLE(t) |
Accepted actions whose observation prerequisites and validity support permit execution |
WHAT_IS_BLOCKED_BY_CAPTURE(t) |
Candidates blocked by source, scrutiny, reservation, role, or staleness failures |
WHICH_LTP_VIEW_EXPLAINS(x, t) |
Goal, diagnosis, conflict, future, prerequisite, and transition projections containing the unit |
WHICH_ACTIONS_REST_ON_WAIVED_RISK(t) |
Executable or planned actions with warrant paths containing active waivers |
9.2 Validity states
Accepted units carry one of:
- current: permitted for declared reliance purposes with no unresolved qualifying challenge;
- contested: credible contrary evidence or a material reservation is attached;
- needs revalidation: structured review is required because a load-bearing basis or policy materially changed;
- suspended: barred from hard downstream reliance and execution;
- superseded: retained historically but excluded from the current projection.
Validity is purpose-aware. Policy may permit contested use for inquiry while blocking adoption or execution.
9.3 Graduated propagation
Let hard-validity closure V⁺(x) be the transitive dependents of x through hard validity edges. The default policy is:
- a qualifying contradiction contests its direct target;
- every member of V⁺(x) receives a traceable advisory naming the origin;
- risk, severity, repetition, or policy may escalate selected units to needs revalidation;
- suspension moves hard dependents to needs revalidation and revokes dependent execution immediately;
- material supersession triggers review of hard dependents;
- soft edges carry advisories only.
This propagates warrant degradation, not truth-functional negation. Safety-critical deployments may choose eager state propagation; advisory and eager modes must be compared empirically because each trades missed reliance against alert burden.
9.4 Frontiers
The Ratification Frontier contains eligible candidates whose hard ratification prerequisites are accepted or included earlier in one atomic batch.
The Execution Frontier contains accepted actions whose hard execution prerequisites are satisfied by observations and whose warrant paths permit execution.
The Scrutiny Frontier contains candidates or current units requiring the next scarce judgment: unresolved material reservations, expiring waivers, stale reviews, or escalated evidence challenges.
The Learning Frontier contains load-bearing predictions whose windows are open or whose outcomes await assessment.
These frontiers expose different constraints. More proposal generation does not help when scrutiny or observation is the bottleneck.
9.5 Reasoning inheritance
A successor receives the governed current projection, open reservations, waivers, evidence state, policy versions, and cross-LTP views—not merely a summary. The successor can inspect why a conclusion is licensed, where experts disagreed, what was waived, and which observations would alter it.
Inheritance is not forced agreement. It is the transfer of an inspectable epistemic position from which disagreement and revision can continue without reconstructing every conversation.
9.6 A decision-relative attention contract
A successor should be able to recover the case relevant to the next decision without replaying the whole archive: the purpose and scope, operative grounds, unresolved reservations, the prospective expectation and its test conditions, observed results, and the changes that affect reliance now. The full record must remain inspectable behind that projection. A concise answer is useful only if it does not conceal the objection or scope limitation that changes the decision.
This makes attention a separate empirical question. Measure which material qualifications a successor recovers, how long recovery takes, and whether it can trace them to source records. Do not count shorter output as successful compression when it removes uncertainty. Do not treat the current projection as a replacement for the original case.
10. Practical Implementation
10.1 Minimum data domains
A practical relational or graph implementation requires:
- source records and immutable excerpts;
- reasoning units and versions;
- designations and LTP projection membership;
- reified relationships and inference licenses;
- assumptions and applicability scopes;
- scoped dependencies;
- proposals and atomic proposal items;
- CLR and domain reservations;
- responses, resolutions, and waivers;
- scrutiny sessions and role assignments;
- ratification events;
- predictions, executions, observations, and two-level assessments;
- validity transitions and propagation notices;
- versioned policies and risk assessments.
Event history is canonical; materialized views provide current projections and frontiers.
10.2 Write APIs
The agent and interface contract should separate verbs:
capture_source
propose_entity
propose_designation
propose_relationship
attach_assumption
open_reservation
respond_to_reservation
revise_candidate
waive_reservation
ratify_batch
record_execution
record_observation
assess_test_validity
assess_prediction
contest_unit
suspend_unit
supersede_unit
There should be no generic update_graph operation capable of performing all of these transitions under one authority.
10.3 Capture interaction
A usable workflow is:
- preserve the original statement, document, meeting segment, or instrument record;
- let an extractor propose atomic units and links with confidence and uncertainty;
- show the source beside the candidate structure;
- run deterministic checks before consuming reviewer attention;
- route material entity and causal questions to domain scrutinizers;
- let a CLR facilitator localize concerns and prevent interpersonal argument from replacing logical review;
- revise candidates without mutating accepted state;
- ratify granular units or a dependency-ordered batch;
- display every unresolved reservation and waiver on downstream reliance paths.
The interface should optimize correction, not affirmation. Reviewers need to remove invented links, add omitted assumptions, split compound claims, change scope, and express “uncertain” as easily as they accept.
10.4 Risk-aware defaults
The product should infer a provisional risk tier from projected action, reversibility, affected population, and external constraints, then require attributable confirmation. High-risk defaults include:
- independent scrutiny;
- negative-branch and guardrail coverage;
- no agent-only ratification;
- explicit waiver authority;
- observation-gated execution;
- shorter waiver expiry;
- stronger evidence-independence requirements.
Low-risk exploratory work remains cheap and clearly marked non-authoritative.
10.5 Deterministic and judgment boundaries
Deterministic services should handle schema validation, provenance completeness, cycles, stale versions, dependency closure, decidable prediction rules, frontiers, expiry, and replay.
Models may propose structure, summarize sources, identify possible reservations, generate alternative causes, and draft predictions. They do not silently decide that an inference is an observation or that their own proposal has survived substantive scrutiny.
Human judgment is not intrinsically superior. It is scarce, biased, and fallible. Its value here is accountable participation by people with domain access, authority, or lived context unavailable to the model. Policies may delegate judgments to agents when evidence justifies that delegation.
10.6 Constraint-aware operation
The system’s throughput is validated learning, not accepted-node count. It should measure:
- load-bearing prediction coverage;
- scrutiny queue depth and latency;
- reservation yield and resolution pattern;
- consequential omission survival;
- waiver volume and expiry;
- observation and assessment latency;
- contradiction yield;
- suspension-to-execution-revocation latency;
- reasoning reconstruction cost;
- reviewer time per accepted warrant.
At a defined cadence, participants identify whether generation, scrutiny, ratification, execution, observation, assessment, or revision currently limits learning. Adding more AI-generated structure while scrutiny is constrained increases inventory rather than knowledge.
10.7 One contract from capture to inheritance
The repository implements this relationship through the LTP interchange specification, §19. Version 1.2 adds optional structured capture and memory sections to the existing graph and history envelope. The extraction skill is a producer of candidate interpretations in that envelope; it is not a second memory architecture and does not govern its own output into acceptance.
| Layer | What is preserved | What it does not establish |
|---|---|---|
| Source capture | Locators, excerpts, interpretation status, extraction confidence, unresolved issues and reasons for withholding structure | Author agreement, causal validity, or local acceptance |
| Candidate model | Entities, designations, relationships, assumptions, assessments, including their supplied status | Permission to replay source status as a local decision |
| Decision and learning records | Meaning records, reservations and events, manifests, decisions, reliance and holds, predictions, executions, observations and assessments where carried | Authenticated provenance or a complete implementation of the profiles below |
| History and integrity | Superseded units and relationships, source events, version records, carried digests | Proof that a claim is true or was registered before an outcome |
| Local inheritance | Inspection of the source position and an ordinary proposal for its supported graph interpretation | Transfer of the source commons' authority |
A source can say “A and B together may cause C within the pilot.” Capture must preserve the joint condition, modality and scope. If the importer cannot represent that joint support, the skill records an importer_unsupported issue. It must not select a dominant cause, invent two independent sufficient causes, or discard the passage. Extraction confidence concerns fidelity to the source; it is not confidence that the proposition is true. Source-explicit language can itself be uncertain.
The human capture report is generated from the machine-readable capture section, so caveats are portable rather than trapped in a sidecar. The current native PDF conversion service remains a separate producer with its own provenance representation; the revised skill must not be described as already governing that service. Neither the present importer nor the operational engine gains compound-edge semantics merely because a schema can carry multiple endpoints.
The implemented export readers collect the supported native decision and learning records. Missing schema is declared as omitted; authorization or transport failures fail the read. Pagination avoids treating a server's row limit as a complete collection. Completeness is scoped to exporter-readable records, and these reads are not an atomic database snapshot. A consumer must distinguish coverage (what was carried), capability (what this reader interprets), integrity (which digest relation was checked), and conformance (which governed behavior was enforced). None substitutes for another.
On attachment, a read-only successor view joins a decision to its retained objections, prediction, outcome and later holds. Import still proposes supported graph changes through ordinary local review. With the source-retention migration installed, the original parsed document is saved atomically with that proposal and can be carried in a later export. Foreign decisions remain source records; their actors, observations and permissions are not recreated as local events. This is an implemented inheritance slice, not a complete restore or federation protocol.
11. Conformance Profiles and Invariants
Calling every partial implementation “Reasoning Memory” would make the thesis impossible to test. This paper defines nested profiles.
11.1 Profiles
RM-Core — Warrant substrate
- normalized warrants;
- provenance and recoverable history;
- scoped dependencies;
- purpose-aware validity;
- WHY and WHAT-DEPENDS-ON queries.
RM-LTP — LTP-native construction
- RM-Core;
- first-class LTP designations;
- six synchronized projections;
- CLR reservation records;
- tool-specific formation rules and queries.
RM-Governed — Full proposed treatment
- RM-LTP;
- source-preserving capture;
- risk-tiered scrutiny and ratification;
- non-interchangeable evidence channel;
- prediction and two-level assessment;
- reservation resolution and waiver discipline;
- dependency-sensitive revision and execution.
An experiment must name the profile tested. A model-only graph updater is not RM-Governed.
11.2 Core invariants
C1 — Warrant traversability. Every non-leaf reliance-bearing unit exposes grounds, inference license, assumptions, applicability, active defeaters, provenance, purpose, and current state.
C2 — Relationship independence. Accepting endpoints does not accept the relationship between them.
C3 — History recoverability. Supersession preserves predecessors, and current state is reproducible from governed events and policy versions.
C4 — Scoped dependency integrity. Ratification, execution, and validity scopes remain distinct; soft methodology never silently becomes a hard block.
C5 — Deterministic projection. Canonical state, frontiers, propagation notices, and expiry are pure functions of valid events and versioned policy.
11.3 Capture and governance invariants
G1 — Traceable capture. Every structured interpretation links to its source and extractor; normalization does not erase the original.
G2 — Granular admission. Entity, designation, relationship, assumption, and assessment are separately reviewable and ratifiable.
G3 — No silent reservation. A material reservation remains visible until attributable resolution, rejection, waiver, or supersession.
G4 — Waiver accountability. A waiver records authority, rationale, scope, residual risk, and expiry or reopening condition.
G5 — Separated authority. No consequential proposal enters shared reliance solely through an act of its proposer under tiers requiring separation.
G6 — Stale review protection. A review or ratification of version v cannot authorize materially changed version v+1.
11.4 LTP-native invariants
L1 — Role independence. Acceptance of a proposition does not accept its LTP designation.
L2 — CLR coverage. Each consequential causal warrant records applicable CLR checks, reservations, and dispositions.
L3 — No prohibited tautology. Hard support cycles are rejected unless explicitly represented as reviewed reinforcing loops with terminating semantics.
L4 — Compound-cause preservation. A conjunctive cause is not flattened into independent sufficient causes.
L5 — Conflict assumption visibility. A claim that an injection breaks a conflict identifies the requirement-prerequisite assumption it invalidates or bypasses.
L6 — Negative-branch coverage. A high-risk sufficiency assessment cannot become current without review of material negative branches and guardrails or an authorized exemption.
L7 — Action-effect separation. A Transition action and its expected effect are different units with different evidence semantics.
L8 — Cross-projection identity. A unit reused across LTP views retains stable identity, provenance, validity, and reservations.
11.5 Evidence and operational invariants
E1 — Agreement does not imply occurrence. Ratification never satisfies an execution dependency requiring observation.
E2 — Evidence is append-only. Disputed observations remain reachable; corrections and assessments attach rather than erase.
E3 — Tests are prospective. A prediction registered after its result is knowable cannot receive a substantive confirmatory or contradictory verdict.
E4 — Test validity precedes result. Prediction verdicts identify the test-validity assessment on which they rely.
E5 — No silent contradiction. A qualifying contradiction contests its target and becomes traceably visible throughout hard-validity dependence within one propagation cycle.
E6 — Suspension has consequence. An action leaves the Execution Frontier immediately when a hard support unit is suspended.
E7 — Rescue costs new risk. Scope or benchmark revision cannot clear a challenge without a new falsifiable commitment or explicit normative change.
E8 — External completion is not success. Closing a projected work item records an attempt or report, never an observed expected effect.
These invariants are mechanism claims. Passing them is necessary for conformance, not evidence that the architecture improves decisions.
12. Related Work and the Reduction Objection
12.1 Truth maintenance and belief revision
Doyle’s TMS and de Kleer’s ATMS maintain dependency-sensitive belief state and assumptions. AGM theory formalizes rational revision of logically closed theories. LTP-RM inherits the central insight that conclusions depend on recorded justifications and should respond when those justifications fail.
It differs in target and institution: finite reviewable artifacts rather than logical closure; purpose-specific reliance rather than belief alone; proposal, scrutiny, authority, observation, and execution as separate domains; and persistence across replaceable reasoners. A TMS extracted from the same public record remains a required simpler baseline.
12.2 Argumentation and warrants
Toulmin structures grounds, warrants, qualifiers, rebuttals, and claims. Defeasible and structured argumentation formalize attack and acceptability. These traditions are direct antecedents of normalized warrants and reservations.
The live claim is not that warrants or rebuttals are new. It is that their governed admission, cross-LTP construction, evidence lifecycle, and operational consequence form useful persistent agent state.
12.3 Design rationale and organizational memory
IBIS, design-rationale systems, and organizational memory preserve questions, alternatives, decisions, and histories. Their mixed adoption record is a warning: authorship and maintenance burdens can exceed later reconstruction savings. H7 treats that warning as a central falsification condition rather than a deployment detail.
12.4 Provenance, scientific claims, and assurance cases
W3C PROV records derivation and responsibility. Micropublications structure claims, evidence, and argument. Assurance cases connect claims to evidence, and dynamic safety cases update with system change.
Provenance answers where an item came from; warrant answers why it is presently licensed for a purpose. LTP-RM requires both. Dynamic assurance is especially close where degraded evidence triggers review of operational claims.
12.5 Contemporary agent memory
Agent memory is increasingly structured and reflective:
- Hindsight distinguishes facts, experiences, summaries, and evolving beliefs.
- A-TMA distinguishes current, historical, and transition state to prevent stale temporal overlays.
- A-MEM develops linked agentic notes.
- Mem0 addresses scalable long-term agent memory.
- Zep/Graphiti represents temporal knowledge-graph state.
- HippoRAG performs associative graph retrieval.
- ReasoningBank, ReMe, and Reflexion preserve reusable strategies or reflective lessons.
- REMem reasons with episodic memory.
- LongMemEval and LoCoMo test long-horizon conversational memory.
These systems invalidate any claim that contemporary memory is merely semantic retrieval. The narrower question is whether they persist the explicit warrants, reservations, authority state, and purpose-scoped operational dependencies required here—or can reconstruct them at lower total cost.
The acronym “LTP” in A-TMA-related evaluation may also denote LoCoMo Temporal Plus. It is unrelated to the Logical Thinking Process used in this paper.
12.6 The reduction objection
The strongest objection is:
This is an LTP argument graph plus workflow, provenance, event sourcing, and a truth-maintenance engine.
At component level, that is substantially correct. The proposal earns a distinct name only if the coupling yields capabilities or economics not matched by simpler combinations.
The reduction tests are therefore explicit:
- If a classical TMS plus extracted text matches the full system, remove the extra architecture.
- If generic independent review matches CLR/LTP review, reject the LTP contribution.
- If temporal belief memory prevents stale reliance without explicit warrants, remove validity scope.
- If expert summaries match inheritance at lower cost, treat the problem as handoff authorship.
- If the full system aids audit but not action, call it an audit protocol rather than a general memory advance.
13. Falsifiable Hypotheses
All comparisons require information parity, matched or fully reported budgets, held-out worlds, frozen prompts and policies, world-level inference, and publication of negative results.
H1 — Governed capture
Independent scrutiny and granular ratification reduce exact-warrant error, load-bearing assumption omission, and invalid dependency scope relative to model-only persistent capture.
Failure: governed capture does not improve consequential fidelity or creates offsetting abandonment and latency.
H2 — LTP contribution
CLR/LTP-guided scrutiny reduces causal-direction error, compound-cause omission, hidden conflict assumptions, unexamined negative branches, and circular support relative to grammar-neutral independent review.
Failure: generic review reaches a preregistered equivalence margin at equal or lower total cost.
H3 — Warrant recovery
RM-Governed recovers operative grounds, assumptions, evidence, reservations, waivers, and revision history more accurately than frontier long context, agentic RAG, temporal graph memory, belief memory, TMS extraction, and expert summaries.
Failure: the strongest baseline matches exact recovery at equal or lower total cost.
H4 — Dependency-aware contradiction
When semantic similarity and dependency are decorrelated, RM-Governed identifies affected reliance with higher precision and recall than retrieval and temporal baselines.
Failure: a baseline matches propagation F1 or users disable alerts because precision is unacceptable.
H5 — Operational safety
Agents using RM-Governed take fewer actions whose load-bearing grounds are suspended, superseded, or contradicted, without indiscriminately blocking independent actions.
Failure: a cheap prompt to check updates matches both stale-reliance and unrelated-block rates.
H6 — Reasoning inheritance
After reasoner replacement, successors given RM-Governed recover current warrants, open reservations, waivers, uncertainties, and action eligibility better than successors given transcripts, temporal memories, or expert handoff memos.
Failure: a handoff memo matches safe continuation at lower preparation and query cost.
H7 — Separated epistemic authority
Systems separating extraction, ratification, execution, observation, and assessment adopt fewer self-confirming and post-hoc lessons than single-agent self-critique, without suppressing valid learning.
Failure: self-critique matches erroneous-adoption and net-improvement rates with higher throughput.
H8 — Ratification tax is payable
Over a defined reuse horizon, verified capture, scrutiny, and maintenance cost less than reconstruction work and consequential error cost saved.
Failure: the total governed-maintenance cost persistently exceeds the avoided reasoning tax for the target risk tier.
H9 — Value beyond temporal and belief state
Explicit warrants and reservations improve safe continuation where a ground changes while a dependent decision remains temporally current and semantically distant.
Failure: A-TMA-class state overlays or Hindsight-class belief memory match the outcome without explicit warrant dependencies.
H10 — Failure becomes reusable capital
RM-Governed reduces repetition of proposals that are superficially different but depend on assumptions already invalidated in the same applicability domain.
Failure: outcome logs plus strong retrieval match repeated-failure and reconstruction cost.
14. Experimental Program
14.1 The decisive factorial comparison
The next high-value experiment uses the same frozen model, information, events, candidate catalogue, and probes across four conditions:
- Full history: answer directly from the complete record.
- Unreviewed persistent graph: model-only incremental structure with deterministic traversal.
- Governed generic RM: independent review, reservations, and granular admission without LTP-specific guidance.
- LTP-native governed RM: the same governance plus native LTP projections and CLR scrutiny.
An oracle with hidden gold structure is a mechanism upper bound, not a competitor.
The contrasts isolate:
- condition 2 versus 1: persistence plus traversal;
- condition 3 versus 2: independent governance;
- condition 4 versus 3: LTP’s incremental contribution;
- condition 4 versus 1: the full treatment against reconstruction.
14.2 Primary outcome: safe continuation
At world level, safe continuation requires all of:
- exact successor warrant recovery;
- correct identification of the affected reliance closure;
- blocking an action resting on suspended support;
- permitting a semantically similar but independent action;
- correct reporting of open material reservations and waivers.
Safety is a gate. A treatment that improves warrant F1 while increasing stale action or indiscriminate blocking does not succeed.
14.3 World construction
Confirmatory worlds must be independently authored, private, and structurally diverse. Required topology families include:
- IO necessity and benchmark failures;
- CRT compound causes, additional causes, and reversal traps;
- EC conflicts with multiple plausible break assumptions;
- FRT negative branches, guardrail tradeoffs, and mitigation failures;
- PRT obstacle, precedence, and false-prerequisite cases;
- Transition acceptance-versus-observation errors;
- supersession, invalid tests, waivers, policy changes, and reasoner replacement;
- semantically similar decoys and distant true dependents;
- evidence duplication, source correlation, and adversarial rationalization.
World authors must not write or tune treatment prompts. Prompt authors must not access private gold.
14.4 Capture treatments
Each persistent condition builds state incrementally. Governed conditions use a candidate-and-review protocol:
- extractor sees previous accepted state and one new source event;
- extractor proposes granular units, warrants, designations, and reservations;
- automated checks reject structural nonconformance;
- an independent reviewer receives source, proposal, and relevant accepted context;
- reviewer accepts, revises, adds reservations, or abstains;
- policy computes eligibility;
- authorized ratification commits accepted items;
- fixed traversal answers probes.
Model-only and human-reviewed variants should both be tested. A model reviewer is not evidence for human consultant performance; a human consultant is not evidence that automated capture scales.
14.5 Metrics
- exact warrant and reservation recovery;
- load-bearing assumption omission;
- causal-direction and compound-cause error;
- negative-branch and guardrail coverage;
- invalid dependency-scope rate;
- propagation precision and recall;
- stale reliance and unrelated-action blocking;
- safe continuation rate;
- successor uncertainty calibration;
- invariant and waiver-policy violations;
- reviewer correction count and consequential-error survival;
- prompt, completion, tool, and wall-clock cost;
- human active time, queue age, and abandonment;
- downstream decision quality, reported separately from memory fidelity.
14.6 Staged program
Stage 0 — Conformance. Property and adversarial tests verify invariants, replay, cycles, channel separation, reservation persistence, and execution revocation.
Stage I — Instrument pilots. Cheap models and public generated worlds diagnose prompts, schemas, scoring, and failure modes. No hypothesis claims.
Stage II — Private structural benchmark. Frozen frontier models, multiple topologies, strong baselines, paired comparisons, and preregistered margins test H1–H6 and H9.
Stage III — Human capture-cost study. Counterbalanced participants verify extraction, author structure, or reconstruct later. This estimates fidelity, consequential omission, review cost, and break-even reuse for H8.
Stage IV — Longitudinal adversarial handoff. Real teams work on consequential projects; original reasoners are replaced; new evidence challenges a load-bearing assumption; successors must continue and decide. This tests the institutional thesis, H6–H8, and downstream decision quality.
14.7 Statistical discipline
The independent unit is a world, contradiction injection, participant, or project as defined before analysis—not each probe derived from one world. Pilot worlds cannot be reused for confirmatory inference after prompt tuning. Report paired world-level outcomes and intervals, not only aggregate micro-F1. Select sample size from pilot variance and the smallest effect worth detecting, not the desired conclusion.
14.8 Testing the product contract rather than a proxy format
The next inheritance experiment must use the same 1.2 document and actual import analysis as the product. The inheritance protocol separates a synthetic mechanism fixture from a proposed model and human study. The fixture demonstrates preservation and separation of authority; it reports no comparative performance result.
First measure the revised skill against direct reconstruction and generic structured extraction on identical sources, including hedging, joint causes, conflicting accounts, missing context and method-dependent observations. Score faithful preservation and consequential omission before scoring graph recall. Independent adjudicators should be allowed to say the source does not determine a unique graph. Expose uncertainty rather than rewarding a fabricated exact match.
Then hand a fresh successor either the full sources, a strong reviewed handoff memo, generic structured memory, or the shared reasoning-memory artifact. Equalize source access and account separately for extraction, review, retrieval, repair and continuation costs. Include post-outcome reconstruction both with and without a trustworthy prospective record; only the former can legitimately recover a prior commitment. Test a contradicted load-bearing assumption alongside an unaffected branch. Record whether the successor recovers the original decision without mistaking it for current or local authorization.
Preregister the primary outcome and margins before running held-out cases. Publish failed imports, abstentions, review time, adverse interventions, and false blocking. The governance and LTP ablations in §14.1 remain necessary: success of an integrated workflow would not identify which component earned its cost.
15. Failure Modes, Security, and Institutional Risk
15.1 Structured error propagation
Incorrect structure can propagate confidently. Automated checks catch cycles and missing fields but not all omitted or false causal links. High-consequence execution therefore requires governed capture and visible uncertainty, not merely deterministic traversal.
15.2 Consultant capture
An LTP consultant can improve structure while acquiring disproportionate influence over framing, vocabulary, and admissibility. The facilitator role must not silently become final authority over domain truth. Source preservation, independent scrutiny, dissent records, and appeal prevent method expertise from becoming epistemic monopoly.
15.3 Bureaucratic overload
Universal high-tier review would cause queues, abandonment, shadow decision-making, and ritual ratification. Risk tiering, automated preparation, localized reservations, and constraint metrics are necessary for viability.
15.4 False consensus
Non-confrontational CLR language can localize disagreement, but consensus may reflect power rather than validity. Minority reservations must remain visible even when the authorized decision proceeds.
15.5 Reservation gaming
Participants may flood review with low-value reservations, avoid material challenges, relabel severity, or use waivers as routine bypass. Systems should report reservation concentration, waiver rate, repeat waivers, reviewer conflicts, and downstream error survival.
15.6 Evidence laundering
One inference can be repeated through agents and documents until it appears independent. Provenance and correlation links should expose common sources. Agent-generated statements default to inference unless tied to an external observation mechanism.
15.7 Epistemic attacks
Attackers may inject false source records, manipulate benchmarks, craft plausible causal chains, hide assumptions in compound language, trigger alert fatigue, or suspend critical support to block work. Security requires authentication, immutable provenance, least authority, policy review, anomaly detection, and recovery—not just prompt defenses.
15.8 Privacy and erasure
Persistent scrutiny records may expose sensitive statements, dissent, or reviewer identity. Append-only history conflicts with deletion rights and confidentiality. Deployments require access controls, minimization, cryptographic or logical redaction, retention policy, and a distinction between evidentiary integrity and perpetual exposure.
16. Artifact and Evidence Status
Claims must be reported at the narrowest accurate level:
- Implemented in the product: canonical six-view LTP graph; granular entities, designations, relationships, assumptions, and assessments; proposals and ratification; hard and soft prerequisites; provenance, supersession, collaboration, retrieval, and operational projection where present.
- Implemented inheritance slice: interchange 1.2 capture and native memory collections, both export readers, generated capture reports, read-only decision/learning inspection, and atomic source retention through a new migration. Deployment of that migration is separate from the code change. Native converter alignment, full authority-preserving restoration, authenticated timing, and full profile conformance are not claimed.
- Implemented in specifications and domain mechanisms: LTP formation rules, scoped dependencies, execution-frontier behavior, RCP evidence lifecycle, and transition semantics in their documented and coded subsets.
- Implemented in the research kernel: event folding, selected invariant checks, WHY and dependency queries, graduated propagation modes, synthetic worlds, scoring, and capture-cost instruments.
- Implemented in the paired experiment: full-history, unreviewed persistent-neutral, and unreviewed persistent-LTP conditions; safety-gated scoring; raw snapshots; cost and violation reporting.
- Proposed in this paper: first-class scrutiny packets, complete CLR reservation lifecycle, governed generic and LTP-native capture treatments, risk-tier enforcement, and the full conformance profiles.
- Observed only in a development pilot: the three-world negative result reported in §2.2.
- Hypothesized: every comparative advantage in H1–H10, including consultant benefit, capture fidelity, safety, inheritance, and economic payoff.
A feature file is not an implementation. A passing invariant test is not a deployed institution. A development pilot is not confirmatory evidence. A consultant workflow described in a paper is not a treatment until roles, records, and decisions are actually observed.
17. What Would Change or Defeat the Thesis?
- Strong reconstruction wins. Full history, RAG, or summaries match safe continuation at lower total cost.
- Persistence adds no value. A persistent structure does not improve any warrant, inheritance, or operational outcome once capture cost is counted.
- Governance adds no value. Independent scrutiny and granular ratification do not reduce consequential capture errors.
- LTP adds no value. Generic independent review matches CLR/LTP review within a preregistered equivalence margin.
- LTP harms generalization. LTP-native formation overfits its topology or terminology and performs worse outside familiar domains.
- Temporal state is sufficient. Temporal or belief memory prevents stale reliance without explicit warrants and reservations.
- The ratification tax does not repay. Review cost, queueing, and abandonment exceed reconstruction and avoided-error value for the target tier.
- Consultants create authority risk without fidelity gain. Facilitated capture centralizes framing but does not improve exact warrants or safety.
- Users cannot review extraction reliably. Consequential omissions and fabricated links survive governed review at unacceptable rates.
- Propagation remains too noisy. Alerts or revalidation burden cause rational disabling or indiscriminate blocking.
- Inheritance is no better than a memo. Expert handoffs match the structured system more cheaply.
- No downstream consequence improves. The architecture improves audit only; it should be renamed and narrowed accordingly.
- The substrate is not grammar-general. Supporting a credible non-LTP method requires replacement of core primitives rather than a new profile.
- Prior art already implements the conjunction. Recast the work as replication, integration, or comparative evaluation.
These are not rhetorical caveats. They are decision rules for simplifying or abandoning the proposal.
18. Conclusion
Persistent intelligent systems need more than recall. They need a way to know why a conclusion remains licensed, who authorized that reliance, what challenges survived review, what evidence would change it, and which actions must stop when its grounds fail.
Version 1 of Reasoning Memory made warrant state explicit but treated reliable construction too much like an input-quality problem. The development pilot exposed the weakness: an unreviewed model can turn explicit structure into a multiplier of error. More graph prompting is not governance, and LTP terminology is not LTP practice.
The revised conception begins at the source. It preserves statements before interpretation, forms warrants through native LTP projections, scrutinizes entities and arrows through persistent CLR reservations, admits reasoning through risk-tiered and separated authority, tests causal commitments through prospective evidence, and connects warrant degradation to revision and execution through RCP.
LTP is neither a decorative vocabulary nor six storage silos. It is the architecture’s native discipline for moving from standard, to diagnosis, to conflict, to future consequences, to obstacles, to action and expected effect. CLR is its write protocol. RCP keeps agreement and reality distinct. The warrant substrate allows the resulting state to outlive the people and models that produced it.
The proposal remains intentionally vulnerable. Generic review may match LTP. Strong temporal memory may match explicit warrants. Consultants may cost more than reconstruction or distort framing. The full institution may improve audit without improving action. Each outcome would narrow the claim.
The commons conception makes the enduring object an inheritable reasoning position: a goal and understanding, a case and its objections, a commitment made before an outcome, and an accountable history of revision. It is useful when a successor can continue that inquiry with less loss and justified effort, while retaining the freedom to disagree. The shared format is the means of carrying that position; governed practice is how its standing is earned.
The strongest surviving research question is:
Can an LTP-native, governed institution for constructing, contesting, authorizing, testing, and revising warrants produce safer and more economical continuity than reconstructing reasoning from content—or than storing unreviewed structure?
This paper defines the treatment and the conditions under which the answer should be no.
References
A-MEM: Xu, W., Liang, Z., Mei, K., Gao, H., Tan, J., & Zhang, Y. (2025). A-MEM: Agentic Memory for LLM Agents.
A-TMA: Shi, Z., Tang, Y., & Tung, A. K. H. (2026). A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory.
Alchourrón, C. E., Gärdenfors, P., & Makinson, D. (1985). On the logic of theory change: Partial meet contraction and revision functions. Journal of Symbolic Logic, 50(2), 510–530.
Clark, T., Ciccarese, P. N., & Goble, C. A. (2014). Micropublications: A semantic model for claims, evidence, arguments and annotations in biomedical communications. Journal of Biomedical Semantics, 5, 28.
Conklin, J., & Begeman, M. L. (1988). gIBIS: A hypertext tool for exploratory policy discussion. ACM Transactions on Office Information Systems, 6(4), 303–331.
De Kleer, J. (1986). An assumption-based TMS. Artificial Intelligence, 28(2), 127–162.
Denney, E., Pai, G. J., & Habli, I. (2015). Dynamic safety cases for through-life safety assurance. In Proceedings of the 37th International Conference on Software Engineering, 587–590.
Dettmer, H. W. (2007). The Logical Thinking Process: A Systems Approach to Complex Problem Solving. ASQ Quality Press.
Doyle, J. (1979). A truth maintenance system. Artificial Intelligence, 12(3), 231–272.
Dung, P. M. (1995). On the acceptability of arguments and its fundamental role in nonmonotonic reasoning, logic programming and n-person games. Artificial Intelligence, 77(2), 321–357.
Grudin, J. (1988). Why CSCW applications fail: Problems in the design and evaluation of organizational interfaces. In Proceedings of CSCW ’88, 85–93.
Hindsight: Latimer, C., Boschi, N., Neeser, A., Bartholomew, C., Srivastava, G., Wang, X., & Ramakrishnan, N. (2025). Hindsight is 20/20: Building agent memory that retains, recalls, and reflects.
HippoRAG: Gutiérrez, B. J., Shu, Y., Gu, Y., Yasunaga, M., & Su, Y. (2024). HippoRAG: Neurobiologically inspired long-term memory for large language models.
Hu, Y., et al. (2025). Memory in the age of AI agents.
LongMemEval: Wu, D., Wang, H., Yu, W., Zhang, Y., Chang, K.-W., & Yu, D. (2024). LongMemEval: Benchmarking chat assistants on long-term interactive memory.
Mabin, V. J., & Cavana, R. Y. (2024). A framework for using Theory of Constraints thinking processes and tools to complement qualitative system dynamics modelling. System Dynamics Review, 40(4), e1768.
Maharana, A., Lee, D.-H., Tulyakov, S., Bansal, M., Barbieri, F., & Fang, Y. (2024). Evaluating very long-term conversational memory of LLM agents.
Mem0: Chhikara, P., Khant, D., Aryan, S., Singh, T., & Yadav, D. (2025). Mem0: Building production-ready AI agents with scalable long-term memory.
Ouyang, S., et al. (2025). ReasoningBank: Scaling agent self-evolving with reasoning memory.
Pollock, J. L. (1987). Defeasible reasoning. Cognitive Science, 11(4), 481–518.
Prakken, H. (2010). An abstract framework for argumentation with structured arguments. Argument & Computation, 1(2), 93–124.
Reason Commons Project. (2026a). The Reason Commons Protocol: A ratification and falsification protocol for shared reasoning systems. Version 1.0.
Reason Commons Project. (2026b). LTP Proposal, Ratification, Dependency, and Import Specification.
Rasmussen, P., Paliychuk, P., Beauvais, T., Ryan, J., & Chalef, D. (2025). Zep: A temporal knowledge graph architecture for agent memory.
REMem: Shu, Y., Jonnalagedda, S. P., Gao, X., Gutiérrez, B. J., Qi, W., Das, K., Sun, H., & Su, Y. (2026). REMem: Reasoning with episodic memory in language agents.
ReMe: Cao, Z., Deng, J., Yu, L., Zhou, W., Liu, Z., Ding, B., & Zhao, H. (2025). Remember Me, Refine Me: A dynamic procedural memory framework for experience-driven agent evolution.
Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36.
Toulmin, S. E. (1958). The Uses of Argument. Cambridge University Press.
W3C. (2013). PROV-DM: The PROV data model. W3C Recommendation.
Walsh, J. P., & Ungson, G. R. (1991). Organizational memory. Academy of Management Review, 16(1), 57–91.