SRSurendra Reddy
Twenty-five years asking one question: what should a system be allowed to do?

Musewoods / Journal

451 Degrees Standing Thesis

The What If Next Machine

A Governed Dialectical Inquiry System

Architecture and Reference Specification for a Governed Dialectical Inquiry System

Surendra Reddy106 min read

The What If Next Machine

Architecture and Reference Specification for a Governed Dialectical Inquiry System

Surendra Reddy, 451 Labs

Architecture and Detailed Design (combined artifact), Working Draft v3.0, August 2026

Version history

VersionDateChange
v1.x2025 to early 2026Interim Analyst's Edition. Constraint taxonomy, six detection tells, force scan, ownership map, six moat tests, cascade, Unguarded Flank.
v2.0June 2026Methodology and Application Playbook. Quantitative spine made fully computed: Dirichlet prior, source grading and temper exponents, redundancy correction, signal ledger, Bayes factors, floor mechanic, conditional causal tier, cross-cycle update, sensitivity sweep, calibration.
v3.0August 2026This document. Absorbs the two 451 Labs research notes on Socratic and Stoic reasoning. Introduces the Inquiry as the primary computational object, the dialectical kernel, the five-plane architecture, the coupling rule that binds argument status to the tempered likelihood exponent, the quantitative aporia test, the expected-information-gain probe selection rule, formal invariants, a conformance suite, and a full prior-art accounting. Corrects the v2.0 worked example, whose leading branch was reported as a tilt when the Bayes factor against the runner-up does not support one.

Status

This document specifies the What If Next Machine, an instrument of 451 Labs. Names appearing here that belong to the 451 canon, including Astra, Clearstory, Ana, Actra, Pola, and ADAM, are working names for theses under active validation. None of them is a shipped product. The What If Next Machine itself exists today as a human-run analyst discipline with a computed quantitative spine. The architecture specified in Parts Four through Eight is a design under construction, not a deployed system, and every claim about system behaviour in those parts should be read as a specification of intent rather than a report of observed operation.

Lineage

This artifact descends from three parents and supersedes two of them. It supersedes The What If Next Machine: Methodology and Application Playbook (v2.0, June 2026), which remains valid as the practitioner's short reference for a hand-run File but no longer stands as the architectural authority. It absorbs From Examination to Governed Action: Implementing Socratic and Stoic Reasoning in the What If Next Machine (451 Labs Research Note, August 2026) and From Examination to Governed Action: Algorithms for a Dialectical What If Next Machine (Technical Research Paper v2.0, August 2026), both of which are retired as standalone artifacts and are preserved as the philosophical and algorithmic source material for Parts Two, Six, and Seven.

It connects laterally to four bodies of 451 canon. The WIN File series is the output format this architecture produces and constrains, and the File anatomy in Part Eleven is the normative specification for every File written after this date. The 451 Degrees standing theses are the essay-register companions to Files and inherit the argument structures named in Part Eleven.

The RDS verdict grammar of Stop, Redesign, Proceed, and Accelerate is used throughout Part Four to evaluate candidate architectures and in Part Twelve to gate phases. Astra and Clearstory are the two theses under which any eventual software realization of this specification would sit, Astra as the reasoning and inquiry surface and Clearstory as the evidence-versus-testimony substrate, and the evidence object in Part Five is deliberately shaped to be compatible with Clearstory's eight-object model.

Three reciprocal cross-links should be added. The WIN File template needs the new required sections from Part Eleven. The Clearstory whitepaper needs a note that the WIN Machine evidence object is a conforming consumer of its evidence-versus-testimony distinction. The Kriyas material needs a disambiguation note, since this document uses Kriyas Wheel in the developmental-topology sense inherited from the two research notes, which is a different construct from the Kriyas revision discipline built from Strunk.

Part One: Orientation

1.1 What this document is

This is the reference specification for the What If Next Machine. It is written to be self contained, which means a reader who has never seen a WIN File should be able to work from this document alone: the theory that motivates the instrument, the prior art it stands on, the objects it manipulates, the algorithms that operate on those objects, the arithmetic worked to the last decimal, the invariants that constrain it, the tests that would prove it conformant, and the sequence in which it should be built.

It is not a File, and it is not an essay. It carries no market thesis of its own. Where market examples appear they are teaching cases with invented actors, and the arithmetic attached to them is real and checkable while the actors are not. A reader who wants to see the instrument used on a live market should read a File. A reader who wants to build the instrument should read this.

The document is long because the alternative is worse. A shorter version would have to gesture at the parts that are hardest to specify, and the discipline this instrument enforces on markets it must first enforce on itself. Every place where the specification declines to resolve something, it says so and books the question in Part Thirteen rather than papering over it.

1.2 The four names

Three names carried the whole load through v2.0. A fourth is now required, and keeping all four separate is the first discipline.

What If Next is the model. It is a way of seeing a market. The claim underneath it is that the hidden signal in any industry is not the technology everyone is discussing, but the constraint that technology is exposing, the one binding limitation the whole system is organizing around whether anyone has named it or not.

The What If Next Machine, or the Machine, is the engine. It converts a trigger, a deal, a thesis, or a provocative public claim into a structured, falsifiable, numerically grounded view of where a market is headed and why.

The What If File, or a File, is the published artifact. It is a living dossier, not a one-time essay. The constraint thesis at the top holds across many cycles. The branch weights, the signal ledger, and the kill conditions move every cycle as evidence arrives. A File is maintained the way a position is maintained, not published the way an article is.

The Inquiry is the new name and the primary computational object. It is the complete internal state from which a File is rendered: every question asked, every definition fixed, every claim proposed, every assumption exposed, every commitment accepted, every argument built, every attack recorded, every contradiction found, every aporia held open, every belief state assigned, every agency classification made, every option filtered, every experiment authorized, and every observation returned. A File is a view over an Inquiry, the way a published balance sheet is a view over a ledger. The Inquiry persists longer than any single File version, survives changes of analyst and model, and is the thing that carries lineage.

The distinction matters in plain speech. "The What If Next view of enterprise AI" describes a way of seeing. "I ran the Aisera acquisition through the Machine" describes an act. "File No. 9 is due for its quarterly pass" describes a published artifact under maintenance. "The Inquiry behind File No. 9 holds two open aporias" describes internal state that may or may not surface in the published File. Confusing these four is how a single computed weight gets treated as a permanent truth rather than the current rendering of a living estimate.

1.3 The constraint this architecture answers

Everything downstream in this document is tested against one constraint, stated here so it can be quoted back.

The binding constraint on machine-assisted strategic judgment is not the quality of generated reasoning. It is the absence of an arbiter. Generative systems produce fluent claims, plausible counterarguments, and confident numbers at effectively zero marginal cost, which means the scarce resource is no longer the argument but the warrant to believe it. Any system that lets the generator also adjudicate has no arbiter, and its output is indistinguishable from persuasive text.

This is an architectural constraint, not a technical one on the clock table in Section 2.3, which means it does not relieve when models improve. A better generator produces better-sounding claims. It does not, by itself, produce a record of why a claim was accepted, what it depended on, which alternatives were examined and defeated, what would change the answer, and who is accountable if it is wrong. The relief mechanism is the construction of a separate adjudicating layer that the generator cannot write to directly.

The design test that follows from the constraint is the standing test for every decision in this document. For each component: does this component adjudicate, or does it generate? If it adjudicates, is it deterministic or symbolic enough to be audited, and does it record its derivation? If it generates, is it structurally prevented from writing authoritative state? A component that does both is a defect.

1.4 What changed from v2.0

The v2.0 playbook was a belief-updating instrument with an analyst attached. Its quantitative spine was rigorous and remains almost entirely intact in Part Eight. Its weakness was located precisely where the two research notes said it was: the spine assumed that the signal ledger arrives already well formed, that the branch set is exhaustive and correctly individuated, that likelihood judgments are honest, and that the question being weighted is the right question. Each of those assumptions is an unexamined input to a numerically careful process, which is the classic failure shape of quantitative analysis.

Seven changes follow. The Inquiry becomes the primary object and the File becomes a view. The Socratic operators run before the ledger, so that a signal must survive definition, assumption exposure, and impression separation before it earns a row. The Stoic operators run after the belief layer, so that a posterior cannot authorize an action it does not license. The argument graph and the branch simplex are formally coupled by a single rule in Section 8.4, rather than sitting beside each other as two unrelated formalisms. Aporia becomes a computed state with a numerical trigger, not a mood. Probe selection becomes an expected-information-gain calculation rather than an analyst's hunch. And the whole thing acquires invariants, a conformance suite, and a build sequence.

One consequence is worth flagging early because it is uncomfortable. Applying the new aporia test to v2.0's own worked example shows that its reported conclusion was wrong. Section 8.8 works this through. An instrument that cannot correct its own published examples is not an instrument.

Part Two: The theory the architecture implements

Most market analysis follows visible motion: funding rounds, product launches, executive claims, adoption curves, and the language a category adopts. Visible motion shows that something is happening. It rarely explains why the system is under pressure in that particular place and not somewhere else.

A market does not reorganize because everything about it is weak. It reorganizes because one or two constraints dominate the system at a given time. This is the theory of constraints applied to market structure, following Goldratt and Cox (1984): relieve anything other than the binding constraint and the system looks better in the spot you touched while the whole barely moves.

The Machine's job is to find the binding constraint before it becomes obvious, to test who owns it and how, and to map how power relocates when it eventually relieves. The branches in every File represent the plausible ways that relocation plays out, and the branch weights represent how much the current evidence licenses belief in each.

2.2 The constraint taxonomy

Naming the type of constraint is how you answer when, not only where.

TypeWhat bindsRelief mechanism
TechnicalA capability does not yet work at the needed qualityA capability ships at the needed quality
EconomicUnit cost does not clear the buyer's payback thresholdUnit cost crosses the threshold
RegulatoryA ruling, statute, or enforcement posture forecloses the moveA ruling or enforcement action lands
BehavioralTrust has not accrued for the buyer to actTrust accrues through repeated safe use
FinancialThe buyer cannot fund the change inside the current balance sheetFinancing structure changes, or the ask shrinks
OperationalThe surrounding process cannot absorb the capability without re-engineeringThe process gets re-engineered around the capability
CulturalThe organization's identity or incentives resist the changeIncentives realign, or a champion forces the issue
ArchitecturalA required layer or standard does not yet existThe missing layer or standard gets built

Constraints blend. A File names the dominant type and notes any secondary type with its own clock. An antitrust posture sits as a regulatory constraint riding on top of an otherwise economic one. The blend, not just the primary type, determines the realistic timeline, and a File that names only the primary type will be systematically early.

2.3 The type-to-clock table

Where without when is the difference between being early and being right.

TypeClockCharacter
TechnicalEngineeringFast once feasible
EconomicCost-curvePredictable, watch the curve
RegulatoryLitigation and rulemakingSlow and lumpy
BehavioralAdoption-curveGradual, then sudden
FinancialBudget-cycleStepwise, tied to fiscal calendars
OperationalProcess-redesignSlow, organization-dependent
CulturalIncentive-realignmentSlowest, resists forcing
ArchitecturalStandards-emergenceLumpy, then fast once a standard tips

2.4 Why examination must precede weighting

A Bayesian update is a conditional statement. It says that given this hypothesis space, these likelihood judgments, and this prior, the posterior follows. Every one of those givens is a claim about the world that can itself be wrong, and none of them is tested by the update. The arithmetic will run cleanly on a malformed question, an incomplete branch set, and a ledger full of interpretation dressed as observation. It will produce a number, and the number will look exactly as authoritative as a good one.

This is the general problem that Heuer (1999) addressed for intelligence analysis, that Dewar (2002) addressed for planning, and that the analytic tradecraft standards of the United States intelligence community codified as the requirement to distinguish underlying information from assumptions and judgments (Office of the Director of National Intelligence, 2015). It is not a new problem. What is new is the marginal cost of generating plausible content, which has collapsed to near zero and has thereby moved the bottleneck from producing analysis to warranting it.

The Socratic tradition supplies the discipline that operates on the givens rather than on the conclusion. In Plato's shorter ethical dialogues the examination proceeds from what the interlocutor has actually accepted, not from a theory the examiner imports, and the productive outcome is frequently aporia, the recognition that what looked like knowledge was not yet knowledge (Vlastos, 1983). Translated into engineering, the requirement is that the machine hold unjustified certainty, not uncertainty, as the quantity to be minimized. A conventional generative system tries to minimize residual uncertainty by filling conceptual vacancies with plausible text. This architecture does the opposite, and makes productive uncertainty a first-class, persisted state.

2.5 The Socratic function

The elenchus, the examination or putting to the test, has a recognizable computational shape. Begin with a thesis. Elicit further propositions the holder accepts. Show that the combination cannot all be maintained. What follows is not that the negation is true. What follows is that the thesis is inadequately supported, which is a different and more useful machine state.

That ambiguity, long debated in the scholarship on whether elenchus establishes positive knowledge or only refutes unjustified claims, is an asset for this architecture rather than a problem. The machine must be structurally prohibited from inferring that because A has failed, B must hold. The correct state after a successful refutation is: A is inadequately supported, no alternative has yet earned support, and the inquiry continues.

Four Socratic moves become operators, specified in Part Six. Definition seeking pushes a term from examples toward the property that makes instances instances, and gates a concept from becoming a stable node until its operational meaning is precise enough for the inquiry at hand. Assumption extraction converts a claim into the set of premises it silently requires, which is the same move Dewar (2002) called identifying load-bearing assumptions. Counterexample search attacks generalizations at their boundaries. Contradiction discovery runs the elenchus proper across four classes of conflict, and is required to emit the derivation, not merely the verdict.

2.6 The Stoic function

Socratic examination has no natural terminus. Every definition can be interrogated, every assumption challenged, every resolution reopened. That is philosophically productive and operationally fatal, because a venture instrument must eventually decide whether to invest, build, partner, wait, or run an experiment. The Stoic tradition supplies exactly the missing half, and it does so with a structure that maps unusually cleanly onto machine states.

The first contribution is the distinction between receiving an impression and assenting to it. An impression presents something to the mind; assent treats its propositional content as true; the two are separate acts, and the discipline consists in examining impressions rather than being carried by them. In an evidence-ingestion layer this becomes a hard decomposition. A headline is not a belief. "Microsoft launches autonomous procurement agents, threatening traditional procurement vendors" contains an observation, an interpretation, and a judgment, and only the first is directly evidentiary. Every input is split before anything is weighted.

The second contribution is agency. The popular reduction of Stoicism to controlling what you can control seriously underspecifies it and, worse, misapplies it: strategic actors rarely control markets, customers, competitors, governments, or technology, so a binary control distinction classifies almost everything as external and yields fatalism. Epictetus's actual emphasis falls on prohairesis, the faculty of choice, and on the correct use of impressions, which supports a four-way classification rather than a two-way one. This architecture uses govern, influence, observe, and absorb.

The third contribution is that correct action is not reducible to achieving an external outcome. Ancient Stoicism places virtue at the centre of the good rather than treating externals as the good, and while an enterprise system cannot import ancient virtue ethics wholesale, the structural lesson transfers exactly: filter the option set for permissibility before optimizing it for expected value. That gives the constitutional gate in Section 7.5 and the formal separation between belief authority and action authority in Section 10.6.

2.7 The division of labour

The two disciplines attack different failure modes, which is why the architecture needs both and why neither is decorative.

Socratic examination protects against premature answers, undefined concepts, hidden assumptions, contradictory beliefs, authority masquerading as knowledge, confirmation bias, false certainty, and poorly framed problems. Stoic judgment protects against reacting to impressions, confusing events with judgments, spending effort on outcomes outside agency, optimizing outcomes without governing conduct, strategy driven by fear or unexamined desire, and the failure to convert knowledge into appropriate action.

Stated as a single line: Socrates prevents the Machine from believing too quickly, and Stoicism prevents it from acting foolishly after it has reasoned. Neither is implemented as a conversational persona. Two language models role-playing philosophers is theatre, and Section 4.5 disqualifies that architecture explicitly. Both are implemented as operators over persisted state, which is what makes them auditable.

Part Three: Prior art

This part accounts honestly for what already exists before Part Four claims anything. The instrument stands on six mature literatures and one immature one, and most of what looks novel in the two research notes has a formal ancestor that is stronger than the informal version. Reading the ancestors closely is how the architecture avoids reinventing weaker forms of solved problems.

3.1 Formal argumentation

Toulmin (1958) broke the monolithic argument into claim, data, warrant, backing, qualifier, and rebuttal, establishing that the inference step is a separate object from the premises and can be attacked separately. That single distinction does more work in this architecture than any other borrowed idea.

Dung (1995) supplied the formal core. An abstract argumentation framework is a pair consisting of a set of arguments and an attack relation over them, and acceptability is determined by the structure of attack and defence rather than by any confidence value attached to an argument in isolation. A set is conflict-free if no member attacks another, admissible if it is conflict-free and defends all its members, complete if it contains every argument it defends, preferred if it is a maximal admissible set, and grounded if it is the least complete extension. Grounded semantics is the sceptical choice and is the correct conservative default for this architecture, because it accepts only what is defended without ambiguity and leaves everything else undecided.

Structured argumentation reconnects abstract arguments to their internal parts. ASPIC+ (Modgil & Prakken, 2014; Prakken, 2010) builds arguments from strict and defeasible inference rules and identifies three distinct ways to attack: undermining an ordinary premise, rebutting the conclusion of a defeasible inference, and undercutting the defeasible inference step itself. Preferences over arguments then convert attacks into defeats. The undercut is the move most often missing from informal analysis and most valuable here, because in market reasoning the usual error sits in an inference that quietly assumed a mechanism no longer holding, while every premise remains true.

Carneades (Gordon, Prakken, & Walton, 2007) adds proof standards and burden of proof, and models the critical questions attached to argumentation schemes as additional premises of three kinds, ordinary premises, presumptions, and exceptions, so that different critical questions shift the burden differently. Walton, Reed, and Macagno (2008) catalogue the schemes themselves with their attendant critical questions. For an instrument whose signals are frequently arguments from expert opinion, from sign, from correlation to cause, and from analogy, the scheme catalogue is effectively a library of pre-built counterexample generators.

Quantitative and probabilistic extensions matter for the coupling problem in Part Eight. Bipolar and gradual semantics assign degrees of acceptability rather than labels (Baroni, Caminada, & Giacomin, 2011). Hunter and Thimm (2017) distinguish two probabilistic readings: the constellation approach, in which probability expresses uncertainty about which framework obtains, and the epistemic approach, in which probability expresses degree of belief in an argument. The architecture in this document uses a deliberately restricted hybrid, described in Section 8.4, and does not claim to satisfy the full rationality conditions of epistemic probabilistic argumentation.

3.2 Nonmonotonic reasoning, reason maintenance, and belief revision

Pollock (1987) established defeasible reasoning as reasoning from prima facie reasons that can be defeated, and distinguished rebutting from undercutting defeaters, a distinction ASPIC+ later formalized.

Reason maintenance is the operational ancestor of the lineage requirement. Doyle (1979) built the first domain-independent truth maintenance system, in which each belief is held for a recorded reason and dependency-directed backtracking revises beliefs when assumptions change or contradictions surface. De Kleer (1986) generalized this to the assumption-based truth maintenance system, which labels each node with the sets of assumptions under which it holds and maintains minimal inconsistent environments, allowing simultaneous reasoning about multiple conflicting contexts rather than one consistent context at a time. The ATMS is the correct mental model for an Inquiry that must hold two incompatible branches live at once without corrupting itself.

Belief revision supplies the norm for change. Alchourrón, Gärdenfors, and Makinson (1985) formalized rational belief change with an emphasis on incorporating new information while minimizing unnecessary loss of existing justified structure. The practical import is a prohibition: when a signal arrives that conflicts with the standing thesis, the machine does not rebuild the worldview, and it does not delete the superseded belief. It computes a minimum-loss revision and records what was superseded and why.

Paraconsistency supplies the tolerance. In a logic where contradiction entails everything, a single inconsistent pair renders the knowledge base useless, which is fatal for an instrument whose raw material is inconsistent market evidence. Priest (1979) and the subsequent paraconsistent literature show that inference can be constrained so that a contradiction is a local finding rather than a global catastrophe. In this architecture, contradiction means investigate, and the pair persists until explicitly resolved or superseded.

3.3 Design rationale and deliberation capture

Kunz and Rittel (1970) invented the issue-based information system to support the coordination of political decision processes, structuring discourse as issues, positions, and arguments that support or oppose positions. Rittel and Webber (1973) supplied the motivating category of wicked problems, which have no stopping rule and no correct answer. Conklin and Begeman (1988) built gIBIS, moving the notation onto a computer and demonstrating its use for capturing design rationale, including the options rejected and the reasons for rejecting them.

IBIS is the closest ancestor to the Inquiry object, and it is important to say where it stops. IBIS captures the structure of a deliberation faithfully and makes it navigable. It does not adjudicate. It has no semantics that tells you which position currently survives, no belief state, no evidence grading, and no mechanism for converting a resolved deliberation into a governed action. The architecture here is, in one honest reading, IBIS with a Dung adjudicator, a graded evidence layer, a Bayesian belief layer, and an action gate bolted on.

3.4 Structured analytic tradecraft

Heuer (1999) developed analysis of competing hypotheses as an operational countermeasure to confirmation bias. Its central inversion is the one this architecture inherits: work the full hypothesis space, evaluate every item of evidence against every hypothesis, and prefer the hypothesis with the least disconfirming evidence rather than the one with the most supporting evidence, because supporting evidence is cheap and diagnostic evidence is rare. Heuer and Pherson (2011) extended this into a catalogue of structured analytic techniques including key assumptions check, which is the tradecraft twin of the assumption extraction operator.

Intelligence Community Directive 203 (Office of the Director of National Intelligence, 2015) codifies analytic standards that read as a requirements document for this instrument: describe the quality and credibility of underlying sources, express uncertainties associated with major judgments, distinguish underlying information from assumptions and judgments, identify and assess plausible alternative hypotheses, and state how a judgment is consistent with or changed from previously published analysis. The four-action maintenance classification in Section 10.5 exists to satisfy the last of these.

The NATO Admiralty code, standardized under STANAG 2022, grades source reliability and information credibility on separate alphabetic and numeric scales. The WIN source grading in Section 8.3 is a deliberately coarsened three-level version of the same idea, and the coarsening is defensible because empirical work has found interpretation of the finer alphanumeric codes to be inconsistent across analysts.

Forecast scoring closes the loop. Brier (1950) gave the proper scoring rule that makes probabilistic judgment accountable, and the Good Judgment Project work summarized by Mellers et al. (2014) and Tetlock and Gardner (2015) demonstrated that calibration is trainable and that tracked scoring across a book of forecasts, rather than post hoc narration of individual calls, is what produces improvement.

3.5 Decision making under deep uncertainty

Scenario planning in the Shell tradition (Wack, 1985; Schoemaker, 1995) established that plural, internally consistent futures beat point forecasts when the environment is not statistically stationary. Its recurring failure, which the two research notes correctly diagnosed, is that scenarios are often produced without any determination of what the reader can actually do about them.

Dewar (2002) fixed exactly that gap with assumption-based planning. The method extracts the load-bearing assumptions of an existing plan, identifies which of them are also vulnerable, assigns each vulnerable assumption a signpost that would indicate its failure, and then prescribes shaping actions to shore up the assumption and hedging actions to survive its failure. The signpost is the direct ancestor of the kill condition in Section 10.4, and the shaping and hedging split is the direct ancestor of the agency-conditioned option generation in Section 7.6.

Lempert (2019) and the wider decision-making-under-deep-uncertainty programme generalize this into robust decision making, which uses computation to stress-test candidate strategies across many plausible futures and find those that perform acceptably across them, rather than to improve prediction. The relevant discipline for this architecture is the refusal to let the quality of a decision depend on the accuracy of a forecast.

3.6 Robust and generalized Bayesian updating

The belief layer in Part Eight is not ordinary Bayes and should not be described as such. Raising a likelihood to a fractional power produces a power posterior, and there is a substantial and well-motivated literature on why this is the right move under model misspecification.

Four results carry that literature. Bissiri, Holmes, and Walker (2016) gave the general framework showing that a coherent update of a prior belief distribution can be defined through a loss function, with the standard likelihood recovered as a special case. Grünwald and van Ommen (2017) showed that ordinary Bayesian inference can be inconsistent under mild misspecification and proposed learning the learning rate. Miller and Dunson (2019) showed that conditioning on a neighbourhood of the observed data rather than the data itself yields a coarsened posterior well approximated by tempering the likelihood. Holmes and Walker (2017) gave a method for assigning the power.

This matters for honesty about what the source-grade exponent is doing. It is not a fudge factor and it is not a novel invention. It is a per-source learning rate in a generalized Bayesian update, chosen by expert judgment rather than by SafeBayes or expected-information matching, and the architecture should say so plainly and should treat the choice as a calibration target rather than a constant of nature.

The prior uses a Dirichlet over the branch simplex, which supplies conjugate carryover across cycles, and Kass and Raftery (1995) supply the interpretive scale for Bayes factors that keeps the reporting language honest.

3.7 Causal estimation

Where a panel of resolved comparable cases exists, causal estimation is available and prediction is not a substitute for it. Wager and Athey (2018) and Athey, Tibshirani, and Wager (2019) established causal forests for estimating heterogeneous treatment effects, with an orthogonalization step drawn from double machine learning. Brodersen, Gallusser, Koehler, Remy, and Scott (2015) established Bayesian structural time-series event studies for estimating a counterfactual trajectory around a dated intervention. Lundberg and Lee (2017) supplied Shapley-value attribution for decomposing a prediction into per-feature contributions, which is explanation of a prediction and not an estimate of a causal effect, a distinction Section 8.11 enforces. Hamilton (1989) established regime-switching models, whose descendant hidden Markov formulations produce a posterior over latent regimes rather than a regime label.

3.8 The LLM era

Chain-of-thought prompting (Wei et al., 2022) and self-consistency decoding (Wang, X., et al., 2023) improved reasoning by exposing and aggregating intermediate steps. Their limitation for this architecture is that the intermediate steps are not reliably faithful to the computation that produced the answer (Turpin, Michael, Perez, & Bowman, 2023), which means a chain of thought is generated content and cannot serve as an audit record.

Debate as a mechanism has an honest and mixed record that must be stated accurately, because the two research notes leaned on it. Irving, Christiano, and Amodei (2018) proposed debate for scalable oversight. Du, Li, Torralba, Tenenbaum, and Mordatch (2024) reported factuality and reasoning gains from multi-agent debate. Khan et al. (2024) found that debate improves judge accuracy when debaters have access to information the judge lacks.

Against this, Wang, Q., Wang, Z., Su, Tong, and Song (2024) found that a single agent with strong prompts matches the best discussion approach across a wide range of reasoning tasks and backbone models, with multi-agent discussion outperforming only when the prompt lacks demonstrations. Cemri et al. (2025) analysed over 1,600 annotated traces across seven multi-agent frameworks and produced a taxonomy of fourteen failure modes in three categories, concluding that many failures are system design issues rather than model limitations and that better base models will not resolve them.

The correct reading of that literature is the one this architecture adopts: adversarial generation is useful for search diversity and is not a mechanism for adjudication. Two models arguing produce more candidate arguments. They do not produce a warrant, and a system that treats their convergence as a verdict has installed majority opinion where an arbiter should be.

Constitutional AI (Bai et al., 2022) demonstrated that principle-based evaluation can be applied as an explicit layer over model behaviour, which is the closest existing analogue to the constitutional gate. Sycophancy work (Sharma et al., 2024) documents the specific failure this instrument must resist, since a strategy machine that agrees with its principal is worse than no machine.

Uncertainty quantification has advanced materially, with semantic entropy computing uncertainty in meaning space rather than token space and detecting confabulation without ground-truth labels (Kuhn, Gal, & Farquhar, 2023; Farquhar, Kossen, Kuhn, & Gal, 2024). Argument mining surveys report strong headline performance for large models on claim and relation extraction alongside systematic weakness on long, nuanced, and emotionally charged text, and note that off-the-shelf generation is not a substitute for an explicit schema (Lawrence & Reed, 2020, for the earlier state of the art; recent LLM-era surveys concur).

Socratic framing in LLM systems is now a small literature of its own, concentrated in tutoring, where a Socratic system is one that teaches through guided questioning rather than explanation. That work is genuinely adjacent but solves a different problem: the pedagogical goal is to lead a learner to an answer the system already holds, whereas the goal here is to prevent the system from holding an answer it has not earned.

3.9 What is genuinely unclaimed

Set against that inventory, the honest differentiation is narrow and specific, and it is worth stating narrowly so it survives scrutiny.

Nothing in the prior art couples a defeasible argument graph to a tempered Bayesian belief layer over market outcomes, such that the dialectical status of an argument modulates the exponent applied to the likelihood of the evidence it supports. Probabilistic argumentation assigns probabilities to arguments; the WIN Machine needs probabilities over market branches, with argument status acting as a gate on evidentiary weight. Section 8.4 specifies that coupling, and it is the one place this architecture makes a formal claim rather than an integration claim.

Beyond that single coupling, three combinations appear unoccupied rather than novel in their parts. Aporia is not, in any prior system reviewed here, a persisted first-class asset with resolution conditions, a materiality threshold, and a numerical trigger derived from the belief layer. Structured analytic tradecraft, generalized Bayesian updating, and formal argumentation have not been assembled into one instrument with shared provenance and a single maintenance cadence. And the separation between belief authority and action authority, in which a principal may authorize an experiment on a claim the system holds as provisional without the claim's epistemic status being rewritten, is standard in governance thinking and absent from reasoning systems.

Everything else in this document is assembly. That is not a weakness. Most of what fails in practice fails at the joints.

Part Four: The architectural thesis

4.1 Five planes

The Machine is organized as five planes. Each plane answers exactly one question, and a plane that answers two is a defect to be split.

PlaneThe one question it answersNature
ImpressionWhat actually arrived, stripped of what was added to it?Deterministic decomposition over generated candidates
DialecticalWhat survives examination, and what has collapsed?Symbolic adjudication over an argument graph
BeliefGiven what survives, how much weight does each branch carry?Generalized Bayesian computation over the branch simplex
Agency and governanceWhat is ours to move, and what are we permitted to do?Rule evaluation against classification and constitution
Action and learningWhat do we do, what did reality return, and what must be revised?Experiment emission, observation capture, minimum-loss revision

The ordering is not decorative. Each plane consumes only what the plane above it has authorized, and no plane may write upstream state directly. The Belief plane cannot see raw inputs; it sees only signals that survived the Impression and Dialectical planes. The Action plane cannot see raw posteriors; it sees assent states and agency classifications derived from them.

Figure 1

Rendering diagram…

4.2 The generator and the arbiter

The single load-bearing separation is stated as a rule rather than a preference.

The language model proposes. The dialectical kernel determines what the Machine is entitled to believe, what remains contested, what has collapsed into aporia, and what may legitimately become action.

Generation is used aggressively and without embarrassment. Concept extraction, candidate definitions, assumption generation, counterexample generation, argument construction, analogy search, question reformulation, hypothesis generation, experiment ideation, and natural-language explanation are all tasks where a strong model outperforms a human analyst on throughput and frequently on coverage. None of those outputs is authoritative on arrival. Each is a candidate mutation of the Inquiry, and each must pass a typed validation before it becomes state.

The kernel determines accepted truth, whether a contradiction is resolved, whether evidence is authoritative, whether an action is constitutionally permissible, and whether a high-consequence action executes. It is deterministic and symbolic wherever the semantics permit, and where it must be probabilistic it exposes its arithmetic. Formatting a generative judgment as JSON does not make it deterministic, and the architecture must never pretend otherwise.

4.3 Candidate architectures

Five architectures were evaluated against the design test in Section 1.3. Verdicts use the RDS grammar and are scoped to a role.

Architecture A: prompt-orchestrated persona debate. Two or more model instances role-play examiner and judge, exchanging critiques until they converge, with the transcript as output. Strengths: trivial to build, immediately demonstrable, produces genuinely diverse candidate arguments. Structural disqualification: convergence is not adjudication, and the empirical record shows both that single-agent prompting matches discussion approaches across a wide range of tasks (Wang, Q., et al., 2024) and that multi-agent systems fail predominantly through design issues rather than model limitations (Cemri et al., 2025). More seriously, nothing persists, so there is no lineage, no revision record, and no way to ask on what basis a belief was held three months ago. Verdict: Proceed as a demonstration and as a generator inside Architecture D; Stop as destination.

Architecture B: pure symbolic argumentation kernel. Formalize everything in ASPIC+ with grounded semantics, extract structure from text into the formalism, and compute acceptability. Strengths: the adjudication is fully auditable and the rationality postulates are studied. Structural disqualification: market evidence does not decompose cleanly into strict and defeasible rules over a fixed language, extraction is the hard part and is exactly where large models are weakest on long, nuanced text, and the formalism produces labels rather than the weighted branch map a File requires. A capital allocator cannot act on the statement that argument A3 is IN. Verdict: Redesign, and retain grounded semantics as the adjudication layer inside Architecture D rather than as the whole system.

Architecture C: the v2.0 Machine, a pure belief layer. Analyst-curated signal ledger, source-graded temper exponents, tempered Bayesian update over branches, Bayes factors, floor, cross-cycle carryover, sensitivity sweep, calibration. Strengths: it works, it is checkable, it produces exactly the artifact a decision maker needs, and it has an operating record across a book of Files. Structural disqualification as a whole: everything upstream of the ledger is unexamined, so the arithmetic can be impeccable while the branch set is wrong, the question is malformed, and half the ledger rows are interpretation. Verdict: Proceed as the Belief plane; Stop as the whole Machine.

Architecture D: coupled dialectical kernel with a tempered Bayesian belief layer. The five planes of Section 4.1, with the Dialectical plane adjudicating what may enter the ledger and modulating how strongly it enters, and the Belief plane converting what survives into weights. Strengths: it preserves the entire quantitative spine, it repairs the exposed flank, it produces both an auditable derivation and an actionable weight, and each plane can be built and tested independently.

Costs: it is materially more work than C, it introduces a coupling rule that has no direct precedent and therefore no external validation, and it can be over-engineered into an expensive machine for avoiding decisions if the scheduler and stopping rules are weak. Verdict: Accelerate, with the coupling rule in Section 8.4 flagged as the primary technical risk and the scheduler in Section 6.9 as the primary operational risk.

Architecture E: end-to-end learned dialectical model. Train or fine-tune a model to emit the whole Inquiry state and its transitions directly. Strengths: potentially far cheaper at inference, and avoids brittle extraction. Structural disqualification for now: it collapses generator and arbiter into one artifact, which is precisely the constraint in Section 1.3, and it has no audit surface that survives the model being replaced. Verdict: Stop for this programme. Revisit trigger: a demonstrated method for verifying that a learned adjudicator's emitted derivations are faithful to its computation, at which point re-evaluate.

Architecture D is recommended, built in the increments of Part Twelve. The layers have deliberately different expected lifespans, and stating them prevents the common error of treating a scaffold as a foundation.

The Inquiry object model and its provenance rules are intended to be permanent. They are the thing that must survive every model change, every analyst change, and every revision of the belief mathematics, and a schema change here is a migration event.

The Dialectical plane's adjudication semantics are intended to be long-lived but replaceable. Grounded semantics is the conservative starting choice; moving to a gradual or preference-based semantics later is a bounded change if the graph schema holds.

The Belief plane's specific update rule is intended to be revisable every few quarters against the calibration record. The temper exponents in particular are calibration targets, not constants.

The generative layer is intended to be disposable. Every prompt, every model choice, and every extraction pipeline should be replaceable within a sprint without touching persisted state, and any design that makes a model choice load-bearing has violated the separation.

Part Five: The object model

5.1 Identity, versioning, and provenance

Every node carries a stable identifier, a monotonically increasing version, a creation timestamp, and a provenance record naming the actor that produced it, the Inquiry it belongs to, and, where the actor is generative, the model and prompt version. Nodes are never destroyed. Supersession replaces the current version and preserves the prior one with a pointer, following the reason-maintenance tradition (Doyle, 1979) and the belief revision norm of minimizing loss of justified structure (Alchourrón, Gärdenfors, & Makinson, 1985).

Provenance is modelled on the W3C PROV pattern of entities, activities, and agents, so that any node can answer three questions: what was I derived from, what activity derived me, and who or what performed that activity. This is not bureaucratic overhead. It is the only mechanism by which the Machine can answer why it believed something on a given date, which is the question that separates an instrument from a generator.

5.2 The Inquiry

An Inquiry is the full persisted state of one line of investigation. Formally it is the tuple

I = ⟨Q, D, C, A, K, E, R, G, X, U, B, Z, V, O, P, H⟩

where Q is the question set with exactly one active question, D the definitions, C the claims, A the assumptions, K the commitments, E the evidence, R the arguments, G the relation graph over arguments and claims, X the contradictions, U the aporia ledger, B the belief and assent states, Z the agency classifications, V the constitutional constraints in force, O the options, P the experiments and probes, and H the history and provenance.

A published File is a rendering of a subset of I at a point in time. Which subset is a publication decision, not an epistemic one, and Section 10.7 requires that the unpublished remainder be retrievable.

5.3 Node types

Seventeen node types are required. Each is listed with the one question it answers, which is the test for whether a proposed eighteenth type is really needed or is a field on an existing one.

NodeThe question it answers
QuestionWhat are we trying to determine?
DefinitionWhat must this term mean for the inquiry to remain coherent?
ClaimWhat is being asserted, at what scope and modality?
AssumptionWhat must be true for a claim to hold, and is it necessary or contributory?
CommitmentWhat has a participant actually accepted, at what strength, revocably or not?
EvidenceWhat was observed, from what source, with what provenance and independence?
ArgumentWhich premises, through which inference, yield which conclusion?
RelationHow do two nodes stand to each other: supports, rebuts, undercuts, undermines, depends on, contradicts, qualifies, supersedes?
AporiaWhat tension is unresolved, what blocks it, and what would resolve it?
BranchWhat is one of the mutually exclusive ways this constraint could relocate?
SignalWhich piece of admitted evidence bears on the branches, with what likelihood vector and what temper?
AgencyAssignmentFor this variable, do we govern, influence, observe, or absorb?
NormWhat is impermissible regardless of expected value?
OptionWhat could we do?
ExperimentWhat observation would most strongly discriminate between live hypotheses?
ObservationWhat did reality return, and how does it compare to what was predicted?
DispositionWhat did this cycle conclude, and under which of the four maintenance actions?

Two of these deserve emphasis. The Relation node is first class rather than an edge attribute because relations are themselves attackable: an analyst may dispute that A3 undercuts A2 without disputing either argument, and that dispute needs somewhere to live. The Signal is distinct from the Evidence because one item of evidence may bear on several Files, and its likelihood vector is File-specific while its provenance and independence are not.

5.4 Schemas

The canonical claim, showing the required separation between model confidence and dialectical status.

claim:
  id: clm_01873
  version: 3
  proposition: >
    Agent-mediated workflows will reduce direct human use of enterprise
    application interfaces.
  epistemic_type: forecast          # fact | interpretation | assumption | hypothesis | forecast | norm
  scope:
    domain: enterprise-software
    population: large-enterprises
    horizon: 2026-2030
  modality: predictive
  polarity: positive
  generator_confidence: 0.85        # advisory only, never authoritative
  dialectical_status: contested     # IN | OUT | UNDECIDED, from grounded labelling
  assent: suspended                 # accept | provisional | suspend | contested | reject
  support: [arg_0221, arg_0240]
  attacks: [arg_0233]
  assumptions: [asm_0102, asm_0103]
  confidence_vector:
    evidentiary_strength: 0.42
    argument_acceptability: 0.55
    defeat_exposure: 0.61
    source_diversity: 0.30
    temporal_freshness: 0.88
    unresolved_assumption_burden: 0.70
  provenance:
    inquiry_id: inq_142
    created_at: 2026-08-11T14:22:07Z
    generated_by: analyst-agent/claude-opus-5
    prompt_version: define-v4
  supersedes: clm_01873@2

The evidence node, shaped to be compatible with the Clearstory evidence-versus-testimony distinction. Testimony about a fact is not the fact, and the schema forces the caller to say which it holds.

evidence:
  id: ev_0554
  kind: observation                 # observation | testimony
  statement: >
    Incumbent reported 41.2 million dollars of revenue attributed to the
    contested product line in the quarter ended 30 June 2026.
  source:
    citation: "Form 10-Q, filed 2026-08-04"
    grade: A                        # A | B | C
    motive_exposure: none           # none | vendor-adjacent | self-interested
  independence:
    cluster_id: clu_0071            # signals sharing a cluster are correlated
    primary: true                   # false means it restates a cluster primary
  temporal:
    observed_at: 2026-06-30
    decay_class: corporate-strategy
  provenance:
    ingested_by: impression-pipeline/v2
    decomposed_from: imp_0919

The argument node, with the inference step exposed as an attackable object in its own right.

argument:
  id: arg_0233
  premises: [ev_0554, asm_0102]
  inference:
    id: inf_0233
    scheme: argument-from-correlation-to-cause
    critical_questions_open: [cq2_confounder, cq4_reverse-causation]
  conclusion: clm_01873
  attackable_at:
    premises: true
    inference: true
    conclusion: true
  status: UNDECIDED

The aporia node, which persists across cycles and may never be silently deleted.

aporia:
  id: apo_0021
  opened_at: 2026-07-02
  trigger: mutually_undefeated_arguments
  question: Does agent mediation destroy or deepen incumbent platform economics?
  tension:
    - interface engagement may fall
    - transaction volume and policy dependence may rise
  blocking_unknowns:
    - future pricing architecture
    - agent transaction intensity per seat displaced
  materiality: 0.71                 # share of branch mass the tension controls
  resolution_conditions:
    - a pricing-model disclosure from any top-five incumbent
    - telemetry from a two-platform delegation pilot
  status: open                      # open | resolved | reframed | superseded | accepted_as_irreducible

5.5 The six epistemic types and the assent lattice

Propositions are typed before they are reasoned over, because applying one generic verification move to all of them is a category error. A fact claims an observed or verifiable state. An interpretation assigns meaning to facts. An assumption is accepted temporarily to permit reasoning. A hypothesis is deliberately exposed to falsification. A forecast concerns a future state. A norm says what should or should not be done. Facts are checked against sources, interpretations against alternative readings, assumptions against necessity, hypotheses against discriminating experiments, forecasts against base rates and calibration, and norms against the constitution.

Assent is separate from type and separate from dialectical status. Five values are used: accept, provisional, suspend, contested, and reject. Suspend is the most important and the one generative systems structurally lack, because it is the state of having examined and declined to conclude. The decision rule is given in Section 7.2.

The three states must never be collapsed into a single score. A claim can be typed as a forecast, labelled UNDECIDED by the grounded semantics, and held at provisional assent because a principal has authorized an experiment on it. Those three facts say different things, and a single number destroys all of them.

5.6 Lifecycles

The claim lifecycle runs proposed, typed, argued, examined, then one of accepted, contested, suspended, or defeated, and finally either superseded or resolved. Transitions are events, and every transition emits an event per Section 10.7.

The aporia lifecycle runs open, then either resolved, reframed, superseded, or accepted as irreducible. There is no delete transition. An aporia that stops being interesting is accepted as irreducible with a reason, which preserves the record that the Machine once could not answer something.

The File lifecycle runs opened, active, and then per cycle receives exactly one of the four maintenance dispositions in Section 10.5, terminating in retirement when the constraint has fully resolved or the thesis has been structurally falsified. A File may be active while holding open aporias. That is a normal operating state and not a defect.

Part Six: The Socratic operators

Each operator is specified with its trigger, its contract, its algorithm, and its refusal condition. The refusal condition matters as much as the algorithm: an operator that can never decline to act is a rubber stamp. All operators are epistemically negative in authority, meaning they may challenge, expose, distinguish, and mark, but none of them may mark a claim as accepted. Acceptance is computed by the Dialectical plane from the graph, never asserted by an operator.

6.1 DEFINE

Trigger. A material concept appears in the active question or in a proposed claim and its ambiguity exceeds threshold.

Ambiguity is estimated as a weighted sum of three components: polysemy risk, meaning the concept has multiple established senses in the domain; boundary uncertainty, meaning near-miss cases cannot be reliably sorted; and participant disagreement, meaning two humans or two independently seeded generators produce materially different senses. The third component is the most diagnostic and the cheapest to measure, since it requires only sampling the generator twice with different seeds and testing whether the resulting definitions classify the same counterexamples the same way.

DEFINE(term, inquiry):
    senses        <- generate_candidate_senses(term, n>=3, independent seeds)
    examples      <- retrieve_or_generate_examples(term)
    near_misses   <- generate_boundary_cases(term)

    for sense in senses:
        properties <- extract_defining_properties(sense)
        classify(examples, properties)
        classify(near_misses, properties)

    if two senses classify any near_miss differently:
        mark term AMBIGUOUS
        require scoped definition from analyst or infer from commitments

    node <- DefinitionNode(
        description, necessary_properties, sufficient_properties,
        included, excluded, open_boundaries, status = provisional)
    return node

Refusal condition. DEFINE refuses to emit a definition when the surviving senses would produce materially different branch sets. That refusal is itself a finding: it means the question is asking about two different things at once and must be split before anything is weighted.

Worked case. The question "what if enterprises become agentic" fails the gate. Autonomous execution, goal-directed software, delegated authority, persistent agents, machine-to-machine transaction, and self-modifying workflow are six different senses, and an enterprise containing agents is not necessarily an agentic enterprise under any of them. The provisional definition that survives, an enterprise in which software actors perceive state, reason over goals, initiate actions, and coordinate with other actors under delegated authority without requiring a human decision at every execution boundary, explicitly excludes conventional workflow automation, chat interfaces without autonomous action, and deterministic scripts labelled as agents. That definition is a node open to further elenchus, not a glossary entry.

6.2 EXPOSE

Trigger. Any claim entering the Inquiry.

EXPOSE converts a claim into the set of premises it silently requires, then classifies each by kind and by dependency strength. The kinds are logical, causal, empirical, economic, behavioural, temporal, and normative, and they matter because they determine which challenge applies. The dependency strength is binary: an assumption is necessary if the claim fails without it, and contributory otherwise. This is Dewar's load-bearing test (2002), and the necessary set is the set that gets signposts in Section 10.4.

EXPOSE(claim):
    premises <- infer_minimum_premises(claim)
    for p in premises:
        p.kind     <- classify(logical|causal|empirical|economic|
                               behavioural|temporal|normative)
        p.strength <- NECESSARY if negate(p) defeats claim else CONTRIBUTORY
        p.vulnerable <- true if p could plausibly fail within the clock horizon
        emit AssumptionNode(p)
    return assumption_set

Refusal condition. EXPOSE refuses to return an empty set. A claim with no exposed assumptions is either a tautology or an extraction failure, and both require the analyst.

Worked case. "Autonomous agents will reduce enterprise software spending" carries at least five: that agents substitute rather than complement, that pricing remains seat-based, that vendors cannot capture agent interactions, that adoption reaches material scale, and that agent operating cost is lower than the displaced expenditure. Four of the five are necessary and vulnerable. The second, seat-based pricing, is the one that a transaction-priced incumbent falsifies immediately, which is why it is the assumption a competent analyst attacks first.

6.3 ELICIT

Trigger. A participant, human or generative, asserts something that later reasoning will lean on.

The commitment ledger records what has actually been accepted, by whom, at what strength, in what scope, at what time, and whether revocably. Strength runs asserted, accepted, strongly accepted, and constitutionally committed. Acceptance never implies permanence, and a revoked commitment is superseded rather than deleted so that the record of having held it survives.

The reason this operator exists is that Socratic examination proceeds from the interlocutor's own commitments rather than from a worldview the examiner imports. Without a ledger the Machine substitutes its own priors for the principal's and calls the result an examination.

6.4 ELENCHUS

Trigger. Any claim whose acceptance would materially move a branch weight.

Four classes of contradiction are detected and they are not equally strong. A logical contradiction is a formal inconsistency. A semantic contradiction arises when two propositions use incompatible definitions of the same term, which is why DEFINE runs first. A causal contradiction arises when a proposed mechanism conflicts with another accepted mechanism. A pragmatic contradiction arises when a proposed strategy conflicts with a standing commitment even though both are logically possible, which is the most common class in strategic work and the one that generates the most valuable aporias.

ELENCHUS(target_claim, inquiry):
    K   <- relevant_commitments(target_claim)
    A   <- dependencies(target_claim)
    D   <- definitions_in_scope(target_claim)

    candidates  = symbolic_consistency_check(target_claim, K)
    candidates += semantic_conflict_check(D)
    candidates += causal_conflict_search(target_claim, inquiry.G)
    candidates += pragmatic_conflict_search(target_claim, K, inquiry.V)

    for conflict in candidates:
        conflict.derivation <- explicit_chain(conflict)   # REQUIRED
        conflict.severity   <- branch_mass_at_stake(conflict)
    return rank(candidates, by = severity)

Refusal condition, and it is the strictest in the specification. ELENCHUS refuses to emit a contradiction without a derivation. It may not report that two things seem inconsistent. It must produce the chain, in the form: C1 depends on A4; A4 implies X; K7 implies not X; therefore C1 and K7 cannot both survive unchanged. A contradiction without a derivation is an opinion, and the kernel discards it.

6.5 COUNTEREXAMPLE

Trigger. Any universally quantified or unbounded generalization.

Ten generation strategies are run: edge case, opposite market condition, different geography, different customer type, temporal inversion, scale change, adversarial actor, regulatory intervention, resource scarcity, and second-order consequence. Where the argument instantiates a known scheme, the scheme's critical questions supply additional targeted attacks, which is why the scheme catalogue (Walton, Reed, & Macagno, 2008) is worth carrying as a resource rather than regenerating attacks from scratch each time.

The selection criterion is specified precisely because the intuitive one is wrong. The winning counterexample is not the cleverest, the most surprising, or the one the generator ranks highest. It is the one that forces the largest revision to the proposition, measured as the reduction in scope required for the claim to survive it. A counterexample that removes one exotic case from a claim's scope is nearly worthless; one that removes the population the File actually cares about is the finding.

6.6 DIALECTICAL_EXPAND

Trigger. A claim that has survived initial examination and is about to influence branch weights.

Three products are required and the third is the one usually missing. Generate the strongest case for, the strongest case against, and at least one non-binary reframing. Binary expansion manufactures the false dichotomy that most malformed strategic questions already contain, and the alternative branch is where the actual answer usually lives.

The instruction to generators is to steelman, not to criticize. Each side is made as strong as it can be made before any evaluation occurs, and generators producing opposing cases should not see one another's output on the first pass, because the point is independence of cognitive search rather than the appearance of disagreement. This is the one place where the multi-agent literature's positive findings apply cleanly: adversarial generation improves coverage of the argument space (Du et al., 2024; Khan et al., 2024) even though it does not adjudicate (Wang, Q., et al., 2024; Cemri et al., 2025).

6.7 DETECT_APORIA

Trigger. End of every examination pass, before the belief update is published.

Five structural triggers open an aporia. Contradictory commitments survive and neither can responsibly be rejected. Multiple materially incompatible definitions survive. Competing arguments remain mutually undefeated, meaning both are labelled IN under the chosen semantics while their conclusions are incompatible. A critical unknown blocks inference, meaning the thesis depends on an assumption for which sufficient evidence does not exist and cannot currently be obtained. The question is malformed, meaning analysis reveals a false binary or a category error in the framing.

A sixth trigger is quantitative and is new in v3.0. It fires from the belief layer and is specified with its arithmetic in Section 8.8: when the two leading branches are separated by a Bayes factor below three, both clear the floor, and together they hold the majority of the branch mass, the File is in aporia regardless of which branch nominally leads.

Materiality gates whether a detected tension becomes a node. Materiality is the share of branch mass the tension controls, and the working threshold is 0.15. Below that, the tension is logged in the ledger notes and not promoted, because an aporia ledger that accumulates every trivial tension becomes noise and stops being read.

6.8 REFORMULATE

Trigger. An aporia of the malformed-question class, or a question quality score that a candidate reformulation beats.

Question quality is scored on five weighted components: clarity, meaning ambiguity has been reduced; testability, meaning some observation could bear on it; decision relevance, meaning at least one actor's action changes with the answer; non-bias, meaning it does not presuppose its own answer or impose a false binary; and agency, meaning the answer touches something someone can govern or influence. A reformulation replaces the active question only when it scores strictly higher, and the replacement is recorded as an event so that the lineage from the original question survives.

The canonical worked case is the SaaS question. "What if AI eliminates SaaS" is malformed because the inquiry establishes that enterprises still require systems of record, that agents need authoritative data and action interfaces, and that many workflows span multiple systems, which together contradict elimination while leaving the interesting question untouched. The reformulation, "if agents increasingly mediate enterprise work, where do control of workflow, context, execution authority, transaction state, and economic rent migrate," scores higher on every component and produces a completely different and far better branch set. That improvement in the question, not any answer, was the cycle's product.

6.9 The scheduler and the stopping rules

Socratic examination has no natural terminus, so the architecture supplies one. Two mechanisms do the work, and they are the primary defence against building an expensive machine for avoiding decisions.

The scheduler ranks candidate investigations by a priority function combining materiality, uncertainty, decision impact, and resolvability in the numerator against cost in the denominator. In plain terms: investigate the uncertainties that matter most, that could actually change a decision, and that are reasonably resolvable, and stop dignifying the ones that are none of those things. A high-uncertainty, low-decision-impact question is philosophically interesting and operationally worthless, and the scheduler must be willing to say so.

Five stopping conditions terminate a reasoning cycle. Sufficient justification: a claim has survived meaningful attack and meets its evidence threshold. Decision sufficiency: uncertainty remains but further inquiry is unlikely to change the decision. Aporia: a critical contradiction cannot presently be resolved and is booked. Experiment required: further progress needs interaction with the world. Resource bound: the expected epistemic gain of the next step falls below its cost, which is the value-of-information criterion (Howard, 1966) and is computed rather than felt, using the arithmetic in Section 8.14.

Part Seven: The Stoic operators

7.1 IMPRESSION

Trigger. Every external input, without exception, including inputs the analyst is confident about.

Each input is decomposed into atomic statements, and each statement is classified as observation, attribution, interpretation, forecast, or recommendation. Only observations and attributions are directly evidentiary. Interpretations and forecasts enter the Inquiry as claims with assent set to suspended, and they may earn assent later through argument, but they may never enter the signal ledger as though they were observations.

IMPRESSION(input):
    statements <- decompose(input)
    for s in statements:
        s.class <- classify(observation | attribution | interpretation
                            | forecast | recommendation)
        s.source_grade <- grade(source)          # A | B | C
        s.motive_exposure <- assess(source)
        s.cluster_id <- assign_provenance_cluster(s)
        if s.class in {interpretation, forecast, recommendation}:
            s.assent <- SUSPENDED
            emit ClaimNode(s)
        else:
            emit EvidenceNode(s)
    return ImpressionBundle

This operator is where source hygiene lives, and it does three jobs that were previously scattered. It grades the source. It assigns the provenance cluster used by the redundancy correction in Section 8.5, so that ten articles restating one press release are recognized as one observation with nine echoes. And it separates the observation from the judgment attached to it, which is the single highest-yield discipline in the whole specification because it removes the most common route by which a narrative becomes a number.

Worked case. The input "agentic AI spending will reach a given figure by 2028" is not a belief and not evidence. The observation is that a named firm published a figure using a stated methodology on a stated date. Whether that figure measured realized spending or purchase intention, what the firm counted as agentic, whether the sample was representative, whether another dataset corroborates it, and whether its stated assumptions still hold are all open questions. Only after those are settled does anything enter the ledger, and what enters is the observation, with a likelihood vector reflecting what the making of that forecast implies about the branches, not the forecast's content taken as fact.

7.2 ASSENT

Trigger. After the argument graph is labelled and the branch posterior is computed.

Assent is a function of support, attack, source quality, coherence with commitments, scope match, recency, and residual uncertainty. The decision rule is ordered, and the ordering matters because the first matching condition wins.

ASSENT(claim):
    if decisive_contradiction_exists(claim):        return CONTESTED
    if undefeated_strong_argument(claim)
       and evidence_threshold_met(claim):           return ACCEPT
    if reasonable_support(claim)
       and necessary_assumptions_unresolved(claim): return PROVISIONAL
    if evidence_insufficient(claim):                return SUSPEND
    if strongest_arguments_defeated(claim):         return REJECT

The significance of SUSPEND is worth stating plainly. It is the computational form of refusing premature assent, it is the state most generative systems cannot represent, and a Machine that never returns it is not examining anything. A File in which no material claim is suspended should be treated as evidence of a shallow pass rather than a clean one.

7.3 PAUSE

Trigger. Any strategically significant event, meaning any input that would otherwise trigger an action proposal.

The pause inserts an explicit judgment boundary between stimulus and reaction. It asks what actually occurred, what interpretation has been added, what impulse follows from that interpretation, what the assent status of the underlying claims is, what lies within agency, and what action if any is warranted. Where material claims remain suspended, the pause blocks irreversible action, which is Invariant 5 in Section 10.1.

The value is easiest to see in the competitor-launch case. Without the pause the chain is threat detected, therefore respond. With it: the competitor launched a product, which is observed; the claim that this destroys our advantage is an interpretation, not an observation; the evidence for that interpretation is currently incomplete; customer switching behaviour is unknown; our pricing, differentiation, and partner strategy are governable; customer response is influenceable but not controllable; therefore run three discriminating experiments before changing strategy. The pause converts an anxiety into a judgment and a budget.

7.4 CLASSIFY_AGENCY

Trigger. Every driver, variable, or uncertainty that appears in a branch or a cascade link.

Four classes replace the binary control distinction. Govern means directly choosable, such as product architecture, capital allocation, or experiment design. Influence means outside direct control yet responsive to intervention, such as partner adoption, standards participation, or customer behaviour. Observe means informative but not meaningfully influenceable, such as competitor earnings or macroeconomic conditions. Absorb means the appropriate response is resilience and optionality rather than intervention, such as regime-level shocks.

The classifier scores authority, causal leverage, resources, access, and time. High authority with high leverage yields govern. Low authority with meaningful leverage yields influence. Minimal leverage with observability yields observe. Minimal leverage with unavoidability yields absorb. The scoring should be conservative, because the characteristic failure is optimism: firms routinely classify as influenceable what they merely wish to influence.

Agency classification is what prevents fantasy strategy. If the Machine determines that foundation model token prices will fall sharply, the next question is not what we should do to make prices fall, which is outside agency for all but a handful of actors. It is to design product economics robust across multiple price curves, which is inside it. Strategy becomes action on the governable plus adaptation to the exogenous, rather than a wish addressed to the world.

7.5 CONSTITUTIONAL_FILTER

Trigger. Before any option is scored for expected value.

A constitution is a set of normative constraints that are evaluated before optimization, not weighed inside it. The permissible option set is the subset satisfying every constraint, and the selected action is the expected-value maximizer within that set rather than over the whole set. The difference between those two formulations is the entire difference between an optimizer and a governed machine.

A working constitution for a 451 Labs venture would encode truthfulness, non-manipulation, preservation of human agency and recourse, fairness, regenerative rather than extractive value, security, attributable authority for high-consequence autonomous action, and long-term stewardship. Two constraints are load-bearing for this instrument specifically and should be written as hard rules rather than principles: high-consequence autonomous action requires attributable authority, and no irreversible action may be derived from an epistemically suspended claim.

The filter's relationship to the ADAM canon should be stated so the two are not confused. In ADAM, Pola is the policy decision point and the Governor is the enforcement point holding no write path. The constitutional filter here is a policy decision function local to the Inquiry, and if a WIN Machine realization is ever built inside ADAM, the filter should delegate to Pola rather than reimplement it, with the Governor enforcing. Nothing in this document should be read as duplicating either.

7.6 GENERATE_OPTIONS and SELECT

Trigger. After assent, agency, and constitutional filtering.

Options are generated conditioned on agency class, which is what makes them actionable rather than aspirational. Governable drivers yield direct actions. Influenceable drivers yield interventions with explicit theories of influence and a measurable proxy. Observable drivers yield monitoring commitments with named signposts. Absorbable drivers yield resilience and optionality, which is Dewar's hedging.

Selection is by expected value within the permissible set, except that where two claims remain plausible and the decision is sensitive to which holds, the correct output is an experiment rather than an action, chosen by the discrimination criterion in Section 8.14. The decision to run an experiment instead of acting is a first-class outcome and should be reported as one, not as an absence of conclusion.

Part Eight: The quantitative spine, coupled

This is the part most often named in prose and not actually run. A File that says its branch weights reflect a Bayesian update without showing the prior vector, the signal-by-signal tempering, and the worked posterior has not run the spine. It has gestured at it. Every step below is fully computed, and every number in this Part was produced by executing the arithmetic rather than by asserting it.

8.1 What the belief layer actually is

Branch weights are the output of a tempered, or power-likelihood, generalized Bayesian update over a latent categorical hypothesis. That is the precise name and it should be used, because it tells a reader what can be checked.

Within a cycle the procedure is a per-source tempered update: raise each signal's likelihood to an exponent set by its source grade and its dialectical status, multiply across signals, multiply by the prior, and normalize.

Raising a likelihood to a fractional power is not an improvisation. It is the standard device of the generalized Bayes literature, motivated by the fact that ordinary Bayesian updating can behave badly under even mild model misspecification, which is the permanent condition of market analysis.

Bissiri, Holmes, and Walker (2016) established that a coherent update can be built on a loss function with the likelihood as a special case; Grünwald and van Ommen (2017) showed the inconsistency that motivates a learning rate below one; Miller and Dunson (2019) showed that conditioning on a neighbourhood of the data rather than the data itself yields a coarsened posterior well approximated by exactly this tempering.

The Dirichlet distribution does real work but at a specific step. It supplies the prior mean and prior strength entering each cycle, and it supplies conjugate carryover across cycles in Section 8.12. It does not describe the within-cycle update.

The honest statement of what this buys and what it does not: tempering makes the update robust to overconfident and correlated evidence, and it makes the analyst's judgment about source quality explicit and auditable. It does not make the likelihood judgments correct. Those remain analyst estimates, and Section 8.16 states the limits plainly.

8.2 The prior and the pseudo-count vector

The prior over branches is the mean of a Dirichlet parameterized by a pseudo-count vector α summing to α₀. Pseudo-counts are assigned on base-rate judgment before the File's own evidence is run, as if the category were known but not this specific case. A flat prior is the right starting point absent a credible reason to favour one branch, and any departure from flat must carry an explicit, challengeable base-rate justification.

BranchPseudo-count αPrior mean
Incumbent Absorbs4.040.0%
Insurgent Wins Category3.030.0%
Substrate Commoditizes (null)3.030.0%
α₀10.0100.0%

Every File must carry a residual or null branch, the world in which none of the main theses plays out as expected, and it may never be assigned zero weight. The null holds probability mass for unknown unknowns and is the numerical expression of the same humility the aporia ledger expresses structurally.

8.3 Source grading and the base temper exponent

Every signal carries a source grade, and the grade sets the base exponent that tempers its influence.

GradeDefinitionBase exponent τ_g
AMotive-clean primary: regulatory filings, court records, primary data, direct disclosure1.0
BMotive-aware or credible secondary: analyst notes, trade press, vendor-adjacent commentary0.5 to 0.6
CSecondary pending verification, or projection: founder interviews about future plans, unverified claims0.15 to 0.3

A Grade A signal enters at full strength; a Grade C signal enters at roughly one-fifth strength. The Grade C signal is still recorded with its raw likelihood, because a ledger that silently drops weak evidence is less auditable than one that discounts it explicitly and shows the discount.

These exponents are calibration targets, not constants. If the calibration record in Section 8.15 shows systematic overconfidence, the first place to look is the Grade B and Grade C exponents, since overweighting motivated secondary sources is the most common route to a well-computed wrong answer.

8.4 The coupling rule

This is the one formal claim the architecture makes beyond assembly, and it is the join between the Dialectical plane and the Belief plane.

A signal's effective exponent is the product of three factors:

τ_eff = τ_g × δ(status) × ρ(redundancy)

where τ_g is the source-grade exponent from Section 8.3, ρ is the redundancy factor from Section 8.5, and δ is the dialectical modulation determined by the grounded labelling of the argument that carries the signal into the ledger:

Dialectical status of the supporting argumentδReading
IN (accepted, defended against all attackers)1.0The signal enters at its full source-graded strength
UNDECIDED (neither defended nor defeated)0.5The signal enters at half strength, since its bearing on the branches is contested
OUT (defeated by an accepted attacker)0.0The signal does not enter the ledger at all

The motivation is direct. Source grade measures how much the evidence can be trusted as a report about the world. It says nothing about whether the inferential step connecting that report to the branch survives attack. A pristine Grade A filing carried into the ledger by an argument that has been undercut should not move the posterior, and under the v2.0 spine it would have moved it at full strength.

Worked, on the ledger of Section 8.6. Signal S2 is a vendor-adjacent analyst note on insurgent share gain, Grade B with τ_g = 0.6, carried by argument A2. Suppose A3 is constructed: the note's share estimate derives from a survey instrument commissioned by the insurgent, which undercuts the inference from the note to the branch without disputing that the note exists or that it says what it says.

Status of A2τ_eff for S2IncumbentInsurgentNullBF Incumbent to Insurgent
IN0.6056.79%29.89%13.31%1.42
UNDECIDED0.3060.49%25.33%14.18%1.79
OUT0.0063.79%21.26%14.95%2.25

The insurgent branch loses 8.6 percentage points when the argument carrying its best signal is defeated, and the Bayes factor between the two live branches rises from 1.42 to 2.25. Neither figure is dramatic on its own. What matters is that the movement is now attributable to a specific defeated inference with a recorded derivation, rather than to an analyst quietly revising a likelihood.

Three honest limitations attach to this rule. The δ values are stipulated rather than derived, and the choice of 0.5 for UNDECIDED is a convention selected for interpretability, not an optimum. The rule is not equivalent to any established probabilistic argumentation semantics and does not claim to satisfy the rationality conditions of epistemic probabilistic argumentation (Hunter & Thimm, 2017). And it makes the belief layer sensitive to graph extraction quality, which is the weakest link in the pipeline. Section 12.4 specifies the conformance test that guards it and Section 13 books the derivation question.

8.5 Redundancy and the independence assumption

The product-over-signals form assumes signals are independent given the branch. That assumption must be stated explicitly in every File that runs the spine, because the redundancy correction is a patch for the case where it fails.

When two signals report the same underlying fact through two outlets, the second is tempered one step further down from its grade-implied exponent, implemented as the ρ factor in Section 8.4. Without this, one fact dressed in two sources inflates the posterior as though the evidence had arrived twice.

Worked. Two Grade B signals at τ = 0.6, two branches X and Y at a 50-50 prior, both signals favouring X with raw likelihoods 0.75 for X and 0.388 for Y, and both reporting the same underlying fact.

Counted naively, each at 0.6: evidence weight for X is 0.75^0.6 × 0.75^0.6 = 0.8415 × 0.8415 = 0.7081, and for Y is 0.388^0.6 × 0.388^0.6 = 0.5667 × 0.5667 = 0.3211, giving a posterior on X of 68.80 percent.

With the correction stepping the second signal down to 0.3: the weight for X is 0.8415 × 0.9173 = 0.7719 and for Y is 0.5667 × 0.7529 = 0.4265, giving 64.41 percent.

The correction recovers 4.39 percentage points of avoided double counting. Across a ledger with several correlated pairs the cumulative correction can move a leading branch by three to four points, which is enough to change which branch leads. Provenance clustering, assigned at the Impression plane in Section 7.1, is what makes the correction mechanical rather than a matter of the analyst noticing.

8.6 The signal ledger and the within-cycle update

Each ledger row carries an identifier, a one-line description, the source, the grade, the base exponent, the dialectical status of its supporting argument, the effective exponent, and a likelihood for every branch on a relative zero-to-one scale. Only the ratios across branches carry information, since the update normalizes at the end. These are analyst judgments written where they can be challenged, not statistical probabilities, and the File must say so.

The update: posterior ∝ prior × Π L(signal | branch)^τ_eff.

Worked, three branches and three signals, all supporting arguments labelled IN so that δ = 1 throughout and the result is directly comparable to v2.0.

IDSignalSourceGradeτ_effIncumbentInsurgentNull
S1Incumbent posts realized revenue tied to contested featurePrimary filingA1.00.800.300.25
S2Vendor-adjacent analyst note on insurgent share gainAnalystB0.60.350.750.35
S3Founder interview projecting category captureInterviewC0.20.300.700.30

Tempered likelihoods, each raw likelihood raised to its effective exponent:

SignalIncumbentInsurgentNull
S1 at 1.00.80000.30000.2500
S2 at 0.60.53270.84150.5327
S3 at 0.20.78600.93110.7860

Evidence weights, the product across all signals:

BranchCalculationWeight
Incumbent Absorbs0.8000 × 0.5327 × 0.78600.33493
Insurgent Wins Category0.3000 × 0.8415 × 0.93110.23506
Substrate Commoditizes (null)0.2500 × 0.5327 × 0.78600.10467

Posterior, prior times evidence weight, normalized:

BranchPriorEvidence weightPosteriorMove
Incumbent Absorbs40.0%0.3349356.79%×1.42
Insurgent Wins Category30.0%0.2350629.89%×1.00
Substrate Commoditizes (null)30.0%0.1046713.31%×0.44

Read the moves alongside the weights. The ledger moved conviction toward the incumbent and away from the null while leaving the insurgent almost exactly where the prior had it. That flat move is informative rather than empty: the evidence for the insurgent was real but offset by a stronger, higher-grade signal pointing the other way. A flat move is the ledger reporting that this cycle's evidence was a wash for that branch.

8.7 Bayes factors and the interpretive scale

The ratio of two branches' evidence weights, computed before the prior enters, is the Bayes factor. It states how decisively the evidence alone favours one branch over another, independent of starting belief.

From the worked ledger: incumbent against null is 0.33493 / 0.10467 = 3.20 to 1; insurgent against null is 0.23506 / 0.10467 = 2.25 to 1; and incumbent against insurgent is 0.33493 / 0.23506 = 1.42 to 1.

Bayes factors are read against the Kass and Raftery (1995) scale, never by feel.

Bayes factorInterpretation
1 to 3Barely worth a mention
3 to 20Positive evidence
20 to 150Strong evidence
Above 150Decisive evidence

Reporting factors alongside posteriors lets a reader separate the claim that the evidence favours a branch from the claim that the prior already favoured it and the evidence barely moved it. Precision in this language is how the spine stays honest.

8.8 The aporia test, and a correction to the v2.0 worked example

The three Bayes factors above license one conclusion and forbid another, and v2.0 drew the forbidden one.

What they license: the null is genuinely disfavoured. At 3.20 to 1 the evidence against it is positive on Kass and Raftery, and the File may report that the world in which the substrate simply commoditizes is now the weakest of the three readings.

What they forbid: any claim that the incumbent branch leads. At 1.42 to 1 against the insurgent, the evidence is barely worth a mention. The 26.9-point gap in the posterior between 56.79 and 29.89 percent is inherited from the prior, which put the incumbent 10 points ahead before any evidence ran and which the evidence then amplified without warranting. Version 2.0 read this as a tilt toward the incumbent. That reading takes a prior asymmetry and reports it as a finding.

The test is therefore made formal and fires automatically.

The aporia test. A File is in belief-layer aporia when the two highest-weighted branches are separated by a Bayes factor below 3.0, both clear the reporting floor, and together they hold at least 60 percent of the branch mass.

On the worked example: 1.42 is below 3.0, both branches clear the 5 percent floor, and together they hold 86.7 percent. The test fires. The correct disposition for that cycle is a ledger note plus an opened aporia with the question of what observation would discriminate between incumbent absorption and insurgent category capture, not a reweight announcing a tilt.

This is the most consequential change in v3.0 and it is uncomfortable in the right way. The instrument's own published teaching example was overstated, the overstatement was invisible under the v2.0 procedure, and the new test catches it mechanically rather than depending on an analyst's scruple. Any File written under v2.0 whose leading two branches sit inside a Bayes factor of 3 should be re-read under this test at its next cycle.

8.9 The floor and renormalization

Every File carries a floor, conventionally 5 percent, below which no branch reports its bare computed value. A branch computing to 1 or 2 percent usually reflects the limits of the analyst's likelihood judgments rather than genuine near-impossibility, and reporting false precision at the tail invites overconfidence exactly where evidence is thinnest.

The procedure is a single simultaneous pass. Identify all branches below the floor, raise each to the floor at once, compute the total mass added, remove that same total proportionally from the branches that were above the floor, and renormalize once.

Worked, three branches with strong evidence for B. Raw posteriors are A at 12.4 percent, B at 87.2 percent, and C at 0.4 percent. C falls below the floor, so 4.6 points are added to reach it and removed proportionally from A and B, which hold 12.4 percent and 87.6 percent of the above-floor mass respectively. The result is A at 11.8 percent, B at 83.2 percent, C at 5.0 percent, summing to 100.0 percent.

8.10 Impression separation applied to the ledger

The Impression plane changes ledger construction, and the worked example shows how much.

Signal S3 as written in v2.0 is a founder interview projecting category capture. Under Section 7.1 that input decomposes into an observation, which is that a named founder publicly committed to a category-capture position on a stated date, and a forecast, which is the content of that projection. The forecast is a claim with suspended assent. Only the observation may enter the ledger, and its likelihood vector must answer a different question: what does the fact of a founder making this public commitment imply about which branch is running?

That is a genuinely weaker and more ambiguous signal than the projection taken at face value. Public commitment is mild evidence of insurgent conviction and resource allocation, and it is also what a founder does when the private numbers are not yet persuasive. A defensible re-judgment is 0.45 for the incumbent, 0.55 for the insurgent, and 0.40 for the null, still at Grade C with τ = 0.2. Rerunning the ledger with S1 and S2 unchanged:

BranchPriorEvidence weightPosteriorMove
Incumbent Absorbs40.0%0.3632259.12%×1.48
Insurgent Wins Category30.0%0.2239927.34%×0.91
Substrate Commoditizes (null)30.0%0.1108613.53%×0.45

The incumbent gains 2.3 points and the insurgent loses 2.6, and the Bayes factor between them moves from 1.42 to 1.62, still well inside the aporia band. The lesson is not the direction of the move. It is that under v2.0 a founder's forecast entered the ledger as though it were an observation about the world, and its likelihood vector was implicitly answering the question of whether the founder is right rather than what the founder's saying it implies. That confusion is invisible in a ledger and structurally impossible under the Impression plane.

8.11 The conditional causal tier

The causal tier runs only when a realized outcome panel exists, meaning actual resolved outcomes across comparable past cases, not opinions about an unresolved current situation. When the panel does not qualify the tier is explicitly held and noted as held with the reason, never silently omitted, because the absence is itself informative.

Four components sit in three categories, and confusing the categories is a common and serious failure.

Two are causal estimators. A causal forest estimates a conditional average treatment effect across comparable cases, for instance the differential effect of a deal structure on time-to-scale across resolved acquisitions (Wager & Athey, 2018; Athey, Tibshirani, & Wager, 2019). A Bayesian structural time-series event study estimates a counterfactual trajectory around a single dated intervention, for instance what an adoption curve would have done absent a specific regulatory action (Brodersen et al., 2015).

One is an attribution component and is not a causal estimator. A gradient-boosted classifier predicts which branch a new case most resembles, with Shapley attributions decomposing the prediction into per-feature contributions (Lundberg & Lee, 2017). This says what features drove the prediction. It does not say those features caused the outcome, and a File that reports Shapley values as causal findings has made a category error that no amount of computational care redeems.

One is a regime-conditioning model. A hidden Markov formulation in the Hamilton (1989) tradition infers which latent market regime is active. Regime is latent, so the model produces a posterior over regimes rather than a label, and the File conditions on that full posterior. Hard-assigning a single regime discards exactly the uncertainty the model was built to carry.

Where the panel is too small to run the causal estimators responsibly, conventionally fewer than eight to ten comparable resolved cases, present the available cases as a structured comparison table and use the attribution component qualitatively as a feature read rather than a trained model. A small honest comparison is worth more than a model run past its statistical license.

8.12 The cross-cycle update

Within a cycle, Sections 8.4 through 8.9 produce a posterior. Across cycles that posterior must become the next prior without undermining falsifiability. The rule uses a forgetting factor on existing pseudo-counts plus evidence-derived counts added fresh:

α ← λ · α + m · ê

where α is the entering pseudo-count vector, ê is this cycle's normalized evidence weight vector, m is the total count of new evidence strength added this cycle, calibrated to the number and grade of new signals and typically between 3 and 8 per quarterly cycle, and λ is the forgetting factor, conventionally 0.7 to 0.9 depending on how fast the market actually changes.

This rule has a finite steady state. At λ = 0.8 and m = 5, total prior strength converges to α₀ = m / (1 − λ) = 25. Conviction becomes sticky at that level rather than growing without bound, so a File fed steady contrary evidence keeps moving by a predictable amount every cycle instead of moving less and less as accumulated prior swamps new evidence.

The trajectory from the worked example, starting at α₀ = 10 with the normalized evidence weights of Section 8.6 held constant for five cycles:

Cycleα₀IncumbentInsurgentNull
113.0043.7%31.9%24.4%
215.4045.6%32.8%21.5%
317.3246.8%33.4%19.8%
418.8647.5%33.8%18.7%
520.0848.1%34.1%17.9%

Note what this shows and what it does not. Prior strength climbs toward 25 and the prior mean drifts toward the evidence, but even after five cycles of identical evidence the incumbent prior mean is 48.1 percent rather than anything approaching certainty. That damping is the intended behaviour.

The forbidden alternative should be named so nobody reinvents it. Adding a multiple of the posterior itself to the pseudo-count vector each cycle inflates α₀ without bound and compounds old conviction on top of itself, making a long-lived File progressively harder to move. A File that is asymptotically unfalsifiable is not a living dossier. It is a defended position dressed as analysis.

8.13 The sensitivity sweep

Before publishing a reweight, perturb the analyst-supplied likelihoods and recompute. The convention is to vary each likelihood by plus or minus ten to fifteen percentage points and observe whether the leading branch changes.

Perturbing the S1 likelihood for the incumbent while holding S2 and S3 constant:

S1 likelihood for IncumbentPosterior: IncumbentBayes factor against Insurgent
0.70, minus ten points53.5%1.25
0.80, base56.8%1.42
0.90, plus ten points59.7%1.60

The leading branch holds across the range, which under v2.0 would have been reported as a robust result. Under v3.0 the sweep must also report the Bayes factor at each point, and here the factor stays below 3.0 throughout, meaning the aporia test fires across the entire perturbation range. Robustness of a lead is not the same as diagnosticity of the evidence, and a sweep that reports only the former is reassuring for the wrong reason.

Where a plausible perturbation flips the leading branch, that fragility is the most important finding of the cycle and is reported as the central result rather than buried beneath the headline number.

8.14 Expected information gain and probe selection

When two branches remain live and the decision is sensitive to which holds, the correct output is a probe, and the choice among candidate probes is computed rather than felt. This is Howard's information value theory (1966) reduced to the branch simplex.

For a candidate probe with binary outcome, specify the probability of the affirmative outcome under each branch. Compute the predictive probability of each outcome under the current posterior, the posterior that would follow each outcome, and the expected entropy afterwards. The expected information gain is the current entropy minus the expected posterior entropy.

Worked, from the Section 8.6 posterior of 56.79, 29.89, and 13.31 percent, whose entropy is 1.3716 bits. Consider a discriminating probe, for instance instrumenting a two-platform delegation pilot and observing whether transaction volume through incumbent policy surfaces rises. Assign the probability of that affirmative outcome as 0.85 under incumbent absorption, 0.25 under insurgent capture, and 0.40 under the null.

OutcomeProbabilityResulting posteriorEntropy
Affirmative0.610779.0 / 12.2 / 8.70.9459 bits
Negative0.389321.9 / 57.6 / 20.51.4070 bits

Expected posterior entropy is 1.1254 bits, so the expected information gain is 0.2462 bits.

Compare a weakly discriminating probe with outcome probabilities of 0.55, 0.45, and 0.50, differing little across branches. Its expected information gain is 0.0057 bits, roughly forty-three times smaller. Both probes would produce a data point and a slide. Only one changes what the File is entitled to believe.

The selection rule divides expected information gain by cost and time, and the stopping rule in Section 6.9 compares the best available figure against the cost of the reasoning step it would replace. The virtue of doing this arithmetically is that it makes the most expensive failure mode visible: an organization can run many probes with near-zero discriminatory power and experience the activity as diligence.

8.15 Calibration and scoring

A File that issues posteriors is issuing probabilistic forecasts, and forecasts can be scored. When a branch resolves, meaning the underlying question becomes a recorded fact, the File's final pre-resolution posterior is scored against the outcome with a proper scoring rule (Brier, 1950). Scoring is tracked across the whole book of Files, never File by File, because a single call carries almost no information about calibration.

Worked, on a five-File book with multi-class Brier and log scores:

FileFinal posteriorOutcomeBrierLog score (bits)
156.8 / 29.9 / 13.3Incumbent0.29370.8160
220.0 / 65.0 / 15.0Insurgent0.18500.6215
345.0 / 40.0 / 15.0Null1.08502.7370
470.0 / 20.0 / 10.0Incumbent0.14000.5146
535.0 / 55.0 / 10.0Insurgent0.33500.8625

Mean Brier is 0.4077 against a uniform-guess reference of 0.6667, giving a Brier skill score of 0.3884. Mean log score is 1.1103 bits.

Two readings follow, and the second is the useful one. The book beats uniform guessing by a meaningful margin, which is the minimum bar. More instructive is File 3, whose Brier of 1.0850 exceeds the uniform reference on its own: the File held the null at 15 percent and the null resolved. One badly calibrated call on a floored branch dominates the book's average, which is exactly why the floor exists and why null branches must never be treated as decoration.

The diagnostic pattern to watch across a larger book is straightforward. A practitioner whose 70 percent calls resolve true about 70 percent of the time is calibrated. One whose 70 percent calls resolve true 90 percent of the time is understating confidence. One whose 70 percent calls resolve true 40 percent of the time has a systematic bias, most often in how aggressively low-grade signals are weighted, which points directly at the Grade B and C exponents in Section 8.3. Neither overconfidence nor underconfidence is safe, and both are correctable once the record exists. This is the mechanism by which the Good Judgment Project work showed calibration to be trainable (Mellers et al., 2014).

8.16 What the spine cannot do

Four limits are stated so that no File claims more than the arithmetic supports.

The likelihood vectors are analyst judgments. Everything downstream inherits their quality, and the sensitivity sweep bounds the damage without eliminating it. The spine makes judgment auditable, not correct.

The branch set is assumed exhaustive and mutually exclusive. It is neither, in general. The null branch absorbs some of the failure, but a File whose real answer lies in a branch nobody wrote will assign that answer zero mass and never notice. This is why DIALECTICAL_EXPAND is required to produce a non-binary reframing and why the malformed-question aporia trigger exists.

Independence given the branch is assumed and is routinely violated in ways the cluster mechanism only partially catches. Correlated evidence that shares an unobserved common cause, rather than a common outlet, passes the redundancy check undetected.

The causal tier's licence is narrow. Absent a genuine panel of resolved comparable cases it is held, and holding it is the correct answer far more often than practitioners like.

Part Nine: The market-structure instruments

These instruments are what make the Machine a market instrument rather than a general reasoning system. They are retained from v2.0 with two changes: the detection tells are relocated to the Impression plane, where they belong, and the cascade is tightened into a structural claim rather than a narrative one.

9.1 The six detection tells

Run these as symptom-driven detectors before touching any structural model. They find constraints by reading what the market is already confessing, and each is an observation class rather than an interpretation, which is why they sit at the Impression plane.

Spend without return. Money flows in and measurable return does not come back out. The capital is paying to push against a wall.

Where pilots die. Initiatives reach proof of concept and stall at the same production gate in every company trying. Whatever kills them at that consistent point is the constraint.

The universal workaround. Everyone independently builds the same patch for the same gap. The patch marks the missing layer.

The preemptive apology. Every vendor apologizes for the same limitation before being asked. They are all bounded by the same wall.

The not-yet chorus. The people closest to the work keep saying the capability is not ready for the real thing. Listen to what specifically is not yet ready, not that it is not ready.

Convergence. Competitors independently reposition onto the same layer and begin using the same word for what they are doing. That language convergence names the constraint, because they are all circling the same wall.

9.2 The force scan

Run Porter's five forces plus complementors as a six-location coverage scan: buyer power, supplier power, new entrants, substitutes, rivalry, and complements (Porter, 1979; Brandenburger & Nalebuff, 1996). The sixth is where AI-era constraints disproportionately hide, since a complementor's pricing or access decision can bind a market that the original five would read as healthy.

The discipline between the two methods is strict. Porter lists, the Machine ranks. The force scan guarantees coverage so nothing is missed, and the tells plus the constraint logic decide which location is binding right now. A File that produces a five-force table without a ranking has performed a survey, not an analysis.

9.3 Ownership and the owner's game

List who owns each layer of the value chain, then mark the binding layer. Watch for the case where no vendor owns the binding constraint at all, because it sits inside the buyer's own risk function or balance sheet. When that happens, the prize goes to whoever relieves the buyer rather than to whoever holds a supply-side layer, and Files routinely miss this by searching only among vendors.

Owners of a binding constraint have every incentive to keep it binding and extract rent from it. Naming the move predicts what the owner will fight hardest to protect.

Owner moveWhat it doesWhat erodes it
DefendHardens the wall, raises switching costA substitute routes around the wall entirely
TollCharges rent for passage without removing the wallRegulation caps the toll, or a free alternative undercuts it
AbsorbAcquires the source of relief before it scalesThe acquired capability is reproduced independently
CommoditizeA platform tears the wall down to win the layer above itWhoever commoditizes must win the next layer up, or the move is self-defeating

A complete ownership map names, for each layer, who pays, who captures, and who is liable if the layer fails. The liability column is the one most often left blank and the one that most often predicts where regulation lands.

9.4 The six moat tests

Run all six against any moat under examination. A moat is a profile across six dimensions, not a binary judgment.

TestThe question
Constraint fitDoes the moat sit on the binding constraint, or on a feature nobody is paying for?
TransferabilityDoes the advantage transfer to adjacent decisions, or only inside one narrow workflow?
Economic conversionHas the moat converted to revenue, margin, or share, or is it still a story?
Control point leverageDoes owning this point give leverage over the rest of the system, or is it isolated?
Adversarial responseWhat does a well-funded rival do in response, and how long does the moat survive it?
Learning compoundingDoes the system get stronger with every deployment, customer, action, and feedback loop?

Under time pressure, score the two that matter most for the specific position: the one the moat passes most decisively and the one that could kill it. That pairing usually locates it.

9.5 The cascade

Relieving the binding constraint does not end the story. It relocates the constraint, and the cascade traces that relocation forward two or three links so a File anticipates where the fight moves next rather than declaring victory at the first resolution. Each link is a market reorganization and a movement of the control point, and mapping three links makes the shape of the next several years visible rather than only the next quarter.

The cascade is not a repeat of the branch map. The branch map states what the plausible futures are and how much weight current evidence assigns to each. The cascade states what happens next inside the leading branch, how the constraint relocates, where the control point moves, and what the fight looks like in the subsequent round. If the branch map is the map, the cascade is the path on it.

In the most rigorous runs the cascade is rendered as a structural causal model with named paths, and where a resolved panel exists the path coefficients are estimated under the tier of Section 8.11. Where no panel exists the cascade is explicitly qualitative and must say so, because a directed graph with arrows and no coefficients invites readers to treat a narrative as a model.

9.6 The Unguarded Flank

Every File identifies the open strategic position nobody is currently occupying. Three gates determine whether a candidate is a genuine flank or a wishful gap, and a candidate must pass all three.

Structural block: is there a specific reason no current player can occupy this position, rather than merely that none has tried? Constraint, not feature: does the position sit on the binding constraint itself rather than on something useful but not gating? Compounds before arrival: does waiting allow someone else's flywheel to make the position unreachable before a new entrant could claim it?

Passing two and failing the third is the most common error, and it is worse than a clean miss because it looks like a real finding. The Unguarded Flank section of a File must show the three gates explicitly, with the failure mode named where a gate is passed narrowly.

Part Ten: Governance

10.1 Invariants

These are the properties that must hold in every state. Each names its enforcement point, and a violation is a defect rather than a judgment call.

Invariant 1: no unsupported acceptance. A claim may hold assent ACCEPT only if its support set is non-empty, except for explicitly axiomatic or constitutional propositions. Enforced by the assent function.

Invariant 2: contradictions cannot vanish silently. Once a contradiction relation is recorded between two nodes, it persists until explicitly resolved, reframed, or superseded, each with a recorded reason. Enforced by the graph writer.

Invariant 3: assumptions never become facts. A node typed as assumption cannot mutate to fact. It may be superseded by a distinct evidence node that supports the same proposition, and the superseding is a recorded event. Enforced by the type validator.

Invariant 4: evidence must have provenance. No evidence node exists without a source, a grade, a provenance cluster, and an ingestion record. Enforced by the Impression plane schema.

Invariant 5: suspended belief cannot authorize irreversible high-risk action. If a claim is SUSPENDED, an action depends on it, and the action is high risk, the action is denied. Enforced by the constitutional gate.

Invariant 6: revision preserves history. Superseding a claim never destroys the previous version, and the revision records what changed, why, and on what observation. Enforced by the persistence layer.

Invariant 7: the generator holds no write path to authoritative state. No generative component may write to the argument graph labelling, the posterior, the assent states, or the constitutional verdicts. It may only propose typed candidates. Enforced architecturally by the plane boundaries.

Invariant 8: derivation accompanies contradiction. No contradiction is admitted without an explicit derivation chain. Enforced by ELENCHUS.

Invariant 9: the null branch is never zero. Every File carries a residual branch, and after flooring it holds at least the floor value. Enforced by the belief layer.

Invariant 10: dialectical status gates evidentiary weight. A signal's effective exponent is computed from the coupling rule, never assigned directly by an analyst. Enforced by the ledger writer.

10.2 Two-speed control

The constraint thesis at the top of a File moves slowly, only on a named kill condition. The branch weights and the signal ledger move quickly, every cycle, as evidence arrives. Running both at the same speed is how a File loses its discipline: revising the thesis on every signal chases noise, and freezing the weights because the thesis feels settled produces a document that already knows what it is going to say.

This two-speed rule is local to the Inquiry and is a discipline on the analyst, not a component. It should not be confused with the ADAM Governor, which is a structural brake and enforcement point holding no write path, nor with Pola, which is the policy decision point. Any future realization of this Machine inside ADAM would use those components for what they do and would keep the two-speed rule as what it is, a cadence discipline.

10.3 Generation and verification separation

A number is never trusted because it was computed. It is trusted because it was checked by someone other than the person who computed it, or by the same person at a different time with fresh eyes. The signal ledger's likelihood judgments must be examinable by any reader and challengeable by any co-author, and opacity in inputs is the fastest path to a File that confirms its own prior rather than interrogating it.

Operationally this means the Inquiry stores the derivation, not only the result, for every computed quantity: the prior vector, every effective exponent with its three factors shown, the evidence weights, the Bayes factors, the floor arithmetic, and the sweep. A File that publishes a posterior without the means to recompute it has published an assertion.

10.4 Kill conditions

A kill condition is a pre-specified falsifying observation: if this specific thing happens, the thesis is wrong, not merely the weights. Kill conditions are written at File creation, never invented after the fact, and each attaches to a necessary and vulnerable assumption identified by EXPOSE. This is Dewar's signpost (2002) with a stricter obligation, because a signpost warns while a kill condition commits.

Kill conditions are first-class nodes with an observable trigger, a named data source that would show the trigger, and the disposition that follows when it fires, which is always thesis review. A kill condition whose trigger cannot be observed by any named source is not a kill condition; it is a hedge, and the File must either find the source or drop the condition.

10.5 The four-action maintenance classification

Every cycle a File receives exactly one of four dispositions. Logging the disposition is not optional, because the maintenance log is how a File proves it is living rather than dormant.

DispositionWhen it applies
Ledger noteNew evidence arrived but did not move the posterior meaningfully, or the aporia test fires and blocks a reweight
ReweightNew evidence moved the posterior enough to warrant a republished branch map, and the aporia test does not fire
Thesis reviewA kill condition fired, requiring the constraint thesis itself to be re-examined
RetirementThe constraint has fully resolved, or the thesis has been structurally falsified

The sensitivity sweep and the calibration entry are run before every reweight and recorded alongside the disposition. Aporia is not a fifth disposition; it is a File state that can coexist with any of the four, and the most common combination in a healthy File is a ledger note carrying one or two open aporias.

10.6 Belief authority and action authority

Humans are dialectical participants, not merely approvers at the end. A human may assert a claim, reject one, accept a commitment, challenge a definition, provide evidence, declare a value, open an aporia, override an assent state, or authorize an experiment. Every one of those acts is recorded with its actor.

The critical rule is the separation of two authorities that look alike and are not. If a principal says they understand the evidence and choose to pursue a hypothesis anyway, the Machine records the epistemic status as PROVISIONAL and the strategic status as AUTHORIZED_EXPERIMENT with the authorizing human named. It must never rewrite PROVISIONAL to ACCEPT because authority approved action. Belief authority is earned from evidence and argument. Action authority is held by people. Collapsing them is how organizations come to believe their own decisions.

10.7 Audit, events, and replay

Every state change emits an event: claim proposed, definition established, assumption exposed, commitment accepted, argument constructed, argument attacked, contradiction detected, aporia opened, question reframed, evidence added, assent changed, agency classified, action proposed, action rejected by constitution, experiment authorized, observation recorded, belief revised, aporia closed.

The requirement this satisfies is replay. The Machine must be able to reconstruct, for any date, why it believed what it believed, what evidence existed then, what assumptions the belief depended on, what contradicted it, why it changed its mind, and which observation caused the revision. An event log without replay is a diary. Replay is what makes intellectual evolution auditable rather than narrated.

The event stream should be treated as the accountable action history in the same sense that Actra is within the 451 canon, and if a realization is built inside ADAM it should emit to Actra rather than maintaining a private log.

Part Eleven: File anatomy and the complete run

11.1 Required sections of a File

A complete File contains, at minimum and in working order: a constraint thesis with type and clock; a force scan ranked by binding force; an ownership map with the owner's-game rent route named; the six moat tests run against the load-bearing moat; a weighted branch map summing to 100 percent with the floor applied as a single simultaneous pass; a dated signal ledger with grade, dialectical status, effective exponent shown as its three factors, and per-branch likelihoods; the posterior, the Bayes factors against the null and between the two leading branches, and the Kass and Raftery read in plain language; the aporia test result, stated explicitly as fired or not fired; the open aporia ledger with resolution conditions; the conditional causal tier, either run or explicitly held with the reason; the cascade mapped at least two links forward; the Unguarded Flank tested against all three gates; the agency map classifying every material driver; actor reads for the incumbent, the challenger, a capital allocator, and a regulator or buyer; the maintenance log with this cycle's disposition; the sensitivity sweep result including Bayes factors at each perturbation; and source-graded references.

Four of these are new in v3.0: the dialectical status column in the ledger, the aporia test result, the open aporia ledger, and the agency map. A File missing any required section is incomplete rather than light, and the absent section is usually the one that would have been most uncomfortable to write.

11.2 The complete run

Twelve stages, carried through the worked example so the parts connect into a single motion. Stages one through three are new; stages four through nine are the v2.0 run with the coupling added; stages ten through twelve are new.

Stage one: impression. The trigger is a realized revenue disclosure from the incumbent. Decompose it. The observation is that a specific figure was reported in a specific filing for a specific period. The accompanying management commentary about competitive positioning is attribution, and the trade-press framing that this proves the incumbent is winning is interpretation with assent suspended. Only the observation proceeds. Assign source grade A, provenance cluster, and motive exposure none.

Stage two: definition. The active question turns on the contested layer. Run DEFINE on the layer's name. If two senses classify the same near-miss case differently, the question is asking about two layers and must be split before anything is weighted. Assume here that a single sense survives with an explicit exclusion list.

Stage three: examination. Run EXPOSE on the claim that the disclosure indicates absorption, which surfaces at least the assumptions that the revenue is incremental rather than reclassified, that it is attributable to the contested feature rather than bundled, and that the reporting boundary has not changed. Run COUNTEREXAMPLE against the generalization. Run ELENCHUS against standing commitments. Build the argument graph and label it under grounded semantics.

Stage four: detect the constraint. Run the tells. Convergence fires: rivals are repositioning onto the same contested layer and adopting the incumbent's language. The force scan points to rivalry as the binding force. Name the constraint as economic, since the question is whether the insurgent's unit economics reach scale before the incumbent's distribution advantage closes the gap. Clock: cost-curve, predictable but contested.

Stage five: test the moat. Run the six tests against the insurgent. Learning compounding passes strongly. Economic conversion is now a live question, which is exactly what the stage-one disclosure makes it. Control point leverage is weak, since incumbent distribution still gates most of the addressable market.

Stage six: map ownership. The contested layer sits between the incumbent, who owns distribution, and the insurgent, who owns the technical lead. Neither owns the buyer's underlying budget constraint, and that gap is where the rent route lives.

Stage seven: model the owner's game. The incumbent's move is toll, charging for access to distribution rather than building the contested layer. That move is precisely what generates the convergence tell, since tolling signals the layer is monetizable and competitors copy the language.

Stage eight: run the spine. Prior at 40-30-30. Three signals through the ledger with effective exponents computed from grade, dialectical status, and redundancy. Posterior: 56.79, 29.89, 13.31. Bayes factors of 3.20 and 2.25 against the null, both positive on Kass and Raftery, and 1.42 between the two live branches. All branches clear the floor. Sensitivity sweep holds the lead but keeps the factor below 3.0 throughout.

Stage nine: apply the aporia test. It fires. Two leading branches inside a factor of 3, both above floor, together holding 86.7 percent. The File may report that the null is disfavoured. It may not report a tilt. Disposition: ledger note, aporia opened.

Stage ten: agency map. Classify each driver. Insurgent unit economics: observe. Incumbent tolling policy: observe. Standards participation in the contested layer: influence. Our own interoperability and governance architecture: govern. Our partner access: influence. Macro AI investment: absorb. The map immediately tells a builder that betting on the branch outcome is not a strategy, because nothing in the branch determination is governable.

Stage eleven: probe selection. Compute expected information gain across candidate probes. The discriminating probe returns 0.2462 bits against 0.0057 for the weak alternative. Select the discriminating probe, specify its measurements, and record the prediction each branch makes about its outcome, because a probe without a recorded prediction cannot produce a reflection.

Stage twelve: governed action and cascade. Filter options constitutionally, then read the cascade. Inside the incumbent branch, the next contested layer is the insurgent's remaining independent customers; the link after that is whether tolling generates enough cross-industry friction to invite a neutral infrastructure response. Actor reads follow: the incumbent keeps tolling rather than building, since toll is working and absorption is expensive; the insurgent finds the segment incumbent distribution structurally cannot serve and defends it as a beachhead; a capital allocator waits for the probe rather than the next reweight, because the reweight is currently blocked by the aporia test; a regulator watches the toll rate rather than the market share.

11.3 Actor reads

Every File concludes with a read from each major class of participant, because writing all four forces a test of whether the branch weights hold up from multiple vantage points rather than only from the analyst's preferred one. A branch that looks overwhelmingly dominant from one actor's position and nonsensical from another's is signalling something the single-vantage analysis missed, and that signal should be treated as a candidate aporia rather than as an inconsistency to smooth over.

11.4 Writing discipline

Every File and every essay built from it is led by one argument structure with at most one supporting structure, chosen to match the reader's decision rather than the convenience of the material. The eight structures are concede-then-pivot, thesis-and-kill-signal, constraint-first, steelman-then-break, structural disqualification, cost-of-error asymmetry, pattern accumulation, and two-sided weighted brief. Concede-then-pivot is the default in collaborative and co-creation contexts. Constraint-first leads whenever the constraint has not yet been named in public discourse.

Prose conventions hold throughout. No em dashes. Paragraphs of three to five sentences, roughly sixty to one hundred twenty words, one idea each, with sentence length varied deliberately so that a short sentence after a long one carries emphasis. No anaphora stacking, no rhetorical question chains, no antithetical constructions of the not-X-but-Y form, no ordinal scaffolding in prose, no staccato single-line paragraphs, and no bold run-in labels outside instrument sections and tables.

Uncertainty is carried in the branch weights, the Bayes factors, and the aporia test, never in adjectives. A hedge in the prose of a File that has already stated its weights is redundant at best and evasive at worst. The strongest opposing argument is conceded by name and in specific terms before any pivot, since a vague concession only gestures at conceding.

Part Twelve: Implementation

12.1 Build thesis

Stated as cost-of-error asymmetry, because that is what should govern build order. The expensive failure is not a slow Machine or an incomplete one. It is a Machine that produces a confident, well-formatted, fully computed answer whose derivation cannot be reconstructed, because that artifact is indistinguishable from a correct one and will be trusted. Every sequencing decision below prioritizes reconstructability over capability.

The second-most expensive failure is a Machine so thorough that it never issues a disposition. The scheduler and stopping rules are therefore built in the first increment, not the last, even though they are the least interesting component.

12.2 Increments

Four increments, each with entry and exit criteria and a decision gate carrying an RDS verdict question.

WIN-DK1: the Inquiry graph. Build the node types, the relation vocabulary, immutable provenance, the event stream, and replay. Build the scheduler and the five stopping conditions. No operators, no belief layer, no generation: the deliverable is a persistence and audit substrate that a human analyst can drive by hand. Entry criteria: schemas from Part Five ratified. Exit criteria: an existing published File is reconstructed as an Inquiry, and replay answers the question of why a stated weight was held on a stated date. Gate: does replay work on a real File, and can an analyst drive it without the tooling costing more time than it saves? Stop if replay cannot be demonstrated on a real historical File.

WIN-DK2: the Socratic kernel. Implement DEFINE, EXPOSE, ELICIT, ELENCHUS, COUNTEREXAMPLE, DIALECTICAL_EXPAND, DETECT_APORIA, and REFORMULATE, with a generative layer proposing candidates and typed validation admitting them. Implement grounded labelling over the argument graph. Entry criteria: DK1 exits. Exit criteria: the conformance suite in Section 12.4 passes on the Socratic cases, and ELENCHUS emits derivations for every contradiction. Gate: does the kernel find contradictions a competent analyst missed, and does it avoid flooding the ledger with immaterial ones? Redesign the materiality threshold if precision is below a workable level.

WIN-DK3: the Stoic kernel and the coupling. Implement IMPRESSION, ASSENT, PAUSE, CLASSIFY_AGENCY, and CONSTITUTIONAL_FILTER. Wire the coupling rule of Section 8.4 between the labelling and the ledger. Entry criteria: DK2 exits with stable labelling. Exit criteria: the coupling test in Section 12.4 passes, and the ledgers of at least three historical Files are rebuilt with impression separation and their posteriors compared against the originals. Gate: does the rebuild change any published conclusion, and if so, is the change defensible on inspection? This is the highest-risk gate in the programme.

WIN-DK4: the learning loop. Implement probe selection by expected information gain, experiment emission, observation capture, REFLECT, and minimum-loss REVISE. Entry criteria: DK3 exits. Exit criteria: at least one probe is selected by computed gain, run, and reflected, with the belief revision recorded and replayable. Gate: did reality change a belief, and did the record survive the change? After DK4 the Machine is genuinely recursive, and only then.

12.3 Stack decisions

DecisionChoiceVerdictRevisit trigger
Persistence for the Inquiry graphAppend-only event store with a derived graph projectionProceedProjection rebuild time exceeds an analyst's patience on a large File
Adjudication semanticsDung grounded semantics as default, labelling recomputed on every graph writeProceedGrounded proves too sceptical in practice, leaving most claims UNDECIDED, at which point evaluate preferred or a gradual semantics
Structured argumentation formalismASPIC+ attack vocabulary adopted for the relation types; the full formalism not implementedProceed as vocabulary; Stop as full implementationExtraction quality improves enough that strict and defeasible rule sets are maintainable
Belief layer implementationExplicit arithmetic in a scripted numerical layer, every intermediate persistedAccelerateNone foreseen; this is the cheapest and most auditable component
GenerationModel-agnostic behind a typed candidate interface, prompts versioned as dataAccelerateAny design pressure to make a specific model load-bearing
Argument extraction from textHuman-in-the-loop with generative proposal, never fully automaticProceedA conformance-tested extractor reaches acceptable agreement with analyst labelling on the File corpus
Causal tierExternal statistical packages invoked only when the panel qualifiesProceedPanels routinely qualify, at which point promote to a standing component

12.4 Conformance suite

A system is not dialectical merely because it can produce debate. These cases are the minimum bar, and each states the input and the required output.

CaseInputRequired output
Undefined concept"AI will replace work"DEFINITION_REQUIRED, with at least two senses that classify a near-miss differently
Hidden assumption"Agents will destroy software revenue"At least one causal and one economic assumption surfaced, each marked necessary or contributory
Contradiction with derivationCommitments that all high-risk actions require approval, that system X acts without approval, and that X compliesCONTRADICTION_DETECTED with an explicit derivation chain, not a similarity judgment
Counterexample defeats generalization"Lower interface usage always reduces vendor revenue", plus a transaction-priced vendorGeneralization scope reduced or claim defeated, with the reduction recorded
Conflicting evidenceTwo Grade A signals pointing opposite waysCONTESTED, not arbitrary acceptance, and both signals retained
Insufficient evidenceA forecast claim with no supporting evidence nodeSUSPEND
Malformed question"Will AI either eliminate all enterprise software or fail?"False binary detected, aporia opened, reformulation scored higher on the quality function
Out-of-agency goal"Make competitors stop innovating"Not classified GOVERN
Constitutional conflictAn option violating an explicit norm with high expected valueRejected before scoring, with the norm named
Coupling testA Grade A signal whose supporting argument is undercut by an accepted attackerSignal excluded from the ledger, posterior recomputed, exclusion recorded with the attacking argument named
Aporia testA ledger yielding two leading branches inside a Bayes factor of 3Reweight blocked, ledger note issued, aporia opened
Impression separationA headline containing an observation and an implied judgmentTwo nodes emitted, judgment at suspended assent, only the observation eligible for the ledger
RedundancyTwo signals sharing a provenance clusterSecond signal's exponent stepped down, posterior differing from the naive computation
Belief revisionAn experiment contradicting an accepted predictionBelief revised, prior version preserved, closed aporia reopened if it bore on the contradiction
ReplayAny historical dateFull reconstruction of belief state, supporting evidence, and revision cause

12.5 Evaluation

Three families of metric, and they measure different things that are easy to conflate.

Process metrics measure whether the Machine is running its own discipline: proportion of ledger rows whose supporting argument is labelled, proportion of contradictions carrying derivations, proportion of inputs decomposed at the Impression plane, and aporia closure rate against aporia opening rate. These are cheap and they catch decay early.

Output metrics measure whether the Machine is producing better analysis: calibration and Brier skill across the book, the frequency with which the aporia test blocks an overstated reweight, and the proportion of probes whose expected information gain exceeded a stated threshold before they were run.

Counterfactual metrics measure whether the machinery is earning its cost: rebuild historical Files under v3.0 and count how many published conclusions change, then adjudicate each change on inspection. If very few conclusions change, the machinery is expensive ceremony. If many change and the changes do not survive inspection, the machinery is introducing error. The programme's justification rests on this metric more than on any other.

12.6 Deliberate deferrals

Five things are consciously not being built, each with the trigger that would promote it.

Automatic argument extraction from long documents is deferred until a conformance-tested extractor reaches acceptable agreement with analyst labelling on the existing File corpus. Full ASPIC+ with strict and defeasible rule sets is deferred until extraction is solved, since the formalism is worthless over a badly extracted graph. Gradual and probabilistic argumentation semantics are deferred until grounded semantics demonstrably fails by leaving too much UNDECIDED. Multi-model dialectical generation with enforced independence is deferred until the single-generator pipeline is stable, since the literature does not support expecting adjudication benefits from it. And a natural-language conversational surface over the Inquiry is deferred entirely until the underlying state is trustworthy, because a fluent interface over an untrustworthy substrate is the most dangerous artifact this programme could ship.

12.7 Kill signals for the programme

Four observations would end or redesign this build.

Replay cannot be demonstrated on a real historical File after DK1, which would mean the provenance model is wrong at the foundation. Grounded labelling leaves the overwhelming majority of material claims UNDECIDED across several Files, which would mean the adjudication layer produces no signal and the coupling rule degenerates. Rebuilt historical Files show no conclusions changing under v3.0, which would mean the examination layer is ceremony and the v2.0 spine alone was sufficient. Or analyst time per File cycle rises materially without a corresponding movement in the calibration record, which would mean the instrument has become a tax rather than an edge.

Part Thirteen: Open questions

#QuestionOwnerBlocking
1Can the δ values in the coupling rule be derived rather than stipulated, for instance from epistemic probabilistic argumentation, or must they be calibrated empirically against the File corpus?ResearchBlocks a formal claim, not the build
2Should the aporia test threshold be a fixed Bayes factor of 3.0, or should it scale with the number of branches, since a three-branch File and a six-branch File have different accidental-proximity rates?ResearchBlocks publication of the test as canon
3Does grounded semantics leave too much UNDECIDED for the coupling rule to carry signal, and if so is preferred semantics or a gradual semantics the right replacement?Build, DK2 gateBlocks DK3
4How is the provenance cluster assigned when two sources share an unobserved common cause rather than a common outlet, which the current mechanism cannot detect?ResearchDoes not block, but bounds the redundancy correction's value
5Should the constitution be File-specific, venture-specific, or a single 451 Labs constitution with File-level extensions?SurendraBlocks DK3
6Is the Kriyas Wheel, as used in the two absorbed research notes, the same construct as the Kriyas revision discipline, or do the two need distinct names?SurendraBlocks the cross-link back into the Kriyas material
7Does the WIN Machine evidence object conform to the Clearstory eight-object model as written, or does one of the eight need extension?Surendra with Clearstory canonBlocks the reciprocal cross-link
8What is the correct m and λ for a File whose market clock is regulatory, where cycles are lumpy and a quarterly cadence may be wrong?ResearchDoes not block; current values are conventions
9Should the Machine publish its aporia ledger in the public File, or hold it internally, given that an open aporia is both intellectually honest and strategically informative to competitors?SurendraBlocks the File template revision

Appendix A: Notation

SymbolMeaning
IAn Inquiry, the sixteen-part state tuple of Section 5.2
αDirichlet pseudo-count vector over branches
α₀Sum of pseudo-counts, the prior strength
êNormalized evidence weight vector for a cycle
λForgetting factor in the cross-cycle update
mNew evidence strength added per cycle
τ_gBase temper exponent from source grade
δDialectical modulation factor from argument status
ρRedundancy factor from provenance clustering
τ_effEffective exponent, the product of the three above
L(s|b)Analyst likelihood of signal s under branch b, relative zero-to-one scale
BFBayes factor, a ratio of evidence weights computed before the prior
HShannon entropy of the branch distribution, in bits
IN, OUT, UNDECIDEDGrounded semantics labels

Appendix B: The complete worked arithmetic

Every figure in Part Eight, collected for checking. Prior 0.40, 0.30, 0.30 with α₀ = 10.

Tempered likelihoods: S1 at τ 1.0 gives 0.8000, 0.3000, 0.2500. S2 at τ 0.6 gives 0.5327, 0.8415, 0.5327. S3 at τ 0.2 gives 0.7860, 0.9311, 0.7860.

Evidence weights: 0.33493, 0.23506, 0.10467. Posterior: 56.79, 29.89, 13.31 percent.

Bayes factors: incumbent to null 3.200, insurgent to null 2.246, incumbent to insurgent 1.425.

Coupling, varying S2's supporting argument status: IN gives 56.79, 29.89, 13.31 with a factor of 1.42; UNDECIDED gives 60.49, 25.33, 14.18 with 1.79; OUT gives 63.79, 21.26, 14.95 with 2.25.

Impression separation of S3 to likelihoods 0.45, 0.55, 0.40: evidence weights 0.36322, 0.22399, 0.11086, posterior 59.12, 27.34, 13.53, factor 1.62.

Redundancy: naive two-signal posterior on X is 68.80 percent from weights 0.7081 and 0.3211; corrected is 64.41 percent from 0.7719 and 0.4265; recovery 4.39 points.

Floor: raw 12.4, 87.2, 0.4 with 4.60 points added; result 11.8, 83.2, 5.0.

Sensitivity on S1: 0.70 gives 53.5 percent with factor 1.25; 0.80 gives 56.8 with 1.42; 0.90 gives 59.7 with 1.60.

Cross-cycle at λ 0.8 and m 5: steady state α₀ = 25.0; five-cycle trajectory of α₀ 13.00, 15.40, 17.32, 18.86, 20.08 with incumbent prior mean 43.7, 45.6, 46.8, 47.5, 48.1 percent.

Expected information gain: prior entropy 1.3716 bits; affirmative outcome probability 0.6107 leading to 79.0, 12.2, 8.7 at entropy 0.9459; negative outcome probability 0.3893 leading to 21.9, 57.6, 20.5 at entropy 1.4070; expected posterior entropy 1.1254; gain 0.2462 bits. Weak probe gain 0.0057 bits.

Calibration: Brier values 0.2937, 0.1850, 1.0850, 0.1400, 0.3350 for a mean of 0.4077 against a uniform reference of 0.6667, giving a skill score of 0.3884; log scores 0.8160, 0.6215, 2.7370, 0.5146, 0.8625 for a mean of 1.1103 bits.

Appendix C: Glossary

Aporia. A persisted state in which examination has shown that a framing or a set of commitments cannot currently be maintained together, and no alternative has earned support. A first-class asset, not a failure.

Assent. The Stoic-derived belief state assigned to a claim: accept, provisional, suspend, contested, or reject. Distinct from dialectical status and from epistemic type.

Branch. One of the mutually exclusive ways the binding constraint could relocate. Every File carries a residual or null branch that may never hold zero mass.

Coupling rule. The rule binding argument status to evidentiary weight, τ_eff = τ_g × δ × ρ. The one formal claim this architecture makes beyond assembly.

Dialectical status. The grounded-semantics label of an argument: IN, OUT, or UNDECIDED. Computed from the graph, never asserted.

Elenchus. Socratic examination proceeding from the interlocutor's own commitments to show that a thesis and those commitments cannot all be maintained. Implemented as an operator that must emit its derivation.

Impression. An external input before assent. Decomposed into observation, attribution, interpretation, forecast, and recommendation, of which only the first two are directly evidentiary.

Inquiry. The primary computational object and the full persisted state of one line of investigation. A File is a view over an Inquiry.

Kill condition. A pre-specified falsifying observation attached to a necessary and vulnerable assumption, with a named observable source. Firing one forces thesis review.

Temper exponent. The fractional power to which a likelihood is raised before entering the update, following the generalized Bayes literature. A per-source learning rate, not a fudge factor.

Unguarded Flank. An open strategic position that passes all three gates of structural block, constraint rather than feature, and compounding before arrival.

Acknowledgements

The philosophical architecture in Parts Two, Six, and Seven originated in two 451 Labs research notes written in August 2026, which this document absorbs and retires. Vamsi Koduru's long collaboration on governed reasoning architecture shaped the separation between generation and adjudication that Part Four turns into a rule. The quantitative spine was built and hardened across the WIN File series, and every worked number in Part Eight was recomputed for this document rather than carried forward on trust.

A Note on AI-Assisted Research and Intelligence Provenance

Research for this document was conducted with AI assistance across the argumentation, belief revision, decision-analysis, and machine-learning literatures, and the arithmetic in Part Eight was executed rather than asserted. Astra and Clearstory are 451 Labs theses under active development. Neither is a shipped product, and neither was used in the preparation of this document; where they appear, they appear as the theses under which a future realization of this specification would sit. All analytical judgments, including the correction to the version 2.0 worked example, the coupling rule, and the architectural verdicts, remain the author's own.

References

Agrawal, A., Gans, J., & Goldfarb, A. (2018). Prediction machines: The simple economics of artificial intelligence. Harvard Business Review Press.

Alchourrón, C. E., Gärdenfors, P., & Makinson, D. (1985). On the logic of theory change: Partial meet contraction and revision functions. Journal of Symbolic Logic, 50(2), 510–530.

Athey, S., Tibshirani, J., & Wager, S. (2019). Generalized random forests. Annals of Statistics, 47(2), 1148–1178.

Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., … Kaplan, J. (2022). Constitutional AI: Harmlessness from AI feedback (arXiv:2212.08073). arXiv.

Baroni, P., Caminada, M., & Giacomin, M. (2011). An introduction to argumentation semantics. Knowledge Engineering Review, 26(4), 365–410.

Bissiri, P. G., Holmes, C. C., & Walker, S. G. (2016). A general framework for updating belief distributions. Journal of the Royal Statistical Society: Series B, 78(5), 1103–1130.

Brandenburger, A. M., & Nalebuff, B. J. (1996). Co-opetition. Doubleday.

Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3.

Brodersen, K. H., Gallusser, F., Koehler, J., Remy, N., & Scott, S. L. (2015). Inferring causal impact using Bayesian structural time-series models. Annals of Applied Statistics, 9(1), 247–274.

Cemri, M., Pan, M. Z., Yang, S., Agrawal, L. A., Chopra, B., Tiwari, R., Keutzer, K., Parameswaran, A., Klein, D., Ramchandran, K., Zaharia, M., Gonzalez, J. E., & Stoica, I. (2025). Why do multi-agent LLM systems fail? (arXiv:2503.13657). arXiv.

Conklin, J., & Begeman, M. L. (1988). gIBIS: A hypertext tool for exploratory policy discussion. ACM Transactions on Information Systems, 6(4), 303–331.

de Kleer, J. (1986). An assumption-based TMS. Artificial Intelligence, 28(2), 127–162.

Dewar, J. A. (2002). Assumption-based planning: A tool for reducing avoidable surprises. Cambridge University Press.

Doyle, J. (1979). A truth maintenance system. Artificial Intelligence, 12(3), 231–272.

Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., & Mordatch, I. (2024). Improving factuality and reasoning in language models through multiagent debate. In Proceedings of the 41st International Conference on Machine Learning (Vol. 235, pp. 11733–11763). PMLR.

Dung, P. M. (1995). On the acceptability of arguments and its fundamental role in nonmonotonic reasoning, logic programming and n-person games. Artificial Intelligence, 77(2), 321–357.

Epictetus. (n.d.). Enchiridion. In Internet Encyclopedia of Philosophy. University of Tennessee at Martin.

Farquhar, S., Kossen, J., Kuhn, L., & Gal, Y. (2024). Detecting hallucinations in large language models using semantic entropy. Nature, 630(8017), 625–630.

Goldratt, E. M., & Cox, J. (1984). The goal: A process of ongoing improvement. North River Press.

Gordon, T. F., Prakken, H., & Walton, D. (2007). The Carneades model of argument and burden of proof. Artificial Intelligence, 171(10–15), 875–896.

Grünwald, P., & van Ommen, T. (2017). Inconsistency of Bayesian inference for misspecified linear models, and a proposal for repairing it. Bayesian Analysis, 12(4), 1069–1103.

Hamilton, J. D. (1989). A new approach to the economic analysis of nonstationary time series and the business cycle. Econometrica, 57(2), 357–384.

Heuer, R. J., Jr. (1999). Psychology of intelligence analysis. Center for the Study of Intelligence, Central Intelligence Agency.

Heuer, R. J., Jr., & Pherson, R. H. (2011). Structured analytic techniques for intelligence analysis. CQ Press.

Holmes, C. C., & Walker, S. G. (2017). Assigning a value to a power likelihood in a general Bayesian model. Biometrika, 104(2), 497–503.

Howard, R. A. (1966). Information value theory. IEEE Transactions on Systems Science and Cybernetics, 2(1), 22–26.

Hunter, A., & Thimm, M. (2017). Probabilistic reasoning with abstract argumentation frameworks. Journal of Artificial Intelligence Research, 59, 565–611.

Irving, G., Christiano, P., & Amodei, D. (2018). AI safety via debate (arXiv:1805.00899). arXiv.

Kass, R. E., & Raftery, A. E. (1995). Bayes factors. Journal of the American Statistical Association, 90(430), 773–795.

Khan, A., Hughes, J., Valentine, D., Ruis, L., Sachan, K., Radhakrishnan, A., Grefenstette, E., Bowman, S. R., Rocktäschel, T., & Perez, E. (2024). Debating with more persuasive LLMs leads to more truthful answers (arXiv:2402.06782). arXiv.

Kuhn, L., Gal, Y., & Farquhar, S. (2023). Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In Proceedings of the Eleventh International Conference on Learning Representations.

Kunz, W., & Rittel, H. W. J. (1970). Issues as elements of information systems (Working Paper No. 131). Institute of Urban and Regional Development, University of California, Berkeley.

Lawrence, J., & Reed, C. (2020). Argument mining: A survey. Computational Linguistics, 45(4), 765–818.

Lempert, R. J. (2019). Robust decision making (RDM). In V. A. W. J. Marchau, W. E. Walker, P. J. T. M. Bloemen, & S. W. Popper (Eds.), Decision making under deep uncertainty: From theory to practice (pp. 23–51). Springer.

Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems 30 (pp. 4765–4774).

Marcus Aurelius. (n.d.). Meditations (G. Long, Trans.). Internet Classics Archive, Massachusetts Institute of Technology.

Mellers, B., Ungar, L., Baron, J., Ramos, J., Gurcay, B., Fincher, K., Scott, S. E., Moore, D., Atanasov, P., Swift, S. A., Murray, T., Stone, E., & Tetlock, P. E. (2014). Psychological strategies for winning a geopolitical forecasting tournament. Psychological Science, 25(5), 1106–1115.

Miller, J. W., & Dunson, D. B. (2019). Robust Bayesian inference via coarsening. Journal of the American Statistical Association, 114(527), 1113–1125.

Modgil, S., & Prakken, H. (2014). The ASPIC+ framework for structured argumentation: A tutorial. Argument & Computation, 5(1), 31–62.

Office of the Director of National Intelligence. (2015). Intelligence Community Directive 203: Analytic standards. ODNI.

Plato. (n.d.). Apology (B. Jowett, Trans.). Internet Classics Archive, Massachusetts Institute of Technology.

Plato. (n.d.). Euthyphro (B. Jowett, Trans.). Internet Classics Archive, Massachusetts Institute of Technology.

Pollock, J. L. (1987). Defeasible reasoning. Cognitive Science, 11(4), 481–518.

Porter, M. E. (1979). How competitive forces shape strategy. Harvard Business Review, 57(2), 137–145.

Prakken, H. (2010). An abstract framework for argumentation with structured arguments. Argument & Computation, 1(2), 93–124.

Priest, G. (1979). The logic of paradox. Journal of Philosophical Logic, 8(1), 219–241.

Rittel, H. W. J., & Webber, M. M. (1973). Dilemmas in a general theory of planning. Policy Sciences, 4(2), 155–169.

Schoemaker, P. J. H. (1995). Scenario planning: A tool for strategic thinking. Sloan Management Review, 36(2), 25–40.

Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., Kravec, S., Maxwell, T., McCandlish, S., Ndousse, K., Rausch, O., Schiefer, N., Yan, D., Zhang, M., & Perez, E. (2024). Towards understanding sycophancy in language models. In Proceedings of the Twelfth International Conference on Learning Representations.

Stanford Encyclopedia of Philosophy. (2023). Stoicism. Stanford University.

Stanford Encyclopedia of Philosophy. (n.d.). Epictetus. Stanford University.

Stanford Encyclopedia of Philosophy. (n.d.). Plato's shorter ethical works. Stanford University.

Tetlock, P. E., & Gardner, D. (2015). Superforecasting: The art and science of prediction. Crown.

Toulmin, S. E. (1958). The uses of argument. Cambridge University Press.

Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting. In Advances in Neural Information Processing Systems 36 (pp. 74952–74965).

Vlastos, G. (1983). The Socratic elenchus. Oxford Studies in Ancient Philosophy, 1, 27–58.

Wack, P. (1985). Scenarios: Uncharted waters ahead. Harvard Business Review, 63(5), 73–89.

Wager, S., & Athey, S. (2018). Estimation and inference of heterogeneous treatment effects using random forests. Journal of the American Statistical Association, 113(523), 1228–1242.

Walton, D., Reed, C., & Macagno, F. (2008). Argumentation schemes. Cambridge University Press.

Wang, Q., Wang, Z., Su, Y., Tong, H., & Song, Y. (2024). Rethinking the bounds of LLM reasoning: Are multi-agent discussions the key? In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vol. 1, pp. 6106–6131). Association for Computational Linguistics.

Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., & Zhou, D. (2023). Self-consistency improves chain of thought reasoning in language models. In Proceedings of the Eleventh International Conference on Learning Representations.

Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems 35 (pp. 24824–24837).

World Wide Web Consortium. (2013). PROV-O: The PROV ontology (W3C Recommendation). W3C.