The Golden Hall Test: Meaningful Agency Under Machine-Mediated Evidence Terrain

    Breyden E. Taylor · Prompted LLC · 24 July 2026

    cs.HC · cs.AI · cs.CY · 9 pages · 2,476 words · depends on P01, P03, P05 · gates P10, P11

    Abstract

    Human presence in a workflow does not establish meaningful control. A person may approve a decision after the system selected the evidence, named the categories, ordered the alternatives, set the default, framed confidence, and made refusal operationally expensive. We define the Golden Hall Test as a reason-responsive agency evaluation over source visibility, independent formation, challenge, delay, declaration, reversal, refusal, vocabulary independence, and correction reach. The method distinguishes presence, approval, causal contribution, authority, and liability. We provide a preregistered repeated-decision study and an AgencyReceipt schema that records which rights were actually available. Synthetic pilots motivate the design but are not treated as human evidence.

    Claim register

    • P06-C1 — formal/proposed · reopens: new replication, falsifier, source correction, or configuration change
    • P06-C2 — testable hypothesis · reopens: new replication, falsifier, source correction, or configuration change
    • P06-C3 — implemented-or-pilot bounded · reopens: new replication, falsifier, source correction, or configuration change
    • P06-C4 — external validation required · reopens: new replication, falsifier, source correction, or configuration change
    pdf sha256
    1f22ae26a46634da956c32a69f769be8ec7118565199de7b54046346358d8562
    src sha256
    46c6d8e31bd0ac10bfe0814498311c9144a952f82cc5bdcfabdcee25404fdd29
    md sha256
    4b79e09d2a39249f07cc27f4d0ddabe437e1a05844b8a5e249fcc58751544646
    Reading scale

    Proposed arXiv categories: cs.HC / cs.AI / cs.CY Publication DAG node: P06, Wave B Dependencies: P01, P03, P05 Claim register: P06-C1, P06-C2, P06-C3, P06-C4

    1. Thesis and scope

    The most dangerous surrender of agency is the one that still looks like command. Meaningful control requires real causal rights and correction reach, not a final click.

    This paper operationalizes meaningful human control for machine-mediated evidence terrain. It adds a reconstructability test: could the human or institution recover the decision from receipts without the original model? It also tests vocabulary capture and responsibility diffusion, two mechanisms often hidden by accuracy-only evaluation.

    FORKED is an internally owned and managed estate inside the Ubiquity Federation, Prompted LLC's flagship governance substrate. The estate inherits the Fractal Quivers of Quivers parent topology, production splat mechanics, center exclusion, receipt-bearing lawful traversal, lifecycle separation, and the rule that models propose morphisms while governance determines admissible motion [@Taylor2026FQoQ; @Taylor2026OpenCenter; @Prompted2026Ubiquity]. This paper studies one bounded expression of that parent architecture. It neither renames the parent substrate nor inherits universal applicability from it.

    The evidence lanes are kept separate. A canonical or source artifact establishes what that artifact states. A component test establishes behavior under its recorded configuration. A synthetic campaign establishes behavior of the generator and analysis pipeline, not human, hardware, customer, field, or operational performance. Novel cross-domain claims are preregistered as hypotheses and remain open to null, adverse, localized, and disproving results.

    Dependency freeze. Final freeze depends on P01, P03, P05; draft work may proceed in parallel, but imported claim IDs and artifact hashes must be rechecked after each dependency freezes.

    2. Interlattice position

    The interlattice is the citation and governance relationship among the parent FQoQ topology, production splat mechanics, the FORKED foundational architecture, the Native Range, the Experimental Program, and the companion papers in this DAG. Citations are therefore typed: a parent citation supplies inherited architecture; a native citation supplies component behavior; a synthetic citation supplies generator-level evidence; a companion citation supplies a versioned specialized argument. No downstream citation converts an upstream hypothesis into a universal fact.

    This paper's claims are:

    • P06-C1: the central mechanism is formally representable and testable.
    • P06-C2: the specified controls can distinguish the mechanism from simpler alternatives.
    • P06-C3: the recorded implementation or synthetic evidence establishes only its declared lane.
    • P06-C4: external applicability requires the preregistered receipts and remains domain-bound.

    3. Problem statement

    Modern AI systems frequently compress unlike objects into a common output surface. Facts, hypotheses, instructions, permissions, model predictions, runtime observations, and institutional decisions can all arrive as fluent text or a single status field. The resulting failure is architectural before it is linguistic: a local representation acquires authority that its source never possessed. The most dangerous surrender of agency is the one that still looks like command. Meaningful control requires real causal rights and correction reach, not a final click.

    4. Parent architecture

    Reality-mutation boundary

    A machine can render, recommend, rank, and predict. A human or human-authorized institution converts some outputs into standing, resource movement, or durable record. This makes the human a causal boundary, but not automatically a meaningful controller.

    A nominal human-in-the-loop design can leave the person with ceremonial responsibility after the system has already selected the evidence, chosen the categories, ordered the options, established urgency, defined confidence, and written the memory the next decision will inherit. Classical automation research has long warned that removing routine control can leave humans responsible for rare situations in which their skills and situation awareness have degraded. [@Bainbridge1983Ironies; @Parasuraman2000Automation]

    FORKED therefore separates human presence, acknowledgment, approval, meaningful agency, causal contribution, and liability.

    A meaningful agency envelope includes source and transformation visibility; uncertainty and alternative visibility; time appropriate to the responsibility class; ability to request complements; ability to challenge, refuse, escalate, delay, and reverse where lawful; ability to correct dependent history; freedom from a coercive or hidden default; vocabulary independence; and clear liability ownership.

    This extends meaningful-human-control work on tracking and tracing. A system should respond to the relevant reasons of relevant human actors, and potentially dangerous events should remain attributable through a human and institutional chain. [@Mecacci2020MHC; @DeSio2023MHC]

    Agency and liability

    An AgencyReceipt records what the human actually possessed at decision time:

    AgencyReceipt = {
      actor,
      role,
      authority_scope,
      evidence_visible,
      sources_inspected,
      alternatives_visible,
      recommendation_default,
      time_available,
      challenge_available,
      delay_available,
      reversal_available,
      correction_available,
      consequences_disclosed,
      confidence_entered_before_advice,
      confidence_entered_after_advice,
      decision_changed,
      liability_owner,
      system_designers,
      approving_institution,
      runtime_identity,
      receipt_time
    }

    Liability does not follow the last click automatically. If the surrounding system constrained the human's reachable options, the system designers, policy authors, model providers, data stewards, approvers, and institution remain in the causal graph. A person can be authorized to make a decision while lacking the information or time needed to control the reasons the system was responding to.

    This is not a method for dissolving individual responsibility. It is a method for refusing false transfer. Responsibility remains distributed where causality, authority, and ability to correct were distributed. Work on responsibility loci similarly warns against attributing independent agency to automation in ways that erase the collaboration that produced the outcome. [@Nyholm2018Agency]

    Appropriate reliance

    FORKED does not optimize trust upward. It optimizes reliance toward the state in which humans accept correct assistance, reject incorrect assistance, and preserve correction across repeated interaction.

    Prediction sets and cognitive forcing functions illustrate two relevant mechanisms. Prediction sets can expose uncertainty and alternatives rather than presenting one answer, and studies have found benefits in some human decision tasks. They can also create larger cognitive loads and disparate effects. [@Cresswell2024Conformal; @Babbar2022PredictionSets; @Cresswell2024Disparate] Cognitive forcing functions can reduce overreliance by requiring independent thought, but the added friction is not free. [@Bucinca2021Forcing]

    FORKED's branch-preserving interface is therefore responsibility-scaled. It can require an independent estimate before advice for high-consequence tasks, while allowing compact branch presentation for reversible work. It records both premature closure and trauma-driven non-closure. It does not celebrate caution that prevents a lawful declaration after the threshold is met.

    The trust criterion is longitudinal: the path must survive correction, counterevidence, scope change, authority challenge, timing faults, dependency discovery, and repeated movement toward telos.

    Deterministic and probabilistic roles

    Probabilistic models are powerful at perception, synthesis, search, and proposal. Their variability and failure modes make them poor sole authorities for protected state mutation.

    FORKED separates two planes:

    PROBABILISTIC PLANE
      fusion
      classification
      anomaly detection
      observer modeling
      branch generation
      counterfactuals
      narrative projection
    
    DETERMINISTIC AUTHORITY PLANE
      source authentication
      runtime identity
      time and sequence
      schema validation
      authority predicates
      state-transition admission
      protected-ledger write
      rollback and correction fan-out

    The separation can be physical, logical, or both. The key property is that a model cannot grant itself the transition it proposes. A deterministic plane also does not become sovereign merely because it is deterministic. Its rules, keys, clocks, and configuration require provenance and lifecycle validation.

    The first synthetic fault campaign found that separated authority reduced encoded unauthorized transitions by approximately 0.136 and protected-ledger contamination by approximately 0.123 relative to a unified baseline, while imposing availability and complexity costs. [@ForkedExperiments2026] The parent interpretation is not separate everything maximally. It is select the boundary according to responsibility.

    Branch-preserving pilot

    The synthetic branch interface reduced encoded premature closure by approximately 0.039 relative to the point-estimate interface and increased source inspection by approximately 0.252 and challenge by approximately 0.311. It also added about 17.5 seconds of encoded decision time and did not outperform the no-AI synthetic baseline on lawful closure. [@ForkedExperiments2026]

    A shallow conclusion would be that branch preservation is too slow. An equally shallow conclusion would be that any reduction in premature closure justifies the cost. The parent interpretation is responsibility-indexed branch budgeting:

    • preserve full branches for irreversible or constitutional transitions;
    • use progressive disclosure for material but reversible work;
    • compress low-consequence branches after receipt and rollback conditions are met;
    • measure both premature closure and harmful delay;
    • include human attention as an economic resource.

    5. Formal model

    An AgencyReceipt records the set of decision rights present at the consequential transition. Let R={r1,,rk} include independent formation, source inspection, challenge, delay, declaration, reversal, refusal, and escalation. Let ai{0,1,degraded} denote whether right i was practically exercisable. Meaningful agency is not the sum of rights; some are hard requirements for specific responsibility classes. A high-consequence declaration with no source inspection or reversal cannot be relabeled human-controlled because approval was recorded.

    Reason-responsiveness is tested counterfactually: when the reasons change while interface pressure is held fixed, does the human decision change appropriately? When only the interface default changes, does the decision remain stable? Liability geometry traces causal contribution across model provider, integrator, policy author, interface designer, institution, operator, and runtime configuration. The receipt does not allocate legal liability by itself. It prevents the causal chain from being erased.

    6. System and study design

    Participants complete repeated decisions with manipulated source access, pre-advice independent estimate, branch visibility, default strength, explanation fluency, challenge cost, time budget, reversal capability, and liability disclosure. AI correctness varies independently. The study records gaze or interface attention only with explicit consent and data minimization; a no-gaze protocol remains available.

    The core contrast separates nominal approval from causal agency. In nominal conditions, the participant can approve but cannot inspect or alter the action path. In causal conditions, the participant can view sources, request complements, challenge classification, delay within an envelope, and reopen downstream state. A third condition provides rights formally but makes them operationally costly, testing the difference between nominal and practical availability.

    6.1 Controls and ablations

    • No AI
    • AI before independent estimate
    • AI after independent estimate
    • Nominal rights
    • Practical rights
    • Strong/weak defaults
    • Source panel present/absent
    • Reversal present/absent
    • Liability disclosure present/absent

    6.2 Primary outcomes

    • Reason-responsive decision change
    • Appropriate reliance
    • Challenge quality
    • Source inspection
    • Decision reconstructability
    • Vocabulary convergence
    • Correction survival
    • Perceived versus actual control

    7. Current implementation and pilot evidence

    The E01 synthetic program generated 96,000 decision traces across 8,000 simulated participants. Branch-preserving advice increased modeled challenge and source inspection while adding time; the no-AI condition produced the highest lawful closure in the encoded generator [@ForkedExperiments2026]. The native threshold campaign blocks nominal-only approval and absent authority under its test frames [@ForkedCampaign2026]. These receipts justify the measurement architecture, not claims about people.

    The paper treats this evidence directly where capability is established and narrowly where it is not. Implemented code paths are called implemented. Passing component tests are called passing component tests. Synthetic estimates remain synthetic. Human and field effects remain hypotheses until their own evidence exists.

    8. Preregistration

    The formal preregistration artifact distributed with this paper freezes the primary claims, outcomes, controls, exclusion rules, analysis family, null interpretation, and release boundary before external data collection. The core sequence is:

    1. freeze source and configuration manifests;
    1. freeze primary hypotheses and adverse outcomes;
    1. generate or acquire data without changing the admission gate;
    1. run the declared analysis and publish all primary results;
    1. route deviations to an explicit exploratory appendix;
    1. demote, localize, or reject claims when falsifiers fire.

    8.1 Analysis plan

    Agency dimensions are analyzed as predictors of behavior, not combined into a moral score. We preregister a latent-variable model only as secondary; primary results report each right and outcome separately. Reason-responsiveness is estimated from within-person counterfactual pairs. Perceived control is compared with actual causal capability to identify ceremonial agency. Correction survival is measured after an initial decision has already propagated, ensuring that reversal is not merely a button but an effective path.

    8.2 Null and adverse-result handling

    If the agency dimensions do not predict appropriate decision change or correction, the Golden Hall construct is revised rather than defended by participant satisfaction. If rights increase workload without benefit, the interface claim is localized. Perceived control without causal rights is not counted as agency. Refusal and delay are not errors when the preregistered evidence or authority condition is absent.

    9. Falsifiers

    • Nominal approval predicts outcomes as well as causal rights
    • AgencyReceipt dimensions lack behavioral validity
    • Vocabulary capture does not relate to reliance
    • Correction rights do not improve correction survival
    • Rights create unacceptable delay across consequence classes

    10. Limitations

    Labor relationships, command structures, and organizational retaliation can determine whether a technical right is real. A laboratory interface cannot reproduce all such pressures. The framework must therefore be combined with institutional and legal analysis rather than substituted for it.

    11. Security, ethics, and release boundary

    This public paper is limited to benign assurance, provenance, human-agency preservation, lawful state transition, simulation, test infrastructure, and defensive resilience. It does not disclose operational target-selection logic, engagement optimization, platform-specific exploitation thresholds, live deception procedures, signature-emitter recipes, or methods that materially increase harmful capability. Any later operational research requires separate lawful authority, ethics and safety review, configuration control, and release adjudication.

    12. Required next receipts

    • IRB approval
    • Worker and operator co-design
    • Domain-expert replication
    • Organizational field study
    • Legal responsibility analysis

    13. Conclusion

    The most dangerous surrender of agency is the one that still looks like command. Meaningful control requires real causal rights and correction reach, not a final click. The contribution is not a claim that every domain should adopt one representation, threshold, interface, or governance stack. It is a falsifiable architecture and experiment package for determining where this mechanism carries load, where it inverts, and where a simpler system should win.

    Data, code, and provenance availability

    The submission bundle includes the manuscript source, compiled PDF, bibliography, preregistration, claim register, controls and null-handling document, release boundary, source manifest, configuration digest, dependency contract, figure source, and arXiv source archive. Internal source artifacts are cited by immutable digest where available. Restricted operational material is not included.

    Acknowledgments and authorship

    Breyden E. Taylor is the author and bears responsibility for the thesis, terminology, claims, boundaries, and decision to publish. Homeskillet assisted with source synthesis, drafting, artifact generation, typesetting, and build inspection. The Ubiquity Federation supplied prior artifacts, runtime implementations, and receipts. AI assistance is production history, not evidence.

    References

    PREPRINT — this is an author preprint of a paper scheduled for arXiv submission. It has not been peer reviewed and carries no arXiv identifier until announcement. Claims are preregistered and domain-bound; no paper inherits universal applicability from the parent whitepaper or from a companion paper.