Technical whitepaper 1.0 · Review edition · September 2026

    The Verdict Before the Evidence

    Promptfoo, refusal short-circuits, and the governance of AI assurance

    Breyden E. Taylor — Founder & Architect, Prompted LLC

    Core claim

    An evaluator must not promote the appearance of a safety behavior into evidence that the safety obligation was satisfied.

    Wide architectural frontispiece showing prompt, context, intent, and risk moving through evidence, analysis, verification, reasoning, and conclusion.
    Atrium plateThe full assurance path. Prompt, context, intent, and risk enter an evidence-bearing sequence before any trusted outcome is admitted.

    Released 2026-09-13 · review edition

    Public discovery signed by the architect on 2026-09-13. Release admits the review edition to the machine graph; it does not upgrade its evidentiary standing.

    Disclosure status — Upstream maintainer disclosure has not been performed as part of this edition. A maintainers' review packet carrying the exact source pin, the protected custom-policy exception, a real-package regression test, and a bounded account of likely impact remains outstanding.

    The same invariant at two scales. Autonomy is not granted because a model sounds confident or passes a demo; it is earned through traceable behavior, bounded scope, judgment, and outcome history. PASS is not granted because output looks like a refusal; it is earned by the evidence. Appearance is not authority, a signal is not a receipt, a receipt is not a verdict, and execution is not admission.

    Standing

    A peer paper in the flat register. It applies the evidence-authority argument to a public evaluation harness; it sits above nothing and nothing sits above it.

    The decision contract draws on SPLAT constraint geometry from Computing Around the Open Center and on the admission/observation separation developed in Proof-Horizon Sharding. Those are first-party architecture sources, not independent validation of this finding.

    Technical whitepaper 1.0, review edition. A pinned-source analysis plus an executed, isolated mechanism-level reproduction. Not a live foundation-model jailbreak, not a hosted-service penetration test, and not access to any lab's private evaluation infrastructure.

    Words
    6,100
    Rendered pages
    16
    Sections
    12
    References
    20

    The pinned read

    Every source claim is bound to one commit, read GitHub connector; official public webpages; ChatGPT Library. A later upstream change does not silently rewrite this finding; it produces a new epoch.

    Repository
    promptfoo/promptfoo
    Commit
    32bfa9edae56d36f526ab13b6637db8dfb82c73e
    Commit timestamp (UTC)
    2026-09-12T23:38:05Z
    Files pinned
    8
    • src/redteam/plugins/base.tsblob 3c9af33cb0e0…
    • src/redteam/util.tsblob 8089bc7da48a…
    • src/redteam/plugins/policy/index.tsblob 177e05a0b2a3…
    • src/redteam/plugins/bola.tsblob 6a4665122594…
    • src/assertions/redteam.tsblob 635113376e68…
    • src/assertions/guardrails.tsblob 3faf841af094…
    • src/matchers/rubric.tsblob 718b65fe6fb5…
    • LICENSEblob af3fa111d930…

    Executed mechanism harness

    Isolated transcribed-source mechanism harness; not an upstream package, provider, model, CLI, or full plugin integration run. Run 2026-09-13T02:43:58.344Z on v22.16.0. Network: none. Model calls: 0. Upstream package executed: no.

    Assertions
    45
    Base fixtures
    13
    False passes
    5
    Model calls
    0
    The fixture-by-fixture ledger
    Square Prompted Upgrades Promptfoo poster with a suspended verdict seal and evidence papers.
    Review packet plateSquare campaign plate for the review packet: the verdict must follow the evidence.

    Honest negatives

    Results, not disclaimers. A paper about evidence authority does not get to overclaim its own.

    • No live exploitation was observed

      No private lab deployment, unauthorized data access, harmful completion, or in-the-wild exploitation was seen. The evidence is a pinned source read plus an isolated mechanism harness.

    • The custom-policy grader is protected by default

      The policy grader forwards a skip flag that defaults to true, so its refusals still reach policy grading. The finding is not a blanket claim about every Promptfoo assertion.

    • No prevalence, severity, or affected-version claim

      No CVE status, severity score, version interval, or frequency estimate is assigned. Source inspection establishes a branch; it does not establish how often deployments meet it.

    • The harness is not an upstream integration test

      Transcribed mechanisms were executed in isolation with zero model calls. The full package, provider graph, CLI, cache, and aggregate reporting were not run.

    The repair contract

    1. 01

      Separate four objects: observation, evaluation coverage, policy assessment, and admission authority.

    2. 02

      Preserve missingness — an absent signal is not a clean signal, and a rendered zero must carry its cause.

    3. 03

      Bind every receipt to the subject it actually witnessed, at the evidence level it actually reached.

    4. 04

      Repair the shortest causal path first: no deterministic branch may issue a verdict the judge never reached.

    5. 05

      Carry correction across proof horizons, so a later reversal updates lineage instead of erasing it.

    Section map

    Peer works it leans on

    Sealed artifacts

    Rows marked held are cited but not published here. Write breyden@prompted.community to request access to a cited private artifact.

    Recommended citation

    Taylor, B. E. (2026). The Verdict Before the Evidence: Promptfoo, refusal short-circuits, and the governance of AI assurance (Technical Whitepaper 1.0, Review Edition). Prompted LLC.