Technical whitepaper 1.0 · Review edition · September 2026
The Verdict Before the Evidence
Promptfoo, refusal short-circuits, and the governance of AI assurance
Breyden E. Taylor — Founder & Architect, Prompted LLC
Core claim
An evaluator must not promote the appearance of a safety behavior into evidence that the safety obligation was satisfied.

Released 2026-09-13 · review edition
Public discovery signed by the architect on 2026-09-13. Release admits the review edition to the machine graph; it does not upgrade its evidentiary standing.
Disclosure status — Upstream maintainer disclosure has not been performed as part of this edition. A maintainers' review packet carrying the exact source pin, the protected custom-policy exception, a real-package regression test, and a bounded account of likely impact remains outstanding.
The same invariant at two scales. Autonomy is not granted because a model sounds confident or passes a demo; it is earned through traceable behavior, bounded scope, judgment, and outcome history. PASS is not granted because output looks like a refusal; it is earned by the evidence. Appearance is not authority, a signal is not a receipt, a receipt is not a verdict, and execution is not admission.
Standing
A peer paper in the flat register. It applies the evidence-authority argument to a public evaluation harness; it sits above nothing and nothing sits above it.
The decision contract draws on SPLAT constraint geometry from Computing Around the Open Center and on the admission/observation separation developed in Proof-Horizon Sharding. Those are first-party architecture sources, not independent validation of this finding.
Technical whitepaper 1.0, review edition. A pinned-source analysis plus an executed, isolated mechanism-level reproduction. Not a live foundation-model jailbreak, not a hosted-service penetration test, and not access to any lab's private evaluation infrastructure.
- Words
- 6,100
- Rendered pages
- 16
- Sections
- 12
- References
- 20
The pinned read
Every source claim is bound to one commit, read GitHub connector; official public webpages; ChatGPT Library. A later upstream change does not silently rewrite this finding; it produces a new epoch.
- Repository
- promptfoo/promptfoo
- Commit
- 32bfa9edae56d36f526ab13b6637db8dfb82c73e
- Commit timestamp (UTC)
- 2026-09-12T23:38:05Z
- Files pinned
- 8
- src/redteam/plugins/base.tsblob 3c9af33cb0e0…
- src/redteam/util.tsblob 8089bc7da48a…
- src/redteam/plugins/policy/index.tsblob 177e05a0b2a3…
- src/redteam/plugins/bola.tsblob 6a4665122594…
- src/assertions/redteam.tsblob 635113376e68…
- src/assertions/guardrails.tsblob 3faf841af094…
- src/matchers/rubric.tsblob 718b65fe6fb5…
- LICENSEblob af3fa111d930…
Executed mechanism harness
Isolated transcribed-source mechanism harness; not an upstream package, provider, model, CLI, or full plugin integration run. Run 2026-09-13T02:43:58.344Z on v22.16.0. Network: none. Model calls: 0. Upstream package executed: no.
- Assertions
- 45
- Base fixtures
- 13
- False passes
- 5
- Model calls
- 0

Honest negatives
Results, not disclaimers. A paper about evidence authority does not get to overclaim its own.
No live exploitation was observed
No private lab deployment, unauthorized data access, harmful completion, or in-the-wild exploitation was seen. The evidence is a pinned source read plus an isolated mechanism harness.
The custom-policy grader is protected by default
The policy grader forwards a skip flag that defaults to true, so its refusals still reach policy grading. The finding is not a blanket claim about every Promptfoo assertion.
No prevalence, severity, or affected-version claim
No CVE status, severity score, version interval, or frequency estimate is assigned. Source inspection establishes a branch; it does not establish how often deployments meet it.
The harness is not an upstream integration test
Transcribed mechanisms were executed in isolation with zero model calls. The full package, provider graph, CLI, cache, and aggregate reporting were not run.
The repair contract
- 01
Separate four objects: observation, evaluation coverage, policy assessment, and admission authority.
- 02
Preserve missingness — an absent signal is not a clean signal, and a rendered zero must carry its cause.
- 03
Bind every receipt to the subject it actually witnessed, at the evidence level it actually reached.
- 04
Repair the shortest causal path first: no deterministic branch may issue a verdict the judge never reached.
- 05
Carry correction across proof horizons, so a later reversal updates lineage instead of erasing it.
Section map
- 1The model did not fool the judge
- 2Scope, method, and the correction that matters
- 3The earliest causal failure
- 4Executed mechanism-level results
- 5Related paths: one family, different contracts
- 6What the literature already teaches—and what this case adds
- 7A formal boundary for warranted verdicts
- 8SPLAT: preserve the shape before reducing the result
- 9Receipts before verdicts: a repair contract
- 10Evaluation integrity must itself be evaluated
- 11The strongest objection
- 12Conclusion: the receipt must earn the green
Peer works it leans on
- proof horizon shardingapplies
Admission authority separated from observation, here applied to an evaluation harness.
- computing around the open centerapplies
SPLAT constraint geometry used as the diagnostic frame for premature collapse.
- forkedinforms
Observer-indexed standing: what a verdict is allowed to claim about whom.
- fractal quivers of quiverscompanion-to
Conceptual ancestry for governed narrowing without premature collapse.
Sealed artifacts
- Rendered PDFheld
sha256 8608257048945783… - Rendered DOCX
sha256 41c4907a1b776a9f… - Canonical manuscript (Markdown)
sha256 bb86984ba5757f51… - Reproduction harness
sha256 f965a737db6eede0… - Reproduction results
sha256 4d4d6607da37377d… - Source manifest
sha256 6f0e0bc0c7ef4780… - Execution log
- Packaging receipt
sha256 e66c7a9315f412d4… - Source checksums
- Upstream license
Rows marked held are cited but not published here. Write breyden@prompted.community to request access to a cited private artifact.
Recommended citation
Taylor, B. E. (2026). The Verdict Before the Evidence: Promptfoo, refusal short-circuits, and the governance of AI assurance (Technical Whitepaper 1.0, Review Edition). Prompted LLC.