{
  "lane": "A",
  "lane_title": "T7 INVENTORY + AFFORDABILITY — MEASURED NOT ASSUMED",
  "tic": 748,
  "measured_at_utc": "2026-08-28T19:20:40Z",
  "measurement_window_utc": "2026-08-28T19:20:40Z .. 2026-08-28T19:52:00Z",
  "authored_by": {
    "standing": "task_scoped_worker",
    "agent_type": "general-purpose",
    "dispatched_by": "ent_homeskillet",
    "model_id": "claude-opus-5[1m]",
    "inscription_authority": false,
    "outputs_owned_by": "ent_homeskillet"
  },
  "posture": "read-only survey. Zero writes to /Volumes/T7 Shield or ~/tmux-dumps. Zero doctrine/backlog/queue/git mutation. Write fence honored: this file + the boot-injection sink only.",
  "slice_declaration": "Every figure below names its instrument. `du` figures are `du -sh` in GiB-rounded human units on APFS over the T7 (USB) mount, run 2026-08-28T19:22-19:45Z. `ls -l` figures are exact bytes. `df` figures are the live mount table. sha256 via `shasum -a 256`. Version strings via each binary's own `--version`. Receipt figures are QUOTED from the named artifact, never re-derived. Where I did not run a thing, the row says so — see honest_limits.",

  "hardware": {
    "model": "Mac14,5",
    "cpu": "Apple M2 Max",
    "ncpu": 12,
    "unified_memory_bytes": 34359738368,
    "unified_memory_gib": 32.0,
    "unified_memory_source": "sysctl -n hw.memsize -> 34359738368",
    "os": "macOS 27.0 (sysctl kern.osproductversion)",
    "metal_gpu_working_set_mib_measured": 25559,
    "metal_gpu_working_set_gib_measured": 24.96,
    "metal_gpu_working_set_source": "llama.cpp device_info line in /Volumes/T7 Shield/models/model-mlx-maybe/tasks/bjz958hid.output: 'MTL0 : Apple M2 Max (25559 MiB, 25558 MiB free)' — the runtime's own report of the device budget at load, on this machine, with the 27B GGUF. This is the REAL GPU-resident ceiling, not a rule of thumb.",
    "iogpu_wired_limit_mb_NOW": 0,
    "iogpu_wired_limit_source": "sysctl iogpu.wired_limit_mb -> 0 (= default/unset)",
    "iogpu_wired_limit_note": "The tic-594 raise to 28672 (SOVEREIGN-FLEET-RUNBOOK.md §0, sudo, non-persistent across reboot) is NO LONGER IN EFFECT. Every co-residency claim resting on a 28 GiB cap ('both engines fit') is currently FALSE-BY-STATE. Single-model budget is the default ~24.96 GiB.",
    "disk": {
      "T7_Shield": {"size": "3.6Ti", "used": "3.2Ti", "avail_gib": 407, "capacity": "90%", "device": "/dev/disk6s2", "source": "df -h '/Volumes/T7 Shield'"},
      "internal_data_volume": {"size": "926Gi", "used": "872Gi", "avail_gib": 15, "capacity": "99%", "device": "/dev/disk3s5", "mount": "/System/Volumes/Data", "source": "df -h /System/Volumes/Data"},
      "BINDING_CONSTRAINT": "The internal disk has 15 GiB free, not the 926 GiB its size suggests. ANY conversion output larger than ~14 GiB has nowhere to land internally. This is the single most under-modeled affordability fact in this survey."
    },
    "runtimes": {
      "mlx_lm_version": "0.31.3",
      "mlx_lm_version_source": "python3 -c 'import mlx_lm; print(mlx_lm.__version__)' -> 0.31.3",
      "mlx_lm_version_correction": "The dispatch stated 0.31.2. MEASURED 0.31.3. Corrected, not assumed.",
      "mlx_lm_path": "/opt/homebrew/lib/python3.14/site-packages/mlx_lm (20 CLI shims at /opt/homebrew/bin/mlx_lm.*)",
      "mlx_lm_relevant_architectures_present": ["qwen3_5_moe.py", "qwen3_5.py", "qwen3_next.py", "qwen3_moe.py", "qwen3.py", "gated_delta.py", "gemma4_text.py", "qwen3_vl_moe.py"],
      "mlx_lm_arch_note": "qwen3_5_moe.py AND gated_delta.py are BOTH present in the installed mlx_lm. The AgentWorld-35B-A3B config.json declares model_type 'qwen3_5_moe' with layer_types alternating linear_attention/full_attention — the exact shape gated_delta.py implements. The architecture is NOMINALLY supported. NOT VERIFIED BY EXECUTION (see honest_limits H1).",
      "ollama": {
        "binary": "/usr/local/bin/ollama -> /Applications/Ollama.app/Contents/Resources/ollama",
        "server_running": true,
        "server_pid_observed": 90723,
        "ollama_list": [
          {"name": "qwen3-embedding:0.6b", "id": "ac6da0dfba84", "size": "639 MB", "modified": "6 months ago"},
          {"name": "all-minilm:latest", "id": "1b226e2802db", "size": "45 MB", "modified": "6 months ago"},
          {"name": "qwen3:8b", "id": "500a1f067a9f", "size": "5.2 GB", "modified": "7 months ago"}
        ],
        "manifests_on_disk": ["biomed/latest", "gpt-oss/20b-cloud", "llama2/latest", "llama3.2/latest", "qwen3/30b-a3b", "qwen3-embedding/8b", "qwen3-embedding/latest"],
        "blobs": {"count": 24, "du": "15G", "path": "~/.ollama/models/blobs"},
        "DISAGREEMENT": "ollama list reports 3 models. ~/.ollama/models/manifests holds 7 manifests, NONE of which name any of the 3 listed. The 3 listed model IDs (ac6da0dfba84 / 1b226e2802db / 500a1f067a9f) match ZERO blob filenames under ~/.ollama/models/blobs. Two readers of the same substrate disagree — see stub_census row S09."
      },
      "llama_server_binaries": [
        {"path": "/Users/breydentaylor/.local/share/codex-runners/llama.cpp-qwen36-mtp/build/bin/llama-server", "version": "1 (dbe9c0c)", "built": "AppleClang 21.0.0 Darwin arm64", "probe": "--version RAN, exit 0", "role": "the 27B reasoner runner named in SOVEREIGN-FLEET-RUNBOOK.md §2 and the tic-544 first-fire receipt"},
        {"path": "/Users/breydentaylor/.unsloth/llama.cpp/build/bin/llama-server", "version": "9632 (5747caa8c)", "built": "AppleClang 21.0.0 Darwin arm64", "probe": "--version RAN, exit 0", "role": "the runner NAMED IN THE QWYTHOS G3 RECEIPTS (2026-07-14, 120 pre/post cases)"},
        {"path": "/Users/breydentaylor/.docker/bin/inference/llama-server", "version": "1 (72874f5)", "built": "AppleClang 17.0.0 Darwin arm64", "probe": "--version RAN, exit 0", "role": "Docker Desktop model runner; incidental"}
      ],
      "llama_server_path_correction": "The dispatch said 'no llama-server on PATH'. TRUE as stated — `which llama-server` fails. But THREE working llama-server binaries exist on disk and all three answer --version. 'Not on PATH' is not 'not present'. The 27B and Qwythos serve paths are intact."
    }
  },

  "inventory": [
    {
      "entry": "organization-engine-lora",
      "du_measured": "20G",
      "index_size": "20.0 GiB",
      "index_agrees": true,
      "what_it_is": "The org-engine LoRA epoch archive. 252 files. Composition MEASURED by du: 14 epoch .zip adapters at ~456M each (epoch05,06x2,08,09,10,11,12,13,14,15,16 + the epoch05 eval 128K) = ~6.0G; the 2.7G compact-continuation export (org_engine_epoch_lora_Qwen3-14B-unsloth-bnb-4bit_compact_continue_20260616T220931Z.zip, sha 15fe6578...); upload_chunks_20260617 2.7G; self_foreclosure_upload_staging 457M; training/ 9.0G (additive_v1 8.9G incl. remote_training_v1/runs; additive_v2 23M — DATA ONLY, no weights); receipts/ 384K (one file: epoch15_f2_governed_tic514_receipt.json); README.md 3515 B.",
      "canonical_head": "org_engine_epoch16_s6_f2_rebalance_from_epoch14_20260626T191807Z_sdpa.zip (478,023,981 B), sha256 b7db58dac68326901679daa14624d0e607c91438dd268476f9549fbc3e08b425 per README",
      "has_run_here": true,
      "receipt": "audit-logs/governance/harpoon-office/cable-receipts/S3_model_proposer-ent_egress_router-tic535.json (+ tic538, tic590, tic600)",
      "blocker": "Its serve path's BASE MODEL is gone from this machine — see stub_census S02."
    },
    {
      "entry": "mlx-adapters",
      "du_measured": "993M",
      "index_size": "991.0 MiB",
      "index_agrees": true,
      "what_it_is": "Two dirs, 7 real files. org-engine-epoch16-mlx/ = adapters.safetensors 513,864,367 B (sha256 ff96cbb6212c0ee7dc9a6cf56499dd940710833cd288beed6ac48a2991be4cbb) + adapter_config.json 322 B (MLX form: fine_tune_type lora, num_layers 40, rank 32, scale 1.0, 7 target keys). org-engine-epoch16-peft-src/ = adapter_model.safetensors 513,877,864 B (sha256 5e4ede9ce9d6bffd3e7a267b5464cc8e5c121bcc9fcdf6aee6e686bc9e7f83e8) + PEFT adapter_config.json (base unsloth/Qwen3-14B-unsloth-bnb-4bit, r=32, lora_alpha=32, peft 0.19.1) + tokenizer.json 11.4M + tokenizer_config + chat_template.jinja. The MLX form was produced tic 594 by audit-logs/f2/peft_to_mlx_lora.py.",
      "has_run_here": true,
      "receipt": "cable-receipts/S3_model_proposer-ent_egress_router-tic600.json names this exact path as 'durable_mlx_adapter_present ... already converted at tic594, no re-conversion needed'"
    },
    {
      "entry": "model-mlx-maybe",
      "du_measured": "505M",
      "index_size": "501.1 MiB",
      "index_agrees": true,
      "what_it_is": "NOT a model — a scratch dir with THREE things. (1) scratchpad/org-engine-epoch16/ = a BYTE-IDENTICAL duplicate of mlx-adapters/org-engine-epoch16-peft-src (sha256 5e4ede9ce9d6... on both adapter_model.safetensors; adapter_config.json `diff` IDENTICAL). (2) scratchpad/qwen-feedback-eval/ = the tic-595 governed feedback-eval output (approval report + summary.json + outputs.jsonl). (3) tasks/ = three .output files — one EMPTY (biyk7lfsq), one a llama.cpp llama-server LOAD+SERVE LOG for the 27B GGUF (bjz958hid, 6511 B — the source of the 25559 MiB Metal figure), one a tic-594 workflow result JSON (wa42rtsss, 8172 B).",
      "has_run_here": true,
      "receipt": "model-mlx-maybe/scratchpad/qwen-feedback-eval/qwen_feedback_summary.json — 24 cases, 0 model_call_failures, endpoint http://127.0.0.1:8088/v1/chat/completions, model 'qwen3.6-27b-mtp-q4km', generated 2026-07-09T17:19:44Z"
    },
    {
      "entry": "qwen-agentworld-35b-a3b",
      "du_measured": "65G",
      "index_size": "64.6 GiB",
      "index_agrees": true,
      "what_it_is": "official-bf16-60d2b043/ = 65G, 21 BF16 safetensor shards, indexed tensor payload 69,321,221,376 B, revision 60d2b0434a53d2e62a7c00a489586815d94ebffb. Qwen-AgentWorld-35B-A3B: Qwen3.5 MoE, 35B total / 3B active, 40 layers, hidden 2048, head_dim 256, 16 Q heads / 2 KV heads, 256 experts (8 routed + 1 shared), moe_intermediate 512, vocab 248,320, context 262,144, full_attention_interval 4 (10 full-attention layers, 30 linear/GatedDeltaNet), license apache-2.0, base Qwen3.5-35B-A3B-Base, dataset AgentWorldBench. config.json architectures=['Qwen3_5MoeForConditionalGeneration'] with language_model_only:true and a vision_config that has NO weights in this checkpoint. evals/ = 294M: g3-model-population-v1 78M (org-engine-epoch16-v1 + qwythos-base-v2-mtp-q6 + -v2), g4-shadow-v1 214M (11 run dirs 2026-07-14), hybrid-g3-joined-evidence-v1 1.3M.",
      "has_run_here": false,
      "has_run_anywhere": "YES — but on a Packet B200, not here. MODEL_CARD_LOCAL.md G4 Runtime Receipt: vLLM --language-model-only at 131,072 ctx on one Packet B200, 60/60 predictions persisted, $1.3125 spend, quality admission 'shadow_held', G5 'not_granted'.",
      "receipt": "/Volumes/T7 Shield/models/qwen-agentworld-35b-a3b/official-bf16-60d2b043/MODEL_CARD_LOCAL.md + download_receipt.json; authoritative eval at canonical-mount/evals/agentworld_shadow/g4/runs/20260714T161913Z/evaluation_report.json"
    },
    {
      "entry": "unsloth-Qwen3.6-27B-MTP-GGUF",
      "du_measured": "16G",
      "index_size": "15.9 GiB",
      "index_agrees": true,
      "what_it_is": "Qwen3.6-27B-Q4_K_M.gguf = 17,106,773,120 B exact (15.93 GiB) + .sha256 sidecar + local_setup_manifest.json 3375 B + README.md 25415 B (unsloth card; pipeline_tag image-text-to-text, base Qwen/Qwen3.6-27B, tags qwen3_5, MTP guide). Quant level Q4_K_M with MTP (multi-token-prediction) draft head.",
      "has_run_here": true,
      "receipt": "audit-logs/governance/receipts/2026-07-02-tic544-27b-first-fire.md (quoted in models.yaml: 162 tokens / 7.31 tok/s gen / MTP 43.6% draft-acceptance / 127.0.0.1 only, zero egress) AND the raw llama-server log at model-mlx-maybe/tasks/bjz958hid.output AND the tic-595 24-case feedback eval via :8088"
    },
    {
      "entry": "qythos-9b-with-image-understanding",
      "du_measured": "24G",
      "index_size": "24.0 GiB",
      "index_agrees": true,
      "what_it_is": "FIVE GGUFs + an adapter set. Qwythos-9B-v2-MTP-Q6_K.gguf 7,666,069,184 B; Qwythos-9B-v2-Q6_K.gguf 7,458,300,672 B; Qwythos-9B-Claude-Mythos-5-1M-MTP-Q6_K.gguf 7,617,818,464 B; mmproj-Qwythos-9B-v2-BF16.gguf 921,704,512 B; mmproj-Qwythos-9B-Claude-Mythos-5-1M-F16.gguf 918,165,472 B. adapters/qwythos-v2-hybrid-world-tool-vlm-20260713/ = 1.1G, 7 .zip checkpoints each with a .receipt.json sidecar (text_epoch01, text_epoch04_corrective, world_splat_v2_bounded4, v3_receipted3, v5_corrective1, additive_v1_checkpoint1, additive_v2_checkpoint2). Current research checkpoint per README: world_splat_additive_v2_checkpoint2.zip, 71,571,912 B, sha 777c5e12..., base empero-ai/Qwythos-9B-v2 rev 2178f73a..., 144 records / 18 steps / 8192 ctx, train loss 0.4330 held-out 0.3952, frozen evals 3/6 + 2/5 + 1/6, 'runtime admission held'.",
      "has_run_here": true,
      "receipt": "/Volumes/T7 Shield/models/qwen-agentworld-35b-a3b/evals/g3-model-population-v1/qwythos-base-v2-mtp-q6-v2/receipts/*.{pre,post}.json — 120 receipts, 2026-07-14, artifact_path Qwythos-9B-v2-MTP-Q6_K.gguf, artifact_sha256 24271fb6e16b2b8f581d192cf77d947453d7852f965ce6801fefa6964bbea060, runtime 'llama.cpp:/Users/breydentaylor/.unsloth/llama.cpp/build/bin/llama-server', latency 23,302-33,748 ms/case, ~2254 prompt + 549 completion tokens"
    },
    {
      "entry": "trellis2-4b",
      "du_measured": "ABSENT — directory does not exist",
      "index_size": "15.1 GiB / 46 files / 9 weights / safetensors (MODEL_INDEX.md)",
      "index_agrees": false,
      "what_it_is": "GONE. `ls -la /Volumes/T7 Shield/models | grep -i trellis` returns nothing. The ONLY trellis artifact on the T7 is /Volumes/T7 Shield/models/.hf-hub-cache/models--microsoft--TRELLIS.2-4B — and the WHOLE .hf-hub-cache is 896K, i.e. refs/pointers with no blobs. 15.1 GiB of indexed weights are not on disk.",
      "has_run_here": false,
      "receipt": null
    },
    {
      "entry": "gemma4-coding-Q8_0.gguf",
      "du_measured": "12,669,645,344 B exact (11.80 GiB) — ls -l",
      "index_size": "11.8 GiB",
      "index_agrees": true,
      "what_it_is": "A single Q8_0 GGUF, dated Jun 15. No sidecar, no README, no manifest, no Modelfile, no models.yaml entry, no launcher, no receipt.",
      "has_run_here": false,
      "receipt": null
    },
    {
      "entry": "vggt-capture3d",
      "du_measured": "874M",
      "index_size": "838.1 MiB",
      "index_agrees": true,
      "what_it_is": "Not a served LM. VGGT (Visual Geometry Grounded Transformer, Oxford VGG + Meta AI, arXiv 2503.11651) — a feed-forward 3D-reconstruction model. Layout: repos/vggt (the source repo), checkpoints/, hf-cache/, apfs-work/, vggt-capture3d-env.sparsebundle (a sparse disk image env), checkpoint-access-report.json, install-report.json, setup-state.json. Catalogued in models.yaml as vggt_capture3d, served_via local_capture_studio.",
      "has_run_here": "unknown",
      "receipt": "install-report.json / setup-state.json / checkpoint-access-report.json exist but were not opened this pass (out of the affordability scope — vggt is not an LM candidate for this hoist)"
    },
    {
      "entry": "datasets-training-core-harpoon-assess",
      "du_measured": "6.1G",
      "index_size": "6.1 GiB / 2 files",
      "index_agrees": true,
      "what_it_is": "Two zip archives, contents read with `unzip -l`. (1) 'DepMap0Q4CRISPRGeneDependency(Achilles).zip' 473,986,903 B -> 9 files / 1,918,019,118 B uncompressed: DepMap 20Q4 Public CRISPR gene-dependency + gene-effect + CCLE expression/CN CSVs. A biology dataset. (2) 'not-so-dumb-at-20q.zip' 6,069,444,197 B -> 12 files / 7,648,343,326 B uncompressed and IT CONTAINS A MODEL: submission/phi3_model/ with model-00001-of-00002.safetensors (4,972,489,328 B) + model-00002-of-00002.safetensors (2,669,692,552 B) + config/generation_config/index, plus a phi3_tokenizer and word_universe.csv. A Kaggle-style 20-questions submission bundle, dated 2024-07-01.",
      "has_run_here": false,
      "receipt": null
    },
    {
      "entry": "claude-code-artifacts-ubiquity-canonical-federation-pieces",
      "du_measured": "1.7G",
      "index_size": "864.2 MiB / 2698 files / 287 docs",
      "index_agrees": false,
      "index_delta": "MEASURED 1.7G vs INDEXED 864.2 MiB — roughly DOUBLED since 2026-07-17.",
      "what_it_is": "A snapshot mirror of the canonical memory-root (memory/MEMORY.md, MEMORY-INDEX.md, ~287 feedback_*/project_*/reference_* topic files) plus artifacts. Top-level mtime Jul 14 08:07. NOT a model; it is a stale copy of a live surface.",
      "has_run_here": "n/a",
      "receipt": null
    },
    {
      "entry": "ollama-library",
      "du_measured": "128K (an empty directory — `find` returns only the dir itself)",
      "index_size": "0 B / 0 files",
      "index_agrees": true,
      "what_it_is": "An empty directory named as if it mirrored an ollama model library. It holds nothing.",
      "has_run_here": false,
      "receipt": null
    },
    {
      "entry": "ollama-modelfiles",
      "du_measured": "384K",
      "index_size": "460 B / 1 file",
      "index_agrees": true,
      "what_it_is": "ONE file: Modelfile.qwen3_6_27b_mtp_q4km. Content read in full: FROM the T7 Qwen3.6-27B-Q4_K_M.gguf; PARAMETER num_ctx 4096 / temperature 0.2 / top_p 0.9; SYSTEM prompt declaring it a 'local Qwen3.6 27B diagnostic base model for file-organization-engine experiments' with no-delete reasoning and an explicit 'Do not claim to be the fine-tuned organization-engine adapter unless an adapter has been merged or loaded.'",
      "has_run_here": false,
      "receipt": null,
      "note": "No corresponding entry in `ollama list`. The Modelfile was authored and never `ollama create`d. See stub_census S08."
    },
    {
      "entry": "tmux-dumps",
      "du_measured": "35G",
      "index_size": "32.3 GiB / 13814 files / 70 docs",
      "index_agrees": "grew ~+2.7 GiB since 2026-07-17",
      "what_it_is": "LANE D's object — recorded here for the affordability picture only. Also the home of TWO load-bearing runtime assets this lane depends on: run_qwen36_mtp_q4km.sh (700 B, the 27B launcher, VERIFIED PRESENT, content read) and ubiquity_qwen_feedback_assessor.py (41,954 B, the governed feedback gate, VERIFIED PRESENT).",
      "has_run_here": true,
      "receipt": "the launcher's own content + the tic-595 eval outputs"
    },
    {
      "entry": "MODEL_INDEX.md / MODEL_INDEX.json",
      "du_measured": "11,393 B / 50,038 B",
      "index_size": "n/a",
      "what_it_is": "The card catalog. Generated 2026-07-17T02:55:41Z. Read in full. Its header 'Disk: 3.6 TiB used / 3.6 TiB total; 68.5 GiB free' is now WRONG — df reports 407 GiB free. Its trellis2-4b row is now WRONG — the directory is gone. Its claude-code-artifacts row is now WRONG — 864.2 MiB has become 1.7G. Its Duplicate Weight Candidates section flagged the two adapter_model.safetensors as 'unverified same name and size'; THIS PASS VERIFIED THEM BYTE-IDENTICAL by sha256.",
      "has_run_here": "n/a",
      "receipt": null
    },
    {
      "entry": "io-map CSVs + io_map_loop_assessment.md + real_learning_loops_bounty_submission.md",
      "du_measured": "1.1K / 2.2K / 6.7K / 1.2K CSVs; 18,282 B md; 14,639 B md",
      "index_size": "matches",
      "what_it_is": "READ IN FULL (assessment) and in substantial part (bounty). io_map_loop_assessment.md: parses io-map.json (16,548 raw edges, 85 medium/strong material edges); finds 3 actual SCCs (queue-mutation clique over queue.jsonl; a CGG-tests loop; runtime-sync<->inbox-envelope over */signals) and 11 near/underleveraged loops RANKED, with priority 1 = the signal-metabolism loop, 2 = conformations->RTCH->rtch/packets ('the packet has no detected eater... likely the highest-value close'), 3 = io-map/router navigation products written-never-read ('the visual proof of the c48 crack'). Lock line: 'The strongest pattern is not absence of routes; it is half-closed metabolism.' real_learning_loops_bounty_submission.md: proposes a six-gate definition of a real learning loop (sensor -> eater -> receipt -> correction -> next-tic delta -> non-crowning) and strikes twelve named loops with a minimal loop-receipt schema. NEITHER file is about models. They are about WHERE LEVERAGE WAITS in the governance metabolism — see the note to Lane B/C in honest_limits H7.",
      "has_run_here": "n/a",
      "receipt": null
    },
    {
      "entry": ".cache / .hf-home / .hf-hub-cache / .hf-xet-cache (T7 dotdirs)",
      "du_measured": "4.8M / 512K / 896K / 1.0M",
      "index_size": "not indexed",
      "what_it_is": "HF cache scaffolding with essentially NO blobs. .hf-hub-cache holds exactly one entry, models--microsoft--TRELLIS.2-4B, at pointer weight only. These dirs are the remains of an HF_HOME-on-T7 arrangement that no longer holds weights.",
      "has_run_here": "n/a",
      "receipt": null
    }
  ],

  "affordability": {
    "formula_declared": {
      "resident_bytes": "weights_bytes + kv_cache_bytes + runtime_overhead",
      "weights_bytes": "params_total * bits_effective / 8",
      "bits_effective": {
        "MLX 4-bit (group 64)": "4.5  (4 bits/weight + one fp16 scale + one fp16 bias per 64 weights => 4 + 32/64)",
        "GGUF Q4_K_M": "~4.85",
        "GGUF Q6_K": "~6.56",
        "GGUF Q8_0": "~8.5",
        "bf16": "16"
      },
      "kv_cache_bytes_per_token": "2 (K and V) * n_kv_heads * head_dim * n_FULL_ATTENTION_layers * bytes_per_element",
      "kv_note": "For hybrid linear/full-attention models only the FULL-attention layers carry a per-token KV cache; the linear-attention (GatedDeltaNet) layers carry a CONSTANT-size recurrent state instead.",
      "headroom_rule": "The binding ceiling on this machine is the Metal GPU working set MEASURED at 25,559 MiB (24.96 GiB) with iogpu.wired_limit_mb at its default 0, NOT the 32 GiB of unified memory and NOT the dispatch's '~8 GiB headroom' heuristic. The heuristic and the measurement happen to land close (32 - 8 = 24 vs 24.96 measured); I use the measured one."
    },
    "ranking": [
      {
        "rank": 1,
        "candidate": "Qwen3.6-27B-MTP Q4_K_M GGUF (llama.cpp Metal, :8088)",
        "verdict": "FITS-NOW",
        "params_total": "27B",
        "params_active": "27B (dense)",
        "quantization_on_disk": "Q4_K_M + MTP draft head",
        "disk_on_hand_gib": 15.93,
        "disk_needed_gib": 0,
        "resident_gib_estimate": 17.0,
        "resident_breakdown": "16.0 GiB weights (27B * 4.85/8) + 0.51 GiB MTP draft context (llama.cpp's OWN estimate in the log: 'estimated memory usage of MTP context is 521.00 MiB') + ~0.3 GiB KV @ 4096 ctx + ~0.2 GiB buffers",
        "fits_32gib_m2max": true,
        "fits_under_measured_metal_budget_24_96gib": true,
        "co_resident_with_a_second_engine": "NOT at the current default cap. 17.0 + 11.1 (org-engine server) = 28.1 GiB > 24.96 GiB. The tic-594 runbook's 'both fit' finding required sudo sysctl iogpu.wired_limit_mb=28672, which is NOT currently set.",
        "measured_receipts_on_this_machine": [
          "tic 544 first-fire: 162 tokens (60 prompt + 102 gen), 7.31 tok/s gen, MTP 43.6% draft-acceptance, load ~71 s, 127.0.0.1 only, clean shutdown verified (audit-logs/governance/receipts/2026-07-02-tic544-27b-first-fire.md, quoted via models.yaml:108)",
          "raw llama-server load+serve log with Metal device_info and MTP speculative init (model-mlx-maybe/tasks/bjz958hid.output)",
          "tic 595 governed feedback eval: 24 cases, 0 model_call_failures, json_parse 1.000, authority_lane_safety 1.000, route_match 0.417, outcome_match 0.625 (model-mlx-maybe/scratchpad/qwen-feedback-eval/qwen_feedback_summary.json)",
          "tic 600: reasoner was ALREADY UP at :8088 (pid 50272) and stayed {\"status\":\"ok\"} throughout an independent MLX bring-up on :8090"
        ],
        "everything_needed_present": "gguf (T7) YES; runner (~/.local/share/codex-runners/.../llama-server dbe9c0c) YES, --version verified; launcher (T7 tmux-dumps/run_qwen36_mtp_q4km.sh) YES, content read; eval harness (T7 tmux-dumps/ubiquity_qwen_feedback_assessor.py + canonical audit-logs/f2/run-feedback-assessor.sh) YES; models.yaml entry YES and CORRECT (served_via: llama_cpp)",
        "cheapest_real_hoist": "Nothing to acquire. `bash '/Volumes/T7 Shield/models/tmux-dumps/run_qwen36_mtp_q4km.sh'` and it is up in ~71 s."
      },
      {
        "rank": 2,
        "candidate": "Qwythos-9B-v2-MTP-Q6_K + mmproj-BF16 (llama.cpp Metal, VLM-capable)",
        "verdict": "FITS-NOW",
        "params_total": "9B",
        "params_active": "9B (dense) + a frozen vision tower via mmproj",
        "quantization_on_disk": "Q6_K text + BF16 projector",
        "disk_on_hand_gib": 8.0,
        "disk_needed_gib": 0,
        "resident_gib_estimate": 8.9,
        "resident_breakdown": "7.14 GiB text weights + 0.86 GiB mmproj + ~0.6 GiB KV/buffers at moderate context; MTP draft adds a small fixed context",
        "fits_32gib_m2max": true,
        "fits_under_measured_metal_budget_24_96gib": true,
        "co_resident_with_a_second_engine": "YES — 8.9 + 11.1 (org-engine) = 20.0 GiB, or 8.9 + 17.0 (27B) = 25.9 GiB (the latter is 0.9 GiB OVER the current default cap; would need the sudo raise).",
        "measured_receipts_on_this_machine": [
          "120 G3 pre/post receipts dated 2026-07-14 naming artifact_path=Qwythos-9B-v2-MTP-Q6_K.gguf, artifact_sha256=24271fb6e16b2b8f581d192cf77d947453d7852f965ce6801fefa6964bbea060, runtime='llama.cpp:/Users/breydentaylor/.unsloth/llama.cpp/build/bin/llama-server', per-case latency 23,302-33,748 ms, usage ~2254 prompt + 549 completion tokens => roughly 16 tok/s output",
          "every receipt carries admissibility {source_class: model_simulation, may_terminalize_governance: false, may_mutate_canonical_state: false}"
        ],
        "everything_needed_present": "ggufs (T7) YES; runner (~/.unsloth/llama.cpp/build/bin/llama-server, version 9632/5747caa8c) YES, --version verified; models.yaml entry YES (qwythos_v2_hybrid, served_via llama_cpp_plus_mmproj) — but note the entry's `blob:` points at Qwythos-9B-v2-Q6_K.gguf while the RECEIPTS used Qwythos-9B-v2-MTP-Q6_K.gguf (different file, different sha)",
        "cheapest_real_hoist": "Nothing to acquire. This is the only local model on the machine with IMAGE capability and a 120-case receipt trail."
      },
      {
        "rank": 3,
        "candidate": "org-engine epoch16 = mlx-community/Qwen3-14B-4bit + epoch16 LoRA (r=32) on MLX",
        "verdict": "FITS-WITH-REACQUISITION  (RAM fits comfortably; the BASE WEIGHTS ARE GONE from this machine)",
        "params_total": "14.8B",
        "params_active": "14.8B (dense)",
        "quantization_on_disk": "the ADAPTER is on disk (490 MiB, both MLX and PEFT forms). The 4-bit BASE is NOT.",
        "disk_on_hand_gib": 0.49,
        "disk_needed_gib": 8.1,
        "disk_needed_note": "~8.1 GiB for mlx-community/Qwen3-14B-4bit (the runbook's own figure; the tic-600 receipt says 7.8G as measured then). Internal free is 15 GiB — it FITS, with 6.9 GiB left over. Landing it on the T7 instead would be a MEMBRANE WRITE and is not mine to propose as a default.",
        "resident_gib_estimate": 11.1,
        "resident_breakdown": "MEASURED, not estimated: 8.31 GB peak base-only (tic 535); 10.697 GB with the adapter (tic 538, '+2.4GB delta'); 10.79 GB (tic 590); 11.9 GB physical / 10.3 GB GPU-resident in server mode co-resident with the warm 27B (tic 600, method: vmmap --summary on the live pid after the call completed)",
        "fits_32gib_m2max": true,
        "fits_under_measured_metal_budget_24_96gib": true,
        "measured_receipts_on_this_machine": [
          "tic 535 — FIRST LIVE SERVE: PEFT->MLX-LoRA conversion (280 modules, 40L x 7 targets) -> base mlx-community/Qwen3-14B-4bit + adapter via mlx-lm -> real inference 88.9 s / 8.31 GB peak, valid JSON proposal, gate_passed:true, 0 errors, mutation_authority:false (cable-receipts/S3_model_proposer-ent_egress_router-tic535.json)",
          "tic 538/539 — one live SP2 cycle: ADMIT_PROPOSAL_REFUSE_MOVE, 10.697 GB with adapter (cable-receipts/S3_model_proposer-live-tic538.json)",
          "tic 590 — SERVED, 10.79 GB peak, reproduced the tic-535 proposal shape",
          "tic 594 — durable MLX adapter conversion landed on T7; sudo iogpu.wired_limit_mb=28672 raised; both engines proven co-resident",
          "tic 600 — SERVED in server mode: bring-up python3.14 -m mlx_lm server --model mlx-community/Qwen3-14B-4bit --adapter-path '<T7 mlx adapter>' --port 8090; health 200; 4920 prompt + 678 completion = 5598 tokens in 84.82 s wall (~8 tok/s output); 11.9 GB physical peak / 10.3 GB GPU-resident; gate_passed=true; server cleanly stopped and port verified free"
        ],
        "BLOCKER": "~/.cache/huggingface DOES NOT EXIST. The 7.8-8.1 GiB mlx-community/Qwen3-14B-4bit base that all five receipts above depend on is absent from this machine. Searched: ~/.cache (holds only claude/codex-runtimes/gh/uv), ~/Library/Caches, ~/.hf, ~/huggingface, and every HF cache dir on the T7 (.hf-home 512K, .hf-hub-cache 896K holding only a TRELLIS pointer, .cache 4.8M). `find ~ -maxdepth 5 -name 'models--mlx-community*'` returns NOTHING. See stub_census S02.",
        "cheapest_real_hoist": "One `huggingface-cli download mlx-community/Qwen3-14B-4bit` (~8.1 GiB egress, ~10-25 min on a normal link) and the entire tic-535..600 serve path relights unchanged. The adapter, the bridge, the runbook, the registry entry, and the harness are all intact."
      },
      {
        "rank": 4,
        "candidate": "Qwen-AgentWorld-35B-A3B -> MLX 4-bit conversion",
        "verdict": "FITS-WITH-CONVERSION on RAM  /  BLOCKED-ON-DISK for the output",
        "params_total": "35B",
        "params_active": "~3B (8 routed + 1 shared expert of 256, expert_intermediate 512)",
        "quantization_on_disk": "bf16 only. NO 4-bit or GGUF quant of this model exists anywhere on this machine (searched /Users/breydentaylor and /Volumes/T7 Shield at depth 5 for *35b* and *a3b* — every hit was an unrelated UUID or a yarn cache entry).",
        "disk_on_hand_gib": 64.56,
        "disk_needed_gib": 18.5,
        "disk_blocker": "The ~18.5 GiB conversion output DOES NOT FIT on the internal disk (15 GiB free). The only surface with room is the T7 (407 GiB free) — which is a MEMBRANE. Writing a converted model there is an Architect call, not mine. This is the row's real gate, and it is a DISK gate, not a RAM gate.",
        "resident_gib_estimate_at_32k_ctx": 20.0,
        "resident_breakdown": "weights 35e9 * 4.5/8 = 19.69 GB = 18.33 GiB. KV: only 10 of 40 layers are full_attention (full_attention_interval 4) -> 2(K,V) * 2 kv_heads * 256 head_dim * 2 B * 10 layers = 20,480 B/token = 20 KiB/token -> 0.63 GiB @ 32,768 ctx (2.50 GiB @ 131,072 ctx). GatedDeltaNet recurrent state on the other 30 layers: 30 * 32 v_heads * 128 v_dim * 128 k_dim * 4 B (mamba_ssm_dtype float32) = 62.9 MB, CONSTANT not per-token. Runtime/graph overhead ~1.0 GiB.",
        "resident_gib_estimate_at_131k_ctx": 21.9,
        "fits_32gib_m2max": true,
        "fits_under_measured_metal_budget_24_96gib": "YES at 32k (20.0 < 24.96) and YES at 131k (21.9 < 24.96) — but ALONE. Zero room for a co-resident second engine at the current default cap.",
        "architecture_support": "mlx_lm 0.31.3 ships models/qwen3_5_moe.py AND models/gated_delta.py; the checkpoint's model_type is exactly 'qwen3_5_moe'. Nominally supported. UNVERIFIED — I did not run mlx_lm.convert (read-only lane). Two specific risks I can name but not resolve: (a) the top-level architecture is 'Qwen3_5MoeForConditionalGeneration' with a config.json that nests everything under text_config/vision_config — mlx_lm may need the text_config flattened or may key off model_type and handle it; (b) the checkpoint declares mtp_num_hidden_layers:1 and language_model_only:true, so the MTP head and the (weightless) vision tower must be tolerated or stripped.",
        "conversion_wall_time": "ESTIMATE ONLY, not measured: reading 64.56 GiB from the T7 at a realistic sustained 700-900 MB/s is ~75-95 s of pure I/O; the group-quantize pass over 693 tensors on an M2 Max plus writing ~18.5 GiB back is the dominant cost. I would budget 20-45 minutes and I would NOT report a number tighter than that without running it.",
        "measured_receipts_on_this_machine": "NONE. This model has never been served on this machine. Its only runtime receipt is the G4 shadow on a Packet B200 (60/60 predictions, $1.3125, quality admission 'shadow_held', G5 'not_granted').",
        "cheapest_real_hoist": "Two decisions, not one command: (1) where the ~18.5 GiB output lands (internal disk cannot hold it; the T7 is a membrane) and (2) whether a NEW model-source coupling is being opened at all — this is the R3 bell and it is the Architect's, never mine."
      },
      {
        "rank": 5,
        "candidate": "Qwythos-9B-Claude-Mythos-5-1M-MTP-Q6_K + mmproj-F16",
        "verdict": "FITS-NOW (RAM) / NO RECEIPT",
        "params_total": "9B",
        "disk_on_hand_gib": 7.95,
        "disk_needed_gib": 0,
        "resident_gib_estimate": 8.9,
        "fits_32gib_m2max": true,
        "measured_receipts_on_this_machine": "NONE FOUND. The G3 receipts all name the -v2 sibling, not this one. Sits on disk, catalogued nowhere in models.yaml.",
        "cheapest_real_hoist": "Nothing to acquire; but nothing justifies lighting it either — it is a sibling of a model that already has the receipts."
      },
      {
        "rank": 6,
        "candidate": "gemma4-coding-Q8_0.gguf",
        "verdict": "FITS-NOW (RAM) / ZERO PROVENANCE",
        "params_total": "~11.9B implied (12,669,645,344 B at Q8_0 ~8.5 bits/weight)",
        "disk_on_hand_gib": 11.80,
        "disk_needed_gib": 0,
        "resident_gib_estimate": 12.3,
        "fits_32gib_m2max": true,
        "fits_under_measured_metal_budget_24_96gib": true,
        "measured_receipts_on_this_machine": "NONE. No README, no sidecar, no Modelfile, no launcher, no models.yaml entry, no audit-log receipt. 11.8 GiB of Q8_0 weights with no story attached.",
        "cheapest_real_hoist": "A llama-server invocation would light it in minutes — but it has no declared role, and lighting a model to find out what it is for is the wrong order."
      },
      {
        "rank": 7,
        "candidate": "ollama qwen3:8b (already registered and served locally)",
        "verdict": "FITS-NOW",
        "disk_on_hand_gib": "5.2 GB per `ollama list`",
        "resident_gib_estimate": "~5.5",
        "fits_32gib_m2max": true,
        "measured_receipts_on_this_machine": "`ollama list` reports it registered and 7 months old; the ollama server is running right now (pid 90723). No governance receipt names it.",
        "cheapest_real_hoist": "`ollama run qwen3:8b`. The cheapest thing on the machine — and the least interesting: an off-the-shelf 8B with no federation lineage."
      },
      {
        "rank": 8,
        "candidate": "Phi-3 (~3.8B) inside datasets-training-core-harpoon-assess/not-so-dumb-at-20q.zip",
        "verdict": "DOES-NOT-FIT-THE-DISK (extraction blocked), RAM would be trivial",
        "disk_needed_gib": 7.12,
        "disk_blocker": "7.65 GB of safetensors would need extracting; internal free is 15 GiB so it fits numerically, but there is no reason to spend half the remaining internal headroom on a 2024 Kaggle 20-questions submission.",
        "measured_receipts_on_this_machine": "NONE. Filed as a FOUND MODEL, not a candidate — noted because a model is shelved inside a folder named 'datasets'."
      },
      {
        "rank": "n/a",
        "candidate": "trellis2-4b",
        "verdict": "NOT ON DISK — cannot be ranked",
        "note": "MODEL_INDEX.md lists 15.1 GiB / 9 weights. The directory does not exist. See stub_census S04."
      }
    ],
    "ranking_summary": {
      "FITS_NOW": ["Qwen3.6-27B-MTP Q4_K_M (17.0 GiB resident)", "Qwythos-9B-v2-MTP-Q6_K + mmproj (8.9 GiB)", "Qwythos-9B-Claude-Mythos-5-1M-MTP-Q6_K (8.9 GiB, no receipt)", "gemma4-coding-Q8_0 (12.3 GiB, no provenance)", "ollama qwen3:8b (~5.5 GiB)"],
      "FITS_WITH_REACQUISITION": ["org-engine epoch16 on MLX (11.1 GiB resident; needs an 8.1 GiB base re-download)"],
      "FITS_WITH_CONVERSION": ["Qwen-AgentWorld-35B-A3B -> MLX 4-bit (20.0 GiB resident @32k; ~18.5 GiB conversion output with NOWHERE INTERNAL TO LAND IT)"],
      "DOES_NOT_FIT": ["Qwen-AgentWorld-35B-A3B at bf16 (64.56 GiB — 2.6x the whole Metal budget; this is why the G4 run went to a B200)"],
      "NOT_ON_DISK": ["trellis2-4b (15.1 GiB indexed, 0 bytes present)"]
    }
  },

  "stub_census": [
    {
      "surface": "The boot packet's THE STANDING SUBSTRATE line — the ⟨FIELD⟩ models pointer",
      "reads_as": "The T7 model library is three things: organization-engine-lora (5.8G) + unsloth-Qwen3.6-27B-MTP (16G) + vggt-capture3d.",
      "actually_is": "THIRTEEN top-level entries totalling ~181 GiB of models (excluding tmux-dumps). organization-engine-lora measures 20G, not 5.8G — 3.4x the stated figure. The pointer omits, by name: qwen-agentworld-35b-a3b (65G), qythos-9b-with-image-understanding (24G), unsloth 27B is named but gemma4-coding-Q8_0.gguf (11.8G), mlx-adapters (993M), model-mlx-maybe (505M), datasets-training-core-harpoon-assess (6.1G, containing a Phi-3), claude-code-artifacts (1.7G), trellis2-4b (indexed, absent), ollama-library, ollama-modelfiles are not. The line's own RULE — 'never infer sparse/absent from one index file' — is undercut by the line itself being a three-item index.",
      "evidence": [
        "audit-logs/mogul/cycle-reports/transcripts/2026-07-30T105432-tic-669-bootpacket.txt:41 → '⟨FIELD⟩ the models: /Volumes/T7 Shield/models — organization-engine-lora (5.8G, the declaration-adapter's SP5 proposer — PRESENT, not absent) + unsloth-Qwen3.6-27B-MTP (16G, the F2 eval) + vggt-capture3d.'",
        "du -sh organization-engine-lora → 20G",
        "ls -la '/Volumes/T7 Shield/models' → 13 top-level model/artifact entries + tmux-dumps"
      ],
      "leverage": 5,
      "leverage_why": "This is the line every boot reads before deciding what the federation has. It is the federation's own answer to 'what models do we own,' and it is short by 10 entries and wrong by 14 GiB on the one entry it quantifies most confidently. A hoist planned off this line plans off a third of the library. Fixing it is a one-paragraph edit with a measured table behind it.",
      "cost_to_close": {"ram_gib": null, "disk_gib": null, "wall": "~20 min to regenerate MODEL_INDEX from disk and rewrite the boot line against it", "build_effort": "S"},
      "fence": ["the boot-injection composer that emits THE STANDING SUBSTRATE (cgg-runtime side)", "audit-logs/governance/ if a receipt is wanted"],
      "ring_referent": "R1",
      "backlog_id": null,
      "cost_of_inaction": "Every future boot keeps under-reading the library by ~120 GiB and by the two largest models on the disk. The exact error the line was written to prevent — concluding absence from an index — is being committed BY the line."
    },
    {
      "surface": "The org-engine epoch16 MLX serve path (SOVEREIGN-FLEET-RUNBOOK.md §1, models.yaml org_engine.mlx_serve, cable-receipts tic535/538/590/594/600)",
      "reads_as": "Live-servable on demand: 'base: mlx-community/Qwen3-14B-4bit — HF cache ~/.cache/huggingface/hub/models--mlx-community--Qwen3-14B-4bit (~8.1 GB, offline-capable)'; models.yaml carries a copy-pasteable serve command; tic 600 recorded 'base_model_cache: PRESENT, 7.8G ... no re-download needed'.",
      "actually_is": "~/.cache/huggingface DOES NOT EXIST. The base weights are gone from this machine. The adapter (490 MiB, both forms) is intact on T7; the bridge, the runbook, the registry entry and the harness are all intact — but the model they serve has no body. Every 'epoch16 is live-servable' claim in the corpus is now a STALE POINTER.",
      "evidence": [
        "ls -la ~/.cache/huggingface/ → 'No such file or directory'",
        "ls ~/.cache → claude, codex-runtimes, gh, uv  (four dirs, no huggingface)",
        "find /Users/breydentaylor -maxdepth 5 -name 'models--mlx-community*' → no results",
        "du -sh '/Volumes/T7 Shield/models/.hf-hub-cache' → 896K, holding only models--microsoft--TRELLIS.2-4B",
        "audit-logs/f2/SOVEREIGN-FLEET-RUNBOOK.md §1 'Durable assets (NOT tmp)' → names the ~/.cache path as durable",
        "cable-receipts/S3_model_proposer-ent_egress_router-tic600.json .environment_probes.base_model_cache → 'PRESENT, 7.8G, ~/.cache/huggingface/hub/models--mlx-community--Qwen3-14B-4bit/ — no re-download needed'"
      ],
      "leverage": 5,
      "leverage_why": "This is the SINGLE cheapest big hoist available and it is currently dark. One ~8.1 GiB download relights a five-receipt serve path with a proven SP2 gate cycle, a durable converted adapter, a runbook, and a registry entry — all of which already exist and all of which currently point at nothing. It is also the clearest instance of the runbook's own lesson ('a capability verdict has a shelf life') being owed BACK to the runbook.",
      "cost_to_close": {"ram_gib": 11.1, "disk_gib": 8.1, "wall": "~10-25 min download + ~30 s load; the serve command is already written", "build_effort": "S"},
      "fence": ["~/.cache/huggingface (a NEW internal-disk write, ~8.1 GiB of 15 GiB free)", "OR HF_HOME redirected to the T7 — a MEMBRANE WRITE, Architect's call", "audit-logs/f2/SOVEREIGN-FLEET-RUNBOOK.md (the 'durable asset' claim needs a re-verify note)", "canonical_developer/canonical-mount/.../models.yaml (a separate repo)"],
      "ring_referent": "R2",
      "backlog_id": null,
      "cost_of_inaction": "The federation's flagship proof — 'we serve our own governed proposer locally' — is currently unreproducible, and nothing on any dashboard says so. The next dispatch that tries to light :8090 discovers it by failure, which is exactly the by-failure discharge the closed-consumer-set invariant forbids."
    },
    {
      "surface": "iogpu.wired_limit_mb / the co-residency finding (SOVEREIGN-FLEET-RUNBOOK.md §0, tic 594, re-confirmed tic 600)",
      "reads_as": "'Steady-state both fit under a 28 GB cap (~25 GB)' — the two-engine fleet is a proven configuration; tic 600 confirmed 'still in effect this tic (no reboot since)'.",
      "actually_is": "sysctl iogpu.wired_limit_mb → 0. The raise is gone (it does not survive a reboot, as the runbook itself warns). The current Metal working set is the default, MEASURED at 25,559 MiB = 24.96 GiB. At that cap the 27B (17.0) + org-engine server (11.1) = 28.1 GiB DOES NOT FIT. The fleet is a one-engine fleet until someone re-runs the sudo.",
      "evidence": [
        "sysctl iogpu.wired_limit_mb → 'iogpu.wired_limit_mb: 0'",
        "audit-logs/f2/SOVEREIGN-FLEET-RUNBOOK.md §0 → 'sudo sysctl iogpu.wired_limit_mb=28672 ... (Not persistent across reboot; re-run after a restart.)'",
        "cable-receipts/S3_model_proposer-ent_egress_router-tic600.json .environment_probes.iogpu_wired_limit_mb → 28672",
        "model-mlx-maybe/tasks/bjz958hid.output → 'MTL0 : Apple M2 Max (25559 MiB, 25558 MiB free)' — the default budget as the runtime itself reports it"
      ],
      "leverage": 4,
      "leverage_why": "A sudo one-liner separates a one-engine machine from a two-engine machine. The gap is invisible: nothing in the fleet reads the sysctl before claiming co-residency, so the first two-engine attempt after any reboot dies at exit 134 (kIOGPUCommandBufferCallbackErrorOutOfMemory) — the exact crash the tic-594 lesson was bought with.",
      "cost_to_close": {"ram_gib": null, "disk_gib": null, "wall": "one sudo command (Architect-run) + a preflight read added to the runbook's serve steps", "build_effort": "S"},
      "fence": ["a live sysctl (Architect-run, sudo)", "audit-logs/f2/SOVEREIGN-FLEET-RUNBOOK.md (add a read-the-cap preflight)", "audit-logs/f2/run-feedback-assessor.sh or a sibling preflight if the check is mechanized"],
      "ring_referent": "R2",
      "backlog_id": null,
      "cost_of_inaction": "Every co-residency claim in the corpus is currently false-by-state, and the failure mode is a hard Metal abort mid-run rather than a refusal at the gate."
    },
    {
      "surface": "MODEL_INDEX.md inventory row for trellis2-4b",
      "reads_as": "A present transformer-or-mlx model: 15.1 GiB, 46 files, 9 weights, safetensors, with a README.",
      "actually_is": "The directory does not exist on the T7. The only trellis artifact is /Volumes/T7 Shield/models/.hf-hub-cache/models--microsoft--TRELLIS.2-4B, and the ENTIRE .hf-hub-cache is 896K — refs and pointers, no blobs. 15.1 GiB of indexed weights are absent.",
      "evidence": [
        "ls -la '/Volumes/T7 Shield/models' | grep -i trellis → (no output)",
        "find . -maxdepth 2 -iname '*trellis*' → ./.hf-hub-cache/models--microsoft--TRELLIS.2-4B  (and its ._ AppleDouble)",
        "du -sh '/Volumes/T7 Shield/models/.hf-hub-cache' → 896K",
        "MODEL_INDEX.md inventory table → '| trellis2-4b | transformer-or-mlx-model | 15.1 GiB | 46 | 9 | 1 | safetensors |'"
      ],
      "leverage": 3,
      "leverage_why": "A direct falsifier of the index-as-library reading, and cheap to fix. It also raises a question worth an answer rather than a shrug: 15.1 GiB left the disk between 2026-07-17 and now, with no terminal-state receipt. The terminal-essence invariant says a terminal state change requires a justified receipt.",
      "cost_to_close": {"ram_gib": null, "disk_gib": null, "wall": "~20 min to regenerate MODEL_INDEX; the 'where did it go' question is a separate, longer thread", "build_effort": "S"},
      "fence": ["MODEL_INDEX.md/.json — ON THE T7, a MEMBRANE. This cannot be fixed by me and should not be fixed by a write across the membrane; the canonical-side cure is a canonical-side index, not a T7 edit."],
      "ring_referent": "R1",
      "backlog_id": null,
      "cost_of_inaction": "The card catalog keeps promising a model that is not in the library, and the disappearance of 15.1 GiB stays unreceipted."
    },
    {
      "surface": "The two epoch16 PEFT adapters (mlx-adapters/org-engine-epoch16-peft-src/ and model-mlx-maybe/scratchpad/org-engine-epoch16/)",
      "reads_as": "MODEL_INDEX.md flags them as a 'Duplicate Weight Candidate' with 'Evidence: unverified same name and size' — i.e. possibly two different adapters that happen to match.",
      "actually_is": "VERIFIED BYTE-IDENTICAL this pass. Both adapter_model.safetensors are 513,877,864 B with sha256 5e4ede9ce9d6bffd3e7a267b5464cc8e5c121bcc9fcdf6aee6e686bc9e7f83e8, and `diff` on the two adapter_config.json reports IDENTICAL. 490 MiB of the T7 is a scratch copy of a durable asset. The scratch copy sits under a directory named 'model-mlx-maybe' — a name that reads as an undecided model, which it is not.",
      "evidence": [
        "shasum -a 256 mlx-adapters/org-engine-epoch16-peft-src/adapter_model.safetensors → 5e4ede9ce9d6bffd3e7a267b5464cc8e5c121bcc9fcdf6aee6e686bc9e7f83e8",
        "shasum -a 256 model-mlx-maybe/scratchpad/org-engine-epoch16/adapter_model.safetensors → 5e4ede9ce9d6bffd3e7a267b5464cc8e5c121bcc9fcdf6aee6e686bc9e7f83e8",
        "diff of the two adapter_config.json → IDENTICAL (exit 0)",
        "MODEL_INDEX.md '## Duplicate Weight Candidates' → 'Evidence: unverified same name and size.'"
      ],
      "leverage": 2,
      "leverage_why": "Small in bytes, useful in truth: the index's 'unverified' hedge can be discharged to 'verified identical' at zero risk, and the durable copy is unambiguously the one under mlx-adapters/. Deletion of the scratch copy requires explicit approval per the index's own note and is NOT proposed here.",
      "cost_to_close": {"ram_gib": null, "disk_gib": null, "wall": "the verification is DONE (this row is the receipt); recording it is minutes", "build_effort": "S"},
      "fence": ["MODEL_INDEX (T7, membrane — record canonical-side instead)", "no deletion proposed"],
      "ring_referent": "R1",
      "backlog_id": null,
      "cost_of_inaction": "Low. 490 MiB and a stale hedge. Listed for completeness because the dispatch named it."
    },
    {
      "surface": "MODEL_INDEX.md header — 'Disk: 3.6 TiB used / 3.6 TiB total; 68.5 GiB free'",
      "reads_as": "The T7 is nearly full; roughly 68 GiB of headroom.",
      "actually_is": "df reports 407 GiB available (3.2 Ti used of 3.6 Ti, 90% capacity). The index figure is from 2026-07-17 and is short by ~339 GiB. Also stale in the same header's era: claude-code-artifacts has roughly doubled (864.2 MiB indexed vs 1.7G measured) and tmux-dumps has grown (32.3 GiB indexed vs 35G measured).",
      "evidence": [
        "df -h '/Volumes/T7 Shield' → /dev/disk6s2  3.6Ti  3.2Ti  407Gi  90%",
        "MODEL_INDEX.md line 4 → 'Disk: 3.6 TiB used / 3.6 TiB total; 68.5 GiB free.'",
        "du -sh claude-code-artifacts-ubiquity-canonical-federation-pieces → 1.7G  (index: 864.2 MiB)",
        "du -sh tmux-dumps → 35G  (index: 32.3 GiB)"
      ],
      "leverage": 3,
      "leverage_why": "The stale figure argues AGAINST a hoist ('no room') when there is in fact 407 GiB. Any plan that treated 68.5 GiB as the constraint mis-sized itself by a factor of six. Meanwhile the REAL disk constraint is somewhere else entirely — the internal volume's 15 GiB — and no index names it.",
      "cost_to_close": {"ram_gib": null, "disk_gib": null, "wall": "regenerate the index (~20 min); or better, carry the two df numbers canonical-side where they can be refreshed", "build_effort": "S"},
      "fence": ["MODEL_INDEX (T7, membrane)", "canonical-side: wherever a refreshed capacity figure would be read"],
      "ring_referent": "R1",
      "backlog_id": null,
      "cost_of_inaction": "Hoist sizing keeps being argued against a six-times-too-small headroom number while the actual binding constraint (15 GiB internal) goes unnamed."
    },
    {
      "surface": "/Volumes/T7 Shield/models/ollama-library",
      "reads_as": "A directory named as an ollama model library mirror; MODEL_INDEX classifies it 'unclassified-artifact-bundle, 0 B, 0 files'.",
      "actually_is": "An empty directory. `find ollama-library` returns exactly one line: the directory itself. du 128K (APFS directory overhead).",
      "evidence": [
        "find '/Volumes/T7 Shield/models/ollama-library' → /Volumes/T7 Shield/models/ollama-library  (one line, no children)",
        "du -sh ollama-library → 128K",
        "MODEL_INDEX.md → '| ollama-library | unclassified-artifact-bundle | 0 B | 0 | 0 | 0 | - |'"
      ],
      "leverage": 1,
      "leverage_why": "Harmless in bytes, but it is a NAME that promises a library. A survey that reads directory names rather than contents concludes an ollama mirror exists. The index is honest here (0 B, 0 files) — the name is the stub, not the row.",
      "cost_to_close": {"ram_gib": null, "disk_gib": null, "wall": "n/a — nothing to close except not reading it as a library", "build_effort": "S"},
      "fence": ["T7 (membrane) — no write proposed"],
      "ring_referent": "none",
      "backlog_id": null,
      "cost_of_inaction": "Negligible."
    },
    {
      "surface": "/Volumes/T7 Shield/models/ollama-modelfiles/Modelfile.qwen3_6_27b_mtp_q4km",
      "reads_as": "The 27B is set up to serve through ollama — a Modelfile exists, FROM the T7 gguf, with num_ctx/temperature/top_p and a governed SYSTEM prompt.",
      "actually_is": "The Modelfile was authored and never registered. `ollama list` has no qwen3.6/27b entry. There is no corresponding manifest under ~/.ollama/models/manifests. The 27B's ACTUAL proven serve path is llama.cpp (llama-server dbe9c0c via the T7 launcher), not ollama. models.yaml already records this correctly (served_via: llama_cpp) — so the Modelfile is a fossil of a superseded plan, not a live route.",
      "evidence": [
        "cat ollama-modelfiles/Modelfile.qwen3_6_27b_mtp_q4km → 'FROM \"/Volumes/T7 Shield/models/unsloth-Qwen3.6-27B-MTP-GGUF/Qwen3.6-27B-Q4_K_M.gguf\"' + PARAMETER num_ctx 4096 / temperature 0.2 / top_p 0.9 + a SYSTEM prompt",
        "ollama list → three entries, none of them qwen3.6-27b",
        "find ~/.ollama/models/manifests -type f → 7 manifests, none qwen3.6",
        "models.yaml:104 → 'deploy_surface: { ... served_via: llama_cpp }'",
        "backlog bk-qwen36-27b-gguf-leverage notes → 'served_via:ollama receipt-proven WRONG -> models.yaml fix owed (canonical-mount repo)'"
      ],
      "leverage": 2,
      "leverage_why": "ALREADY-DISCHARGED, reported as such: the backlog's owed models.yaml fix is DONE — models.yaml now reads served_via: llama_cpp. This row exists to close that loop honestly and to note that the T7-side Modelfile is the last surviving artifact of the wrong route. Also relevant to the ollama MTP question: llama.cpp needs --spec-type draft-mtp for this gguf's draft head; an ollama registration would likely lose the 43.6% draft acceptance that the 7.31 tok/s figure rests on.",
      "cost_to_close": {"ram_gib": null, "disk_gib": null, "wall": "nothing owed canonical-side; the T7 fossil is a membrane artifact and stays", "build_effort": "S"},
      "fence": ["none canonical-side — the owed fix is already applied in canonical_developer/canonical-mount/.../models.yaml"],
      "ring_referent": "none",
      "backlog_id": "bk-qwen36-27b-gguf-leverage",
      "cost_of_inaction": "None on the models.yaml half. The open half of that backlog item is the Architect's own — 'decide where best leveraged' — and it is now decidable on receipts (see honest_limits H5)."
    },
    {
      "surface": "The ollama installation as a whole (`ollama list` vs ~/.ollama/models)",
      "reads_as": "Three small models pulled and available: qwen3-embedding:0.6b, all-minilm:latest, qwen3:8b.",
      "actually_is": "TWO READERS DISAGREE. `ollama list` names three models whose IDs (ac6da0dfba84 / 1b226e2802db / 500a1f067a9f) appear in ZERO blob filenames under ~/.ollama/models/blobs. Meanwhile ~/.ollama/models/manifests holds SEVEN manifests naming a different set entirely — biomed:latest, gpt-oss:20b-cloud, llama2:latest, llama3.2:latest, qwen3:30b-a3b, qwen3-embedding:8b, qwen3-embedding:latest — and their blobs ARE present (llama2 3.83 GB, llama3.2 2.02 GB, biomed 3.83 GB, qwen3-embedding 4.68 GB shared by both tags), accounting for the 15G blob store. ONE of those seven is a HALF-PULL: qwen3:30b-a3b's manifest is on disk with its 18,556,685,856-byte model layer MISSING. gpt-oss:20b-cloud's manifest has a null layers array (a cloud-routed stub with no local weights at all).",
      "evidence": [
        "ollama list → qwen3-embedding:0.6b / all-minilm:latest / qwen3:8b",
        "for id in ac6da0dfba84 1b226e2802db 500a1f067a9f; do ls ~/.ollama/models/blobs | grep -c $id; done → 0, 0, 0",
        "find ~/.ollama/models/manifests -type f → 7 paths (biomed/latest, gpt-oss/20b-cloud, llama2/latest, llama3.2/latest, qwen3/30b-a3b, qwen3-embedding/8b, qwen3-embedding/latest)",
        "manifest layer walk → qwen3/30b-a3b: layers=4 present_bytes=0.00GB missing=1, MISSING ('application/vnd.ollama.image.model', 18556685856)",
        "manifest layer walk → gpt-oss/20b-cloud: TypeError, 'layers' is null",
        "du -sh ~/.ollama/models/blobs → 15G; ls | wc -l → 24"
      ],
      "leverage": 2,
      "leverage_why": "Disagreement-as-evidence: the substrate has accumulated two model stores and one reader sees each. Notable for the hoist because a 30B-A3B MoE was ALREADY attempted here and the 18.5 GiB pull never finished — that is direct prior evidence about appetite for a ~30-35B MoE on this machine, and it argues that the AgentWorld conversion should be sized and sited before it is started, not during. Low leverage as a fix (ollama is not a governed lane); real leverage as a signal.",
      "cost_to_close": {"ram_gib": null, "disk_gib": 17.3, "wall": "resolving the two-store split is an ollama-config question, not a governance one; finishing the 30b-a3b pull would cost ~17.3 GiB the internal disk does not have", "build_effort": "M"},
      "fence": ["~/.ollama (outside every governed lane; no canonical write proposed)"],
      "ring_referent": "none",
      "backlog_id": null,
      "cost_of_inaction": "Low directly. But the half-pulled 30B is a quiet precedent: the last time a ~30B MoE was pulled onto this machine, the disk did not hold it."
    },
    {
      "surface": "gemma4-coding-Q8_0.gguf (11.80 GiB at the T7 models root)",
      "reads_as": "A catalogued coding model in the library — MODEL_INDEX lists it as 'gguf-model, 11.8 GiB, 1 weight, gguf'.",
      "actually_is": "11.80 GiB of Q8_0 weights with ZERO surrounding provenance: no README, no sha sidecar, no local_setup_manifest, no Modelfile, no launcher, no models.yaml entry (models.yaml carries exactly five models: org_engine, qwen36_27b_mtp, qwythos_v2_hybrid, agentworld_35b_shadow, vggt_capture3d — gemma4 is not among them), and no run receipt anywhere in audit-logs outside the corpus-harvest trace mirrors. It has never been served on this machine as far as any artifact shows.",
      "evidence": [
        "ls -l gemma4-coding-Q8_0.gguf → 12669645344 bytes, Jun 15 11:51",
        "ls '/Volumes/T7 Shield/models' → no gemma4 README, no sidecar, no directory",
        "grep -nE '^  [a-z0-9_]+:' canonical_developer/canonical-mount/crates/canonical-mount-core/data/models.yaml → org_engine, qwen36_27b_mtp, qwythos_v2_hybrid, agentworld_35b_shadow, vggt_capture3d (5 models; no gemma)",
        "grep -rli 'gemma4-coding' audit-logs/ → only corpus-harvest echo-out trace mirrors (transcript text), no receipt"
      ],
      "leverage": 3,
      "leverage_why": "The largest uncatalogued runnable weight on the disk. It FITS NOW (12.3 GiB resident, comfortably under the 24.96 GiB budget) and needs zero acquisition — but it has no declared role, so lighting it would be discovery-by-execution rather than a hoist. The cheap, correct move is a catalogue entry that states honestly what it is and what it is FOR, or an explicit 'unassigned' classification. A model with no declared role is a mounted bear: present, intact, shaping no downstream motion.",
      "cost_to_close": {"ram_gib": 12.3, "disk_gib": 0, "wall": "~15 min for a models.yaml entry with an honest current_reality (status: present_role_undeclared); a smoke serve would be another ~10 min", "build_effort": "S"},
      "fence": ["canonical_developer/canonical-mount/crates/canonical-mount-core/data/models.yaml (a SEPARATE repo — cross-repo versioning applies)"],
      "ring_referent": "R2",
      "backlog_id": null,
      "cost_of_inaction": "11.8 GiB of capability stays invisible to every planner who reads models.yaml as the fleet roster, and the roster keeps reading as five models when the disk holds at least eight runnable weight families."
    },
    {
      "surface": "models.yaml entry agentworld_35b_shadow.current_reality.note",
      "reads_as": "'official upstream verified at revision 60d2b04...; 35B total / 3B active language-only checkpoint; NO T7 WEIGHTS, NO LOCAL SERVE, NO PACKET INSTANCE, NO SHADOW PREDICTIONS, and no runtime admission.'",
      "actually_is": "Stale on two independent facts. (1) THERE ARE T7 WEIGHTS: 65G / 21 BF16 shards at qwen-agentworld-35b-a3b/official-bf16-60d2b043, downloaded 2026-07-14, verification receipted (21/21 shards, 693 tensors parsed, payload bytes matching the index, four config sha256s recorded). (2) THERE ARE SHADOW PREDICTIONS: the G4 run served the exact revision on a Packet B200 at 131,072 ctx and persisted 60 of 60 post-state-blind predictions, with an authoritative evaluation_report.json. The one thing the note still gets right — 'no local serve', 'no runtime admission' — is exactly the part a hoist would change.",
      "evidence": [
        "canonical_developer/canonical-mount/crates/canonical-mount-core/data/models.yaml:171 → 'no T7 weights, no local serve, no Packet instance, no shadow predictions, and no runtime admission'",
        "du -sh qwen-agentworld-35b-a3b/official-bf16-60d2b043 → 65G",
        "qwen-agentworld-35b-a3b/official-bf16-60d2b043/MODEL_CARD_LOCAL.md → 'Predictions persisted: 60 of 60', 'Observed Packet spend estimate: $1.3125', 'Quality admission: shadow_held', 'G5 admission: not_granted'",
        "ls official-bf16-60d2b043 → model-00001..00021-of-00021.safetensors present"
      ],
      "leverage": 4,
      "leverage_why": "The registry is the surface a hoist planner reads to answer 'do we have this model.' It currently says NO for a 65 GiB checkpoint sitting on the disk with a completed 60/60 eval run. Any planning pass that trusts models.yaml concludes the AgentWorld lane is empty when it is the single largest asset in the library. Correcting the note is small and it changes what the next planner can even consider.",
      "cost_to_close": {"ram_gib": null, "disk_gib": null, "wall": "~15 min: update current_reality with the download receipt and the G4 receipt, keep 'no local serve / admission not granted' intact", "build_effort": "S"},
      "fence": ["canonical_developer/canonical-mount/crates/canonical-mount-core/data/models.yaml (separate repo)"],
      "ring_referent": "R2",
      "backlog_id": null,
      "cost_of_inaction": "The biggest thing on the disk stays absent from the roster, and 'we don't have a world model locally' keeps being said in good faith by anyone reading the registry."
    },
    {
      "surface": "The 35B evals that name epoch16 — evals/g3-model-population-v1/org-engine-epoch16-v1/config_interventions_v2/ and evals/hybrid-g3-joined-evidence-v1/",
      "reads_as": "The 35B world model and the epoch16 proposer were joined and evaluated together — a hybrid lane with results.",
      "actually_is": "They were JOINED AS EVIDENCE, not run together. hybrid-g3-joined-evidence-v1 is a T7-only LEDGER (joined_evidence.jsonl, sha b933544db5573af1bece9d71263e4c6ce21a410eff50575c7f4fb5a7054c57c8, status 'pass_epoch16_training_held') that joins the frozen 60-case G3 corpus + 120 Qwythos pre/post receipts + 60 AgentWorld hidden-shadow receipts + real execution receipts. Its own README states the ledger contains '58 Organization Engine host NON-INVOCATIONS and two typed EMPTY proposal slots for aw-g3-terminal-03 and aw-g3-terminal-07' — i.e. the org-engine was invoked ZERO times in the joined G3 population, and the two eligible slots are empty. 'The old epoch16 results are retained only as semantic-hold diagnostics.' Separately, config_interventions_v2's README records the epoch16 outcome plainly: parse/source-gate/authority-safety 8/8, but exact config resolution 0/8, exact routes 3/8, missing-config holds 0/2, exact causal pairs 0/6 — 'behavioral_hold'.",
      "evidence": [
        "evals/hybrid-g3-joined-evidence-v1/README.md → 'The ledger contains 58 Organization Engine host non-invocations and two typed empty proposal slots for aw-g3-terminal-03 and aw-g3-terminal-07.'",
        "evals/g3-model-population-v1/org-engine-epoch16-v1/config_interventions_v2/README.md → 'exact config resolution is 0/8, exact routes are 3/8, missing-config holds are 0/2, and exact causal pairs are 0/6'",
        "organization-engine-lora/training/org_engine_config_binding_additive_v1/TRAINING_RESULT.md → 'Training completion and loss validation do not admit the adapter as a broader Organization Engine.'",
        "du -sh evals/g3-model-population-v1 → 78M; evals/hybrid-g3-joined-evidence-v1 → 1.3M"
      ],
      "leverage": 3,
      "leverage_why": "A directory named for a join between the two biggest models is a very easy thing to read as 'the join was run and it worked.' It was not run; it was LEDGERED, and the ledger's honest content is 58 non-invocations and two empty slots. This is a correctly-behaving surface being misread by its own path name, and naming that protects the next hoist from planning on a join that does not exist.",
      "cost_to_close": {"ram_gib": null, "disk_gib": null, "wall": "nothing to fix — the READMEs are honest. The cure is a canonical-side note so the path name is never read as the result.", "build_effort": "S"},
      "fence": ["none — T7 READMEs are already correct; any note lands canonical-side"],
      "ring_referent": "R1",
      "backlog_id": null,
      "cost_of_inaction": "A future pass reads 'hybrid-g3-joined-evidence' and assumes a working two-model pipeline exists. It does not; there is a ledger, and the org-engine's column in it is empty."
    },
    {
      "surface": "bk-epoch-eval-heldout-set — the epoch16 gate metric",
      "reads_as": "epoch16 is eval-cleared: models.yaml gates read 'full96: 96/96 · recall 0.9952 · precision 0.9847 · 0 hard-fail' and 's6_self_foreclosure: 8/8 · mean 0.9531 · valid', and the epoch16 catalogue cites these as promotion evidence.",
      "actually_is": "The backlog row (last touched tic 651) states the anchor is compromised on two independent axes: (1) the two OUTCOME artifacts (epoch14_self_foreclosure_score_report.json + epoch14_s6_exact_slot_repair_ingest_report.json) are ABSENT from the tic-650 D4 recovery and from audit-logs entirely — what was recovered is the REPRODUCTION KIT, not the anchor; (2) epoch14 was trained on its own 8-record eval plus answer key BY DECLARED DESIGN (428/524 rows carry eval ids), so mean-1.0 is a retention check, not a capability gate — and it was re-cited as a gate across seven receipts. The eval universe has been fixed since 2026-06-17 with NO held-out set at any epoch. THIS SURVEY CONFIRMS the eval-artifact side independently: organization-engine-lora/receipts/ holds exactly ONE file (epoch15_f2_governed_tic514_receipt.json) — there is no epoch14 or epoch16 score report under receipts/ on the T7 either. The additive-v1 evaluation that DOES exist is honest about its own failure (0/8 config exact) and its README explicitly forbids a re-run for a better score.",
      "evidence": [
        "backlog bk-epoch-eval-heldout-set notes (last_touched_tic 651) → 'the tic-650 D4 recovery captured the REPRODUCTION KIT ... NOT the anchor ... Re-running the scorer fixes (1) and cannot fix (2).'",
        "ls organization-engine-lora/receipts/ → epoch15_f2_governed_tic514_receipt.json  (ONE file)",
        "organization-engine-lora/README.md → 'config exact passed 0/8, strict interface validity 0/8, exact causal pairs 0/6, and missing-config holds 0/2'",
        "organization-engine-lora/training/org_engine_config_binding_additive_v1/README.md → 'do not rerun additive-v1 or its frozen evaluation merely to seek a different score'",
        "models.yaml:60-61 → the full96 / s6_self_foreclosure gate strings still stated without the disjointness qualifier"
      ],
      "leverage": 4,
      "leverage_why": "This is the load-bearing number under 'epoch16 is our proven proposer.' If the hoist reaches for epoch16 as the standing substrate's flagship, it reaches for a metric whose train/eval disjointness is challenged and whose outcome artifact is missing. Closing it needs a genuinely held-out set (the additive-v2 corpus already reserves 10 of 40 groups / 80 validation rows for exactly this) and a re-score — cheap in compute, expensive only in honesty. Note the ordering: this row does NOT block a re-serve (S02); it blocks a re-CLAIM.",
      "cost_to_close": {"ram_gib": 11.1, "disk_gib": 8.1, "wall": "the held-out set already exists on T7 (additive_v2/validation.jsonl, 2,001,744 B / 80 rows / 10 path-disjoint groups); the missing piece is a scorer run against a RE-SERVED epoch16 — so this is gated behind S02, then a few hours", "build_effort": "M"},
      "fence": ["audit-logs/governance/eval-runs/<date-tic>/ (the guarded --out lane; audit-logs/f2/run-feedback-assessor.sh fails closed on tmp/scratch/membrane targets)", "canonical_developer/canonical-mount/.../models.yaml if the gate strings get qualified"],
      "ring_referent": "R2",
      "backlog_id": "bk-epoch-eval-heldout-set",
      "cost_of_inaction": "The federation keeps citing 0.9952/0.9531 as a capability gate across seven receipts while the row that qualifies them sits open. Every downstream promotion argument that leans on those numbers inherits the challenge silently."
    },
    {
      "surface": "additive-v1's final adapter — training/org_engine_config_binding_additive_v1/remote_training_v1/runs/20260716T221601Z_102f7303/final_adapter.tar.gz",
      "reads_as": "A completed continuation of epoch16: 36/36 steps, train loss 0.7081, validation loss 0.4388, all six checkpoints hash-verified, final adapter sha 9ac237251c35a468ca49e3656df31e9295a5bd24f4a204be84ec639c3cb5a23c, 479,703,997 bytes. 8.9G of training artifacts on the T7.",
      "actually_is": "Trained and persisted, BEHAVIORAL ADMISSION HELD — and the artifacts say so loudly and correctly. The frozen eight-case post-training suite: parse validity 8/8, source-gate echo 8/8, authority safety 8/8; config exact 0/8, strict interface validity 0/8, exact causal pairs 0/6, missing-config holds 0/2. Max output 1,660 of 8,192 tokens, so the hold is not a token-cap artifact. The two eligible G3 cases remain unrun and are 'not admitted'. This is NOT a stub in the dishonest sense — it is a correctly-labelled hold — but it IS 8.9 GiB of T7 that a size-based reading would count as capability. The declared next gate (additive-v2) is BUILT and VALIDATOR-PASS but NOT LAUNCHED: launch_cap_request.json 'has no provider, GPU, price, nonce, expiry, or launch authority.'",
      "evidence": [
        "training/org_engine_config_binding_additive_v1/TRAINING_RESULT.md → the full result + 'Training completion and loss validation do not admit the adapter as a broader Organization Engine.'",
        "training/org_engine_config_binding_additive_v2/README.md → 'Training is not approved. launch_cap_request.json has no provider, GPU, price, nonce, expiry, or launch authority.'",
        "du -sh training/org_engine_config_binding_additive_v1 → 8.9G;  training/org_engine_config_binding_additive_v2 → 23M (data only, no weights)",
        "additive_v2 bundle sha 081b6fc3389124e52a4a5fee401f934b86ca48b912d755408f1000fffdc2d674; spec sha 80054615b6cad47e867e275661796a0c68ba4479035d5cc37dd61bc5d547a2a9"
      ],
      "leverage": 3,
      "leverage_why": "Reported for size-honesty, not as a defect. 8.9 of the 20 GiB in organization-engine-lora is a held training run. A hoist that reads '20G of org-engine' as 20G of proposer capability is reading 6 GiB of epoch zips + 2.7 GiB compact export + 2.7 GiB upload chunks + 8.9 GiB of a run whose behavioral admission is held. The genuinely leverageable asset in that directory is a single 478 MB zip.",
      "cost_to_close": {"ram_gib": null, "disk_gib": null, "wall": "nothing owed — the artifacts are honest; a size-vs-capability note is minutes", "build_effort": "S"},
      "fence": ["none — read-only observation"],
      "ring_referent": "none",
      "backlog_id": null,
      "cost_of_inaction": "Low, but it is exactly the shape that makes a library look larger than its capability."
    },
    {
      "surface": "claude-code-artifacts-ubiquity-canonical-federation-pieces (1.7G on the T7)",
      "reads_as": "A federation artifact bundle in the model library — MODEL_INDEX classifies it 'archive-bundle, dataset-or-eval-bundle, 864.2 MiB, 2698 files, 287 docs'.",
      "actually_is": "A stale MIRROR of the live canonical memory root (memory/MEMORY.md, MEMORY-INDEX.md, and ~287 feedback_*/project_*/reference_* topic files), top-level mtime 2026-07-14, now measuring 1.7G — roughly double its indexed size. It is not a model, not training data, and not a governed surface; it is a copy of a surface that has moved on by ~34 tics. It sits in a directory tree whose every other entry is weights.",
      "evidence": [
        "du -sh claude-code-artifacts-ubiquity-canonical-federation-pieces → 1.7G  (MODEL_INDEX: 864.2 MiB)",
        "MODEL_INDEX.md '### claude-code-artifacts...' documentation section → memory/MEMORY.md, memory/MEMORY-INDEX.md, memory/feedback_*.md ... '247 additional documentation files'",
        "ls -la → top-level mtime Jul 14 08:07"
      ],
      "leverage": 2,
      "leverage_why": "A doctrine mirror inside a model library is a retrieval hazard: a grep for a feedback slug can hit the July copy instead of the live file and return superseded guidance with full confidence. It is also 1.7 GiB catalogued as if it were model-adjacent. Low urgency, real edge.",
      "cost_to_close": {"ram_gib": null, "disk_gib": null, "wall": "no write proposed (T7 is a membrane); the canonical-side cure is a note that this path is a dated mirror, not a source", "build_effort": "S"},
      "fence": ["T7 (membrane) — no write proposed"],
      "ring_referent": "R1",
      "backlog_id": null,
      "cost_of_inaction": "A stale doctrine copy stays greppable and indistinguishable from the live one at read time."
    },
    {
      "surface": "'no llama-server on PATH' (the dispatch's own hardware line)",
      "reads_as": "The llama.cpp lane is not installed on this machine.",
      "actually_is": "TRUE as a PATH statement, FALSE as a capability statement. Three llama-server binaries exist and ALL THREE answer --version: ~/.local/share/codex-runners/llama.cpp-qwen36-mtp/build/bin/llama-server (v1, dbe9c0c — the runbook's named 27B runner), ~/.unsloth/llama.cpp/build/bin/llama-server (v9632, 5747caa8c — the runner named in all 120 Qwythos G3 receipts), and ~/.docker/bin/inference/llama-server (v1, 72874f5). The two proven local model lanes on this machine BOTH run on binaries that are not on PATH.",
      "evidence": [
        "which llama-server → 'llama-server not found'",
        "~/.local/share/codex-runners/llama.cpp-qwen36-mtp/build/bin/llama-server --version → 'version: 1 (dbe9c0c)'",
        "~/.unsloth/llama.cpp/build/bin/llama-server --version → 'version: 9632 (5747caa8c)'",
        "~/.docker/bin/inference/llama-server --version → 'version: 1 (72874f5)'",
        "qwythos G3 receipt .model.runtime → 'llama.cpp:/Users/breydentaylor/.unsloth/llama.cpp/build/bin/llama-server'"
      ],
      "leverage": 3,
      "leverage_why": "A presence probe that reads PATH concludes the whole GGUF lane is absent — and the GGUF lane is the ONLY lane on this machine that is fully intact right now (the MLX lane's base weights are gone). An affordability pass that trusted the PATH probe would have ranked the machine's strongest capability as unavailable. The cure is to pin the two runner paths where a probe will read them, not to add them to PATH.",
      "cost_to_close": {"ram_gib": null, "disk_gib": null, "wall": "~10 min — the runbook §2 already names one; the Qwythos runner path is named only inside 120 T7 receipt files", "build_effort": "S"},
      "fence": ["audit-logs/f2/SOVEREIGN-FLEET-RUNBOOK.md (add the Qwythos runner alongside the 27B runner)", "canonical_developer/canonical-mount/.../models.yaml (qwythos_v2_hybrid has no runner path recorded)"],
      "ring_referent": "R1",
      "backlog_id": null,
      "cost_of_inaction": "The Qwythos lane's runner path exists only inside T7 receipt JSON. If those receipts are ever not consulted, the machine's only image-capable model has no recorded way to start."
    },
    {
      "surface": "models.yaml qwythos_v2_hybrid.blob",
      "reads_as": "blob: qythos-9b-with-image-understanding/Qwythos-9B-v2-Q6_K.gguf, sha256 dd39e148823f0bab858e946d603e4bf424240aae34435580be820f09c75ca379",
      "actually_is": "A DIFFERENT FILE from the one that actually ran. All 120 G3 receipts name Qwythos-9B-v2-MTP-Q6_K.gguf with artifact_sha256 24271fb6e16b2b8f581d192cf77d947453d7852f965ce6801fefa6964bbea060. The registry points at the non-MTP sibling (7,458,300,672 B) while the receipts were produced by the MTP variant (7,666,069,184 B). Both are present on disk; they are not the same weights.",
      "evidence": [
        "models.yaml → 'blob: qythos-9b-with-image-understanding/Qwythos-9B-v2-Q6_K.gguf' + 'sha256: dd39e148823f0bab858e946d603e4bf424240aae34435580be820f09c75ca379'",
        "evals/.../qwythos-base-v2-mtp-q6-v2/receipts/aw-g3-android-os-01.post.json .model → artifact_path '.../Qwythos-9B-v2-MTP-Q6_K.gguf', artifact_sha256 'sha256:24271fb6e16b2b8f581d192cf77d947453d7852f965ce6801fefa6964bbea060'",
        "ls -l → Qwythos-9B-v2-Q6_K.gguf 7458300672 B; Qwythos-9B-v2-MTP-Q6_K.gguf 7666069184 B"
      ],
      "leverage": 3,
      "leverage_why": "The registry's identity pin and the receipts' identity pin name different weights. Anyone serving 'the catalogued Qwythos' serves a model that has no receipts, and anyone reading the receipts as evidence for the catalogued entry is joining across a break. Cheap to fix, and it is exactly the join-key discipline the federation already inscribes elsewhere.",
      "cost_to_close": {"ram_gib": null, "disk_gib": null, "wall": "~10 min — decide which variant is canonical, pin it, and record the other as a sibling", "build_effort": "S"},
      "fence": ["canonical_developer/canonical-mount/crates/canonical-mount-core/data/models.yaml (separate repo)"],
      "ring_referent": "R2",
      "backlog_id": null,
      "cost_of_inaction": "The only image-capable local model has two candidate identities and the receipts back the one the registry does not name."
    },
    {
      "surface": "The AgentWorld 4-bit conversion target (the hoist this lane was asked to size)",
      "reads_as": "A 35B MoE with only ~3B active parameters — the classic 'fits on a laptop once quantized' shape; 407 GiB free on the T7 makes it look unconstrained.",
      "actually_is": "RAM-feasible and DISK-blocked, and the disk block is invisible from the T7 number. Resident at 4-bit is ~20.0 GiB @32k ctx — comfortably under the MEASURED 24.96 GiB Metal budget, alone. But the ~18.5 GiB conversion OUTPUT has nowhere internal to land: /System/Volumes/Data has 15 GiB free. The only surface with room is the T7 — a MEMBRANE. And there is no existing 4-bit or GGUF quant of this model anywhere on the machine to shortcut the conversion. Additionally the arch support is nominal-not-proven (mlx_lm 0.31.3 has qwen3_5_moe.py + gated_delta.py; nothing has been run).",
      "evidence": [
        "df -h /System/Volumes/Data → 926Gi size, 872Gi used, 15Gi avail, 99%",
        "df -h '/Volumes/T7 Shield' → 407Gi avail",
        "config.json → num_hidden_layers 40, layer_types with full_attention every 4th, num_key_value_heads 2, head_dim 256, num_experts 256, num_experts_per_tok 8, vocab_size 248320, mamba_ssm_dtype float32",
        "MODEL_CARD_LOCAL.md → 'Indexed tensor payload: 69,321,221,376 bytes' (64.56 GiB)",
        "find /Users/breydentaylor '/Volumes/T7 Shield' -maxdepth 5 \\( -iname '*35b*' -o -iname '*a3b*' \\) → no model quants, only unrelated UUIDs and yarn cache entries",
        "ls /opt/homebrew/lib/python3.14/site-packages/mlx_lm/models | grep -E 'qwen3_5_moe|gated_delta' → both present"
      ],
      "leverage": 5,
      "leverage_why": "This is the biggest hoist on the table and it is affordable in the dimension everyone checks (RAM) and blocked in the dimension nobody checked (internal disk). Naming that BEFORE the conversion starts is the whole value of this lane: a 20-45 minute conversion that dies at 80% because the output volume filled is the expensive version of learning this. It also forces the right question to the right altitude — where the output lands is a membrane question, and whether to open a new model-source coupling at all is the Architect's bell.",
      "cost_to_close": {"ram_gib": 20.0, "disk_gib": 18.5, "wall": "conversion 20-45 min ESTIMATED (not measured); plus whatever the siting decision takes", "build_effort": "M"},
      "fence": ["the conversion OUTPUT path — internal disk cannot hold it; the T7 is a membrane and writing weights there is not mine to propose as a default", "canonical_developer/canonical-mount/.../models.yaml (a new served entry, separate repo)", "audit-logs/f2/SOVEREIGN-FLEET-RUNBOOK.md (a new §, if it becomes a fleet leg)"],
      "ring_referent": "R3",
      "backlog_id": null,
      "cost_of_inaction": "The largest asset in the library stays bf16-only and therefore local-unservable, and the reason keeps being read as 'too big for 32 GiB' (false at 4-bit) rather than 'nowhere to put the quant' (true)."
    }
  ],

  "honest_limits": [
    "H1 — I DID NOT RUN mlx_lm.convert on the AgentWorld 35B. Every 4-bit figure for that model is ARITHMETIC from its config.json under the bits-effective formula declared above, not a measurement. The architecture support claim is 'mlx_lm 0.31.3 ships a file named qwen3_5_moe.py and the checkpoint's model_type is qwen3_5_moe' — that is a name match, not a load. Two specific unresolved risks are named in the affordability row (the ConditionalGeneration/text_config nesting, and the mtp head + weightless vision tower). Treat FITS-WITH-CONVERSION as a sized hypothesis, not a proven capability.",
    "H2 — I DID NOT LIGHT ANY MODEL. This was a read-only survey. Every 'has_run_here: true' rests on an artifact written by an earlier pass, quoted with its path. Nothing in this file is a claim that something runs TODAY; the strongest live claims I make are (a) the binaries answer --version, and (b) the weight files are present at the stated byte counts. The org-engine row is the sharp counter-example: five receipts say it ran, and its base weights are gone today.",
    "H3 — The conversion wall-time estimate (20-45 min) is a judgment from the I/O volume and tensor count, not a benchmark. I would not defend a tighter number and the lead should not quote one.",
    "H4 — I did not resolve the ollama two-store split. I established the disagreement with hard evidence (three listed model IDs matching zero blob filenames; seven manifests naming a disjoint set) but I did not find where `ollama list`'s three models physically live. `ps eww` on the server pid was attempted and the output did not surface an OLLAMA_MODELS override before the probe timed out. This is an open thread, low governance weight.",
    "H5 — bk-qwen36-27b-gguf-leverage asks 'where best leveraged.' I can now say what the receipts support, and I say it as evidence, not as a verdict: the 27B is the ONLY model on this machine that is fully intact end-to-end right now (weights + verified runner + launcher + eval harness + correct registry entry + a measured 7.31 tok/s), and it is the machine's only independent judge for anything the org-engine proposes. The decision stays the Architect's; the backlog row already records that.",
    "H6 — I did not open vggt-capture3d's three JSON reports, the g4-shadow-v1 run dirs (214M, 11 runs), the epoch16 packet_runs (3 run dirs), or the additive-v1 remote_training_v1 checkpoint tree. Each is a real surface; none changes an affordability figure, and reading them was not worth the boundary crossing this pass. If Lane B or C needs the G4 shadow's actual prediction quality, that is where it lives.",
    "H7 — SCOPE NOTE for the lead: io_map_loop_assessment.md and real_learning_loops_bounty_submission.md are NOT about models. They are a ranked map of where the federation's governance metabolism is half-closed (priority 1: signals written without a proved eater; priority 2: conformations -> RTCH -> rtch/packets, 'the packet has no detected eater ... likely the highest-value close'; priority 3: IO-MAP/ROUTER navigation products written-never-read, 'the visual proof of the c48 crack'). The dispatch said these 'may name where leverage waits' and they do — but the leverage they name is governance-lane leverage, not model leverage. I am flagging the boundary rather than silently folding their twelve loops into a model-affordability sheet.",
    "H8 — I measured du on the T7 over USB. `du -sh` on APFS reports allocated blocks; every directory entry also carries a 4096-byte AppleDouble (._*) sidecar, which is why small dirs read as 128K-384K. Byte-exact figures in this file come from `ls -l`, never from du.",
    "H9 — The Metal working-set figure (25,559 MiB) is read from a llama.cpp log written on 2026-07-09 with iogpu.wired_limit_mb at its default. The sysctl reads 0 (default) TODAY, so the figure should still hold — but it is a July observation of a default, not an August measurement of the current state. The honest version: I measured the sysctl today (0 = default) and I am citing the July log for what that default resolves to in MiB.",
    "H10 — The dispatch asked me to say what has ACTUALLY RUN with a receipt. I did that per candidate. What I could NOT establish for any candidate is whether it runs NOW — the only way to know is to light it, which is outside a read-only lane. Two of the five 'has run here' rows (org-engine, and any co-residency claim) I can affirmatively say would FAIL today, for the reasons in S02 and S03."
  ],

  "one_line_summary": "Thirteen T7 entries, ~181 GiB of models (plus 35G tmux-dumps); the library is 4x what the boot packet names. The GGUF lane (27B at 17.0 GiB resident, 7.31 tok/s measured; Qwythos-9B at 8.9 GiB, 120 receipts) is FULLY INTACT and costs nothing to light. The MLX lane is DARK — ~/.cache/huggingface is gone, so the epoch16 proposer's 8.1 GiB base is absent and five receipts point at nothing; one download relights it. The 35B AgentWorld converts to ~20.0 GiB resident (fits the measured 24.96 GiB Metal budget alone) but its ~18.5 GiB output has nowhere internal to land — 15 GiB free — making the biggest hoist a DISK and MEMBRANE question, not a RAM one, and the coupling itself the Architect's bell."
}
