# Audit round 5 — external re-audit of PREREG v1.2 (Codex / GPT-5.6 Sol)

Received 2026-07-11. Read-only re-audit of the committed v1.2 objects
(PREREG.md v1.2, deviations.md rows 1–19, design.md, artifacts/, scoring
code). Verdict: **NOT CLEARED** — one finding (R5-M01). Answered by
PREREG v1.3 (design.md alignment; deviations row 20). Report verbatim
below; file paths reference the auditor's checkout
(`/Users/jakubdudek/pollar-labs-wt`).

---

## Verdict

**NOT CLEARED.** v1.2 soundly closes R4-B01, R4-B02, R4-M01, and R4-M03, but R4-M02 is only partial: `design.md` still contains cutoff-only admission and fallback rules that contradict the new stable-alias standard and would trigger the fallback for the production writer whose cutoff is admittedly undocumented. Minimal remaining set: align those stale passages, remove the associated absolute non-exposure claim, log the pre-call deviation, and refreeze. The resulting closure count is **20/22**; Plan-2 implementation gates remain separate.

## Round-4 dispositions

| Finding | Disposition | v1.2 result |
|---|---|---|
| R4-B01 | **CLOSED** | The unsupported global claim is withdrawn. Selection is accurately limited to "outcome-blind with respect to ingest timing only, conditional on present-day membership," with the unreachable historical-membership limitation repeated in the estimand and paper limitations (PREREG.md:155, PREREG.md:815). |
| R4-B02 | **CLOSED** | The second reviewer is a fresh forecast-blind instance receiving only the question packet and named sources; Jakub is barred from YES/NO/VOID decisions; reconciliation and hashing precede scoring. Same-family dependence is honestly disclosed (PREREG.md:603). The earlier sensitive screen remains safe because it must finish before ranking, drafting, and every forecast call. |
| R4-M01 | **CLOSED** | V3 is now an observable two-source publication check, with screenshot/URL evidence. Difficult but source-resolved cases must therefore fall into exhaustive YES/NO rather than discretionary VOID (PREREG.md:408). |
| R4-M02 | **PARTIAL** | PREREG, the dossier header, statusline, and limitations adopt and disclose the weakened stable-alias standard, but two operative `design.md` passages retain the old cutoff-only rule. See R5-M01. |
| R4-M03 | **CLOSED** | Evidence roles are correctly separated: generation metadata establishes identity/route/time/tokens; the session screenshot establishes selected settings; the call-file header establishes intended settings. Settings are invariant within a session, and any change forces a new session and screenshot (PREREG.md:890). |

## 22-item closure

| # | Readiness item | Closure |
|---:|---|---|
| 1 | Immutable protocol, timestamp, deviations | CLOSED |
| 2 | Event eligibility, exclusions, sampling/stopping | CLOSED; sensitive screen and rerank remain Plan-2 gates |
| 3 | Window versus latest deadline | CLOSED |
| 4 | Corpus membership, boundaries, order, cap, merges, hashes | CLOSED, with present-day-membership limitation disclosed |
| 5 | Drafter identity, prompt, parameters, count, deduplication, failures | CLOSED |
| 6 | Keep/edit/strike rubric and decision log | CLOSED |
| 7 | Question template and resolution contingencies | CLOSED |
| 8 | Complete packet to forecasters | CLOSED |
| 9 | Contestant identities, provider, parameters, concurrency, timestamps | **OPEN — R4-M02/R5-M01**; runtime assertions and provenance remain Plan-2 gates |
| 10 | Retries, failures, endpoints, rationale mismatch, denominators | CLOSED; parser tests remain a Plan-2 gate |
| 11 | Status-quo baseline | CLOSED AS REMOVED |
| 12 | Crowd matching, odds, thresholds, priority, missingness | CLOSED |
| 13 | Resolver blinding, review, disagreements, VOID sensitivity | CLOSED |
| 14 | Primary estimand, snapshot, contrast, weighting | **OPEN — R5-M01** |
| 15 | Brier/log formulas, orientation, clipping, ties | CLOSED |
| 16 | Murphy implementation, grouping, bias, uncertainty | CLOSED AS POINT-ONLY |
| 17 | Calibration bins, x-value, minimum n, bootstrap | CLOSED |
| 18 | Signed updating and no-new-article treatment | CLOSED |
| 19 | Aggregate/paired CIs, seed, replicates, multiplicity | CLOSED |
| 20 | Class and crowd-subset claims | CLOSED |
| 21 | Tables, figures, filters, headline rules | CLOSED |
| 22 | Reproducibility bundle | CLOSED AS SPECIFIED; production remains a Plan-2/3 obligation |

## New findings by severity

| Severity | ID | Exact committed text | Finding |
|---|---|---|---|
| Major | R5-M01 | The amended standard is "**documented pre-2026-06-01 training cutoff OR pre-2026-06-01 release of a stable model alias**" (design.md:104). But the same document still says a model whose cutoff is undocumented "**is excluded from the primary comparison**" (design.md:98) and "**if the production model's cutoff cannot be documented, the primary contrast is Fable vs the 0.5 reference**" (design.md:255). It also says pre-June listing means the descriptive aliases "**can[not] contain the June events**" despite the disclosed alias-immutability assumption (design.md:189). | The standard is not stated identically or honestly at every site. The production cutoff is explicitly undocumented, so the stale design rules activate the fallback that PREREG says does not fire. Because PREREG says `design.md` governs intent, execution precedence does not fully cure this contradiction. Replace both cutoff-only rules with the two-branch standard and make the descriptive-arm claim conditional on alias stability; log and refreeze before calling. |

Secret hygiene: **PASS**. The production vendor/model name appears only inside the dossier's `Internal name — NOT FOR PUBLICATION` section. Permitted test: **PASS, 14 tests / 0 failures**.

## Consolidated Plan-2 pre-call checklist

Protocol prerequisite, not a Plan-2 gate:

- Close R5-M01, add deviation row 20, refreeze/tag, and verify the amended snapshot is committed before any RUN call.

Implementation gates:

- Generate the sensitive-screen packet for all 1,667 eligible events from birth material only. Remove ranking/count fields, use opaque randomized review IDs, dual-review every keyword hit, and commit every hit, keep, exclusion, and disagreement.
- Add `sensitive: true` only after both reviews; rerank the canonical manifest offline. Assert exactly 45 unique, eligible, non-sensitive rows under `birthDomains desc / birthArticleCount desc / id asc`.
- Implement the run harness's startup assertion: exact pinned production alias including suffix, deep-equal pinned provider policy, no environment override, abort on mismatch, and ledger the result.
- Add parser tests proving duplicate packet IDs fail that question even when values agree, unknown IDs are ignored and logged, and missing IDs fail only the missing question.
- Generate and hash complete run call files using the full template packet, wrapped corpus, fixed bundling/order, retry rules, and per-arm parameters, including Kimi's 8,000-token cap.
- Preserve every harness attempt, timestamp, raw response, provider metadata, parse disposition, failure, request parameters, cost, and input/output hash.
- For every console attempt, save generation ID, console timestamp, raw output, hashes, and generation metadata. Associate it with a session screenshot showing model, temperature, `max_tokens`, and `response_format`; settings must remain invariant within that session.
- Keep reliability-bin membership fixed before `binOutcomeCIs`; retain seed `20260711`, 10,000 replicates, and the tagged scoring revision. Re-run the frozen test and require 14/14.
- Provision isolated resolution contexts: neither reviewer receives forecasts, ledgers, or arm outputs; Jakub remains process-only. Commit and hash the complete resolution log before scoring.
- Issue no RUN call until the amended tag, final screen log, reranked manifest, startup assertions, parser tests, call files, ledgers, provenance plan, and hashes are committed.
