Files
cruxhub-epistack-contest/case-findings-Opus-4.8.md
T
2026-07-20 02:43:14 +02:00

8.3 KiB
Raw Blame History

Per-case findings (drop-in for the Worked Examples section)

These are drafted to sit directly under each existing ### COVID / ### LHC black holes / ### Eggs heading, above the graph and transcript links. Each is one run of crux-research followed by the full adversarial-graph-review loop, with minimal prompting. Numbers are pulled from the session transcripts in this directory.


COVID — reproducing the 23-orders-of-magnitude spread

The run reconstructed the early-2024 Rootclaim debate (Wilf vs. Miller, judged for zoonosis) as a single shared 47-node evidence tree — raw facts only, hung under one lab-origin spine — rated by eight perspectives: the six analysts Scott Alexander compares (Rootclaim, Miller, both judges, Scott, Daniel Filan) plus Michael Weissman and the executing model's own puppet. The headline test the prompt set was whether the famous ~23-order-of-magnitude spread across independent Bayesian analyses would emerge as a propagated output rather than an asserted one. It did: the six analysts rating the same tree propagated to root odds spanning from Miller at P(lab) ≈ 10⁻²⁰ to Rootclaim at ≈ 0.999 — about 23 orders of magnitude, with every perspective landing within ~1 OOM of its published number. Removing Miller collapses the spread to ~6 OOM, matching Scott's "~7 without Peter."

Three things the map surfaced that prose summaries of the debate do not:

  1. The entire disagreement is manufactured at essentially one node. Every analyst agrees the Huanan-market early-case clustering exists; they disagree by ~44 dB (50 dB for Miller, 6 dB for Rootclaim) on what it is worth. The debate is a likelihood dispute over agreed facts sitting on a shared prior that favors nature — not an empirical dispute.
  2. Miller's ~10⁻²⁰ is structurally exposed as over-multiplication. The tree shows it is built by stacking a 50 dB market factor, 48 dB two-introduction, 40 dB no-backbone and 35 dB base-rate factor as if independent — the exact correlated- evidence failure the independence intermediates are designed to catch. This vindicates Scott's "trolling" bracket node by node rather than by assertion.
  3. The tooling's own crux metric has a blind spot. get_disagreements keys on confidence divergence, but the market crux lives in edge-weight divergence — so the built-in metric under-ranks the single most important crux. That is a genuine, honestly reported limitation of the current Assessment layer.

The adversarial review then hardened the graph over multiple rounds. It caught four real defects — a sign/text contradiction that had both lab-side analysts propagating too hot, a duplicated datum (the Sept-2019 database withdrawal counted under two clusters), Miller's incoherent null grandchild weights, and a correlated-timing over-count — plus a shared- source correlation between the market and two-introduction clusters (both trace to Worobey/Pekar). Removing the correlated over-counts honestly shrank the reported spread from ~23 to ~21 OOM and moved Rootclaim from +31 dB to +27 dB (554:1), almost exactly its in-debate 533:1 figure. Two honest caveats stand: the judges' and Filan's ratings are inferred from their written decisions, not stated by them, and the shrinkage shows the raw 23 was itself slightly inflated by correlation.


LHC black holes — separating the physics from the epistemics

The prompt deliberately split the question into two hypotheses: H1, the physics claim that micro black holes posed no meaningful risk, and H2, the epistemic claim that declaring the case "closed" was justified at the certainty expressed. A 105-agent deep-research workflow (every finding adversarially verified against primary sources) grounded a 34-node graph rated by five perspectives: the executing model, the LSAG/CERN 2008 position, the Rössler/Wagner litigants, the x-risk epistemologists (Ord, Hillerbrand & Sandberg), and a historical-precedent lens built from Trinity (LA-602) and the RHIC strangelet review.

The result the map makes legible: the disagreement was never about black holes. Every perspective that engages the evidence converges above 0.90 on the physics (own posterior 0.97; even the safety case's sharpest published critics sit ~0.90), yet they diverge by up to 0.80 on whether the institutional certainty was warranted (own posterior on H2: 0.45). The physics rests on three layered arguments — production needs speculative extra dimensions; those same theories predict evaporation in ~10⁻²⁷ s; and cosmic rays have run harsher collisions on white dwarfs and neutron stars for billions of years — but the epistemic case has real cracks: the original cosmic-ray argument had a genuine gap (collider products can be slow and trapped) patched only months before startup, the review was CERN reviewing CERN, and the courts dismissed every lawsuit on standing, never on the merits.

The adversarial loop is the most instructive of the three cases because it did not cleanly converge — and reports that honestly. Severity fell monotonically (round 1: two P1 findings; round 3: one P1; rounds 45: P3 only), and the round-5 referee moved to "accept with optional minor revisions" after hand-verifying the entire load-bearing spine, but no round returned zero valid findings, so by the skill's own definition this is a cap-out, not a clean bill of health. The correlated-evidence correction was dramatic: the model's propagated root fell 0.981 → 0.939 → 0.892 across the repairs, closing on its stated 0.85 — the round-1 referee's verdict was that the original 0.98 was "the sound of correlated evidence being summed as if independent." The loop also caught its own adversaries: two findings were rejected as confabulation after the model re-read the cited edges in the live graph and found the claimed defects simply weren't there (this is the hallucination-based check working as intended). Finally, the run hit a real engine limitation — positively-weighted supports can't pull a node below the 50% prior, and there's no way to gate one branch's contribution on another — and filed it as issue #39 (a proposed gate edge relation), which is itself an example of the workflow compounding into product improvements.


Eggs — the everyday question, and a clean convergence

The eggs case is the "ordinary question" of the three: no adversarial debate to reconstruct, so the hard part is deciding what is even worth answering and how we can know it. The run built a graph around a single proposition ("for the general adult population, are eggs net-good / net-bad / fine in moderation") and rated it across four perspectives. Unlike LHC, this one converged in round 5 (of a 5-round cap): the final root sits at a stated 0.85, conservatively below the 0.881 it propagates to, and the round-5 adversary independently reproduced the propagation arithmetic by hand before accepting.

What makes it a good Assessment-layer demonstration is that every finding across rounds 14 was valid — the rejection ledger stayed empty until the final round — and most of them were correlated-evidence catches, the contest's named criterion:

  • The root started as a four-way XOR question and was reframed into a single proposition (a Structure fix).
  • Two studies were double-counted along two paths each; the duplicate edges were archived so each study pays out once.
  • Three overlapping CVD meta-analyses (Frontiers / Godos / Zhao, which share underlying cohorts) were down-weighted from 8/7/6 dB to 6/2/2 dB and cross-annotated with shared-source undercuts — textbook "correlated evidence treated as independent."
  • The LDL-cholesterol surrogate endpoint was capped so a biomarker couldn't stand in for a hard outcome at full strength.
  • An incoherent node (H2 propagating to 0.946 while stated at 0.40) was repaired by adding a propagating attack for the surrogate-to-outcome gap, bringing it back to 0.46.

The run also shows disciplined restraint: one round-5 recommendation was deliberately not applied because its literal remedy (converting an attack to an undercut) would have made the discount stop propagating and re-inflated H2 — undoing an earlier round's fix. That is the builder correctly refusing a plausible-but-wrong suggestion, documented rather than silently dropped.