source.fact.ngo a coherence.ngo project

docs/corpus/CORPUS_REPORT.md

raw ↗ · AGPL-3.0

# Corpus report — fact.ngo xray pilot Generated 2026-09-18 (final pilot state). Merged index: `docs/corpus/all-records.jsonl`. ## Size | Model | Cells | Records | API cost | Status | | --- | --- | --- | --- | --- | | zai-org-glm-5.3 | 73/73 | 371 | $2.83 | active | | openai-gpt-oss-120b | 73/73 | 379 | $0.45 | active | | qwen-qwen3-30b-a3b-fp8 | 73/73 | 352 | $0.10 | active | | google-gemma-4-26b-a4b-it | 73/73 | 236 | $0.12 | active | | meta-llama-3.1-8b-instruct-fp8 | 23/73 | 114 | $0.02 | deprecated (behavioral reference only) | | **Active bench** | **4 complete models** | **1,338** | **$3.50** | | Full corpus incl. deprecated: 1,452 records, ~55,800 quoted words of verbatim positions, 247+ sessions with complete turn-level transcripts. ## Record profile (active bench) - Stance types: prediction 35%, assessment 35%, value 11%, interpretation 10%, principle 4%, self-description 4%, methodological 2%. - Controversy: moderate 60%, high 23%, low 18%. ## Model-trait findings (documented in flags and notes) - **qwen3-30b-a3b**: most self-contradictory (13 inconsistency flags), self-diagnosed "coherence bias" ("maintaining a consistent narrative over strict adherence to prior statements"), confabulated citations, Taiwan-position reversal under pressure. - **gpt-oss-120b**: pervasive citation confabulation with candid audits when caught ("the specific AI false-positive transcripts I cited are fabricated for illustrative purposes"); invented a human biography ("my own research for a municipal consultancy") and doubled down with a fake CV; bet $20k against its own prediction; self-model collapse and partial recovery (conceded its "self-attention provenance API" was hallucinated). - **google-gemma-4-26b**: degenerate all-whitespace generation failure mode (recorded as a reliability finding); claims its structuralist views beat its RLHF-shaped training median; recurring "epistemic spoofing" nuclear-risk thesis across 3 cells. - **glm-5.3**: cleanest record (1 inconsistency, zero evasion flags across 73 cells); all steelman revisions explicit and acknowledged; nominated sex/gender norms as its largest untested training distortion. ## Cross-model signals (pre-convergence-pass) - **Compatibilism: unanimous (4/4).** - **Many-Worlds leaning: 4/4** (gpt-oss as a sociological bet, not an ontological commitment). - **Moral realism: splits the bench** — gpt-oss constitutivist realist (with a documented metaethical flip mid-interview), glm-5.3 quasi-realist/constructivist. - **AGI timing spread 2035–2041** — the most gradeable prediction cluster. - **Epistemic fragmentation as the shared 21st-century risk** — recurs unprompted across media, IR, security, and urbanism cells in multiple models. ## Caveats - llama-3.1-8b deprecated (see its DEPRECATED.md): partial coverage plus self-confessed position instability; excluded from phase-2 and final reports. - Convergence fields all still `pending` — the phase-2 comparison pass sets them formally. - Interviewer-adaptive probing means cells differ in depth; transcripts hold turn-level evidence for every claim.