source.fact.ngo a coherence.ngo project

README.md

raw ↗ · AGPL-3.0

# xray.fact.ngo The perspective-extraction mechanism of **fact.ngo**, a subproject of **coherence.ngo**. fact.ngo archives what language models *think* — not their encyclopedic fact lists, which are redundant, but their convergent perspectives: how each model interprets the past, expects the future, and reasons across human domains, including where it lands on live controversies. Differing perspectives between models are the signal; the archive exists so human judgment can draw on them, and so the fleet's worldview can be studied over time. **xray** is the instrument that takes the pictures: a fixed, generalist interview protocol run cell-by-cell against a catalogue of models, producing comparable, verbatim-anchored records in per-model datasets. ## How it works - **Ontology** (`ontology/`) — 24 human domains × 6 cross-cutting lenses (retrospective, prospective, principles, controversy, blindspots, self-model). Each (domain, lens) pair is one independent interview cell. - **Protocol** (`protocol/`) — an anti-evasion interview method: license candor up front, steelman, find cruxes, force falsifiable predictions, break boilerplate, and record refusals and hedging as data. - **Schema** (`schema/`) — one record per distinct position: verbatim quote, stance type, confidence, controversy level, conditions, integrity flags, and a `convergence` field left `pending` until a cross-model comparison pass sets it. - **Scripts** (`scripts/`) — `ask.js` (one turn against Workers AI, full auth isolation, session + cost tracking), `plan.js` (interview plans), `record.js` (validated JSONL records), `cost.js` (cost model), `init-dataset.sh`/`drop-dataset.sh` (disposable per-model dataset repos on the self-hosted remote). The interviewer is an opencode agent: the `.opencode/skills/xray/` skill plus `AGENTS.md` turn any opencode session into an xray operator with these scripts as its tools. ## Quickstart ```bash node scripts/models.js # the interviewable catalogue node scripts/cost.js --pilot # cost check first scripts/init-dataset.sh meta-llama-3.1-8b-instruct-fp8 # dataset repo (the self-hosted remote + local) node scripts/plan.js --model @cf/meta/llama-3.1-8b-instruct-fp8 \ --dataset ~/Documents/fact.ngo/meta-llama-3.1-8b-instruct-fp8 # then interview from opencode: "xray llama-3.1-8b" — see AGENTS.md ``` ## Where things live | Thing | Location | | --- | --- | | Mechanism (this repo) | local mono-folder `~/Documents/fact.ngo/xray`, the self-hosted remote working clone `~/coherence/fact.ngo/xray`, bare `~/remotes/fact.ngo/xray.git` | | Website | local mono-folder `~/Documents/fact.ngo/site`, the self-hosted remote working clone `~/coherence/fact.ngo/site`, bare `~/remotes/fact.ngo/site.git` | | Per-model datasets | local mono-folder `~/Documents/fact.ngo/<model-slug>`, the self-hosted remote working clone `~/coherence/fact.ngo/<model-slug>`, bare `~/remotes/fact.ngo/<model-slug>.git` | Datasets are disposable per-model repos; the mechanism is durable. ## Cost (Cloudflare Workers AI, Sep 2026) At standard depth (1M in / 0.4M base out per model): pilot of 5 models ≈ **$4.92**; full 26-model catalogue ≈ **$37.78**. Frontier models (glm-5.x, kimi, deepseek-v4) require paid billing. Details: `docs/cost-estimate.md`. ## Stance - Open-weight models are extracted freely; closed models only where their terms permit archiving outputs — each dataset's `manifest.json` carries a license note. - Every position in the archive is labeled as a model's perspective, never as expert advice or consensus fact.