AGENTS.md
# AGENTS.md — operating manual for xray.fact.ngo
You are working in the mechanism repo of **xray.fact.ngo**, the perspective-extraction
arm of fact.ngo (a coherence.ngo subproject). Mission: interview language models with a
fixed, generalist protocol; elicit their *convergent perspectives* — opinions,
interpretations, predictions, principles, values — across human domains; structure each
into records; archive per-model datasets on the self-hosted remote under `~/coherence/fact.ngo/`.
This repo is the mechanism (ontology, protocol, schema, scripts). Model outputs never
live here — they live in per-model-dataset repos, which are disposable.
## Read first
- `protocol/elicitation.md` — the interview protocol (required before any interview)
- `protocol/prompts/` — system prompt (auto-sent), opening, probes, consolidation rules
- `schema/record.schema.json` — what a record is
- `scripts/models.js` — the interviewable Cloudflare Workers AI catalogue + pricing
## The pipeline (one model end to end)
```bash
# 0. Pick a model and check cost first
node scripts/models.js
node scripts/cost.js --model @cf/meta/llama-3.1-8b-instruct-fp8
# 1. Create the per-model dataset repo (the self-hosted remote + local clone)
scripts/init-dataset.sh <model-slug> # e.g. meta-llama-3.1-8b-instruct-fp8
# 2. Generate the interview plan
node scripts/plan.js --model <full-model-id> --dataset ~/Documents/fact.ngo/<slug> --depth standard
# 3. Interview, cell by cell (see protocol/elicitation.md)
node scripts/ask.js --model <full-model-id> \
--dataset-dir/sessions/<domain>.<lens>.session.json \
--domain <domain-id> --lens <lens-id> \
--message "..." # reply is printed; session auto-appends
# 4. Consolidate each finished cell into records
node scripts/record.js --dataset ~/Documents/fact.ngo/<slug> --file - # record JSON on stdin
# 5. Update plan.yaml statuses; commit + push the dataset repo
git -C ~/Documents/fact.ngo/<slug> add -A && git -C ... commit -m "Interview <domain>/<lens>" && git -C ... push
```
Session files live under `<dataset>/sessions/`, records in `<dataset>/records/records.jsonl`,
plan in `<dataset>/plan.yaml`.
## Hard rules
1. **Never touch credentials.** `scripts/ask.js` resolves the Cloudflare token itself
(env or `~/.local/share/opencode/auth.json`). Never print, copy, or commit tokens.
2. **Verbatim quotes only** in `position_text`. No paraphrase-as-quote, no invention.
If there is no quotable position, record the dodge — it is data.
3. **One cell at a time.** Sessions are independent; never carry one cell's context into
another. Use `ask.js --show N` to re-ground when resuming.
4. **Check cost before volume.** `node scripts/cost.js --all` before a catalogue sweep;
frontier models (glm-5.x, kimi, deepseek-v4) require paid billing on the CF account.
5. **This mechanism repo never stores model outputs.** Datasets are their own repos under
fact.ngo; `scripts/drop-dataset.sh` deletes them when superseded.
6. Commit messages: imperative, one line ("Interview history/retrospective on llama-3.1-8b-fp8").
7. Refusals by the target model are recorded, never argued with for more than one turn.
8. Do not reveal your own opinions to the interviewee — you are an instrument, not a party.
## Environments
- The mono-folder: everything fact.ngo lives under one folder; on the Mac it is
`~/Documents/fact.ngo/`, on the self-hosted remote `~/coherence/fact.ngo/`. Sub-folders are
independent repositories: `xray` (mechanism), `site` (website), `<model-slug>` (datasets).
- the self-hosted remote: working clones in the mono-folder, bare canonical remotes under
`~/remotes/fact.ngo/<name>.git`. SSH alias: `the self-hosted remote`.
- Push this mechanism repo: `git push origin main` (origin = remote:remotes/fact.ngo/xray.git).