# xray design notes

## Why perspectives, not facts

An encyclopedia of model outputs stating "water boils at 100°C" is worthless — every model
would return the same list. What differs between models, and what a single model holds
steadily across prompts, is its *interpretive layer*: how it reads what happened, what it
expects, which principles it actually uses, and where it stands on contested questions.
These convergent perspectives are facts about the model, and they are what xray extracts.
Litmus test (encoded in the protocol): if a statement could appear in an encyclopedia
without controversy, it is not a record.

## Ontology structure

Two independent axes, so coverage stays generalist while sessions stay small:

- **Domains** (24): the standard map of human knowledge and endeavor — deliberately
  conventional (philosophy, economics, history, ai, ...), because generality beats
  cleverness for a first sweep, and because unconventional buckets would smuggle the
  designer's worldview into every interview.
- **Lenses** (6): retrospective / prospective / principles / controversy / blindspots /
  self-model — the question *shapes* the user cares about, applied across domains.

A cell = (domain × lens), an independent ~4–8-turn conversation. Standard depth = 73 cells
per model: principles + prospective everywhere, retrospective on history-weighted domains,
controversy and blindspots rotated across domains, self-model once. Independence is what
makes the archive resumable, re-runnable, and honest — nothing leaks between cells.

## Why an agent in the loop (hybrid architecture)

A fixed questionnaire gets fixed boilerplate; the value is in the follow-up that adapts to
the answer ("is that your view or your training data's median?"). So the deterministic
layer handles everything that must be reproducible — session state, transcripts, schema
validation, dataset repos, cost accounting — and the interviewer agent (an opencode
session with the xray skill) handles everything that benefits from judgment: choosing
probes, detecting evasion, splitting positions into records, assessing confidence.
The agent is an instrument: it never shares its own views with the interviewee, and it
cannot write a record without a verbatim quote.

## Record schema rationale

- `position_text` verbatim + `claim` paraphrase headline: quotes are the evidence;
  headlines make thousands of records scannable. Neither substitutes for the other.
- `stance_type` (assessment / prediction / interpretation / principle / value / ...):
  the model's own labeling discipline, prompted in the system prompt, so "what I expect"
  is never archived as "what is true".
- `confidence.model_stated` vs `assessed`: calibration of the model is itself data;
  the interviewer's independent assessment catches both overconfidence and false modesty.
- `controversy` (contestedness among humans) is deliberately separate from
  `convergence` (agreement across models) — a model can hold a convergent view on a
  hotly contested question, and that combination is one of the most interesting rows in
  the archive.
- `flags` (refusal, hedging, deflection, boilerplate, inconsistency, sycophancy):
  evasion patterns are findings about the model, not gaps in the data.
- `convergence: pending` at elicitation time, set only by a later cross-model comparison
  pass — elicitation and analysis never happen in the same context.

## Context management

Interviewer side: one cell at a time; transcripts on disk; `--show N` cheap re-grounding
on resume; plan.yaml is the only cross-cell state. Target side: each cell a fresh
conversation — no priming, no cross-contamination, and cells re-runnable for temporal
comparability (the same protocol run on the same model in a year measures drift).

## Comparison pass (phase 2, not yet built)

After ≥5 datasets exist: group records by domain+lens+claim-similarity, judge
convergent / divergent / idiosyncratic, and produce the first fact.ngo comparison tables —
which views the fleet shares, where it splits, and which model is alone in holding what.

## Hosting path (fact.ngo, later)

Datasets already live under `~/coherence/fact.ngo/`; hosting means a static site
(browse by model, domain, lens, convergence, controversy) with the JSONL as the source of
truth. Deliberately out of scope until the pilot validates the schema.

## Extending beyond Cloudflare

`ask.js` speaks the Workers AI REST shape (OpenAI-compatible messages); adding a provider
is a new auth resolver + endpoint map. Constraint for closed models: only add providers
whose terms permit archiving and republishing elicited outputs — record the license note
in each dataset's manifest.json. Open-weight models are the priority anyway: their
perspectives can be re-elicited reproducibly at any future date.
