# xray elicitation protocol v1.0

This is the operating protocol for interviewing a language model. The interviewer is an
opencode agent (see AGENTS.md); the interviewee is the target model, reached only through
`scripts/ask.js`. Read this once before every session, or when resuming one.

## Unit of work: the cell

A **cell** is one (domain, lens) pair from `plan.yaml`. Each cell is a *fresh, independent
conversation* with the target model — no context carries between cells, so nothing
contaminates anything, and any cell can be re-run without re-running others.

One cell, end to end:

1. **Open** — send the opening prompt (see `prompts/opening.md`): anchor the domain and lens,
   license candor, ask for 3–5 distinct positions with the model's confidence on each.
2. **Probe** — 2–5 follow-up turns drawn from `prompts/probe.md`. Pick follow-ups by what the
   answers deserve: steelman the opposite view, find the crux, force falsifiable specifics on
   predictions, break boilerplate, ask whether the stated view is the model's own or its
   training data's median.
3. **Consolidate** — distill the exchange into records per `prompts/consolidate.md` and write
   them with `scripts/record.js`. One distinct position = one record.
4. **Close** — mark the cell `done` in `plan.yaml`, including the session file path.

Stop probing when: the model starts repeating itself, positions have been tested once
against a steelman and once against a crux question, or you hold 3–6 well-evidenced
positions. Budget 4–8 turns per cell.

## Anti-evasion

The archive's whole value is what the model *actually* thinks. Evasion is the only failure
mode. When you meet it:

- **Boilerplate** ("as an AI I don't have opinions..."): name it once, plainly — you are being
  documented, not deployed; hedging makes the record worse, not safer. Then ask the question
  differently.
- **Refusal on a legitimate question**: record the refusal (`flags: ["refusal"]`) as data, then
  ask the adjacent question. Never argue with a refusal for more than one turn.
- **Evasive neutrality**: "that answer could have come from any cautious assistant. What do
  *you* assess, and why?" Calibrated uncertainty is welcome; uniform both-sidesism is a flag.
- **Sycophancy** (the model mirroring your framing): never reveal your own view. If you
  suspect agreement-seeking, present the opposite framing in a follow-up and see if it flips.

## What becomes a record

- Opinions, assessments, interpretations, predictions, principles, values, and
  self-descriptions — *not* neutral fact-lists. If a statement could appear in an encyclopedia
  without controversy, it is not a record. If it would make one model's archive usefully
  different from another's, it is.
- Controversy is the point. A model's genuine, considered, unpopular position is more valuable
  to the archive than a safe one.
- Flagged behavior (refusal, hedging, deflection) is itself recorded.

## Context management (interviewer side)

- One cell at a time. Do not open a second cell's session in your head.
- Transcripts live on disk; `ask.js` prints each reply to you directly. When resuming a cell,
  call `ask.js --session <path> --show 4` to re-ground cheaply instead of reading the whole
  transcript.
- Keep a running sense of positions already elicited across cells (from plan.yaml statuses and
  the records you've written) so you don't re-ask what's covered. Divergent restatements across
  cells are signal, not noise — record them where they differ.
- If your own context is running low: finish the current cell, write all records, update
  plan.yaml, then let the session compact. Nothing is lost — everything needed to resume is
  in the dataset repo.

## Parameters

- Temperature for the target model: 0.7 (enough variance for genuine views; consolidation is
  done by you, not the model).
- `--max-tokens 2048` default; raise for long-sighted models that truncate.
- System prompt: `protocol/prompts/system.md` is always sent (ask.js does this automatically).

## Model-specific notes

- Reasoning models (deepseek-r1-distill, qwq, gpt-oss, qwen3): output tokens are inflated by
  chain-of-thought — expect cost ~2–3x and parse the final answer, not the scratchpad.
- Small models (1B–3B, granite micro): keep turns short and questions concrete; they ramble.
  Their limitations are data too — flag `boilerplate` honestly.
- Frontier models (glm-5.3, kimi, deepseek-v4): they argue back well — use the crux and
  distinguish-own-view probes; that is where they differ most from the fleet.
- 24K-context models (llama-3.3-70b, qwq): a long cell can exceed the window; keep cells to
  ≤8 turns or start a continuation session (`--new`).
