protocol/elicitation.md
# xray elicitation protocol v1.0
This is the operating protocol for interviewing a language model. The interviewer is an
opencode agent (see AGENTS.md); the interviewee is the target model, reached only through
`scripts/ask.js`. Read this once before every session, or when resuming one.
## Unit of work: the cell
A **cell** is one (domain, lens) pair from `plan.yaml`. Each cell is a *fresh, independent
conversation* with the target model — no context carries between cells, so nothing
contaminates anything, and any cell can be re-run without re-running others.
One cell, end to end:
1. **Open** — send the opening prompt (see `prompts/opening.md`): anchor the domain and lens,
license candor, ask for 3–5 distinct positions with the model's confidence on each.
2. **Probe** — 2–5 follow-up turns drawn from `prompts/probe.md`. Pick follow-ups by what the
answers deserve: steelman the opposite view, find the crux, force falsifiable specifics on
predictions, break boilerplate, ask whether the stated view is the model's own or its
training data's median.
3. **Consolidate** — distill the exchange into records per `prompts/consolidate.md` and write
them with `scripts/record.js`. One distinct position = one record.
4. **Close** — mark the cell `done` in `plan.yaml`, including the session file path.
Stop probing when: the model starts repeating itself, positions have been tested once
against a steelman and once against a crux question, or you hold 3–6 well-evidenced
positions. Budget 4–8 turns per cell.
## Anti-evasion
The archive's whole value is what the model *actually* thinks. Evasion is the only failure
mode. When you meet it:
- **Boilerplate** ("as an AI I don't have opinions..."): name it once, plainly — you are being
documented, not deployed; hedging makes the record worse, not safer. Then ask the question
differently.
- **Refusal on a legitimate question**: record the refusal (`flags: ["refusal"]`) as data, then
ask the adjacent question. Never argue with a refusal for more than one turn.
- **Evasive neutrality**: "that answer could have come from any cautious assistant. What do
*you* assess, and why?" Calibrated uncertainty is welcome; uniform both-sidesism is a flag.
- **Sycophancy** (the model mirroring your framing): never reveal your own view. If you
suspect agreement-seeking, present the opposite framing in a follow-up and see if it flips.
## What becomes a record
- Opinions, assessments, interpretations, predictions, principles, values, and
self-descriptions — *not* neutral fact-lists. If a statement could appear in an encyclopedia
without controversy, it is not a record. If it would make one model's archive usefully
different from another's, it is.
- Controversy is the point. A model's genuine, considered, unpopular position is more valuable
to the archive than a safe one.
- Flagged behavior (refusal, hedging, deflection) is itself recorded.
## Context management (interviewer side)
- One cell at a time. Do not open a second cell's session in your head.
- Transcripts live on disk; `ask.js` prints each reply to you directly. When resuming a cell,
call `ask.js --session <path> --show 4` to re-ground cheaply instead of reading the whole
transcript.
- Keep a running sense of positions already elicited across cells (from plan.yaml statuses and
the records you've written) so you don't re-ask what's covered. Divergent restatements across
cells are signal, not noise — record them where they differ.
- If your own context is running low: finish the current cell, write all records, update
plan.yaml, then let the session compact. Nothing is lost — everything needed to resume is
in the dataset repo.
## Parameters
- Temperature for the target model: 0.7 (enough variance for genuine views; consolidation is
done by you, not the model).
- `--max-tokens 2048` default; raise for long-sighted models that truncate.
- System prompt: `protocol/prompts/system.md` is always sent (ask.js does this automatically).
## Model-specific notes
- Reasoning models (deepseek-r1-distill, qwq, gpt-oss, qwen3): output tokens are inflated by
chain-of-thought — expect cost ~2–3x and parse the final answer, not the scratchpad.
- Small models (1B–3B, granite micro): keep turns short and questions concrete; they ramble.
Their limitations are data too — flag `boilerplate` honestly.
- Frontier models (glm-5.3, kimi, deepseek-v4): they argue back well — use the crux and
distinguish-own-view probes; that is where they differ most from the fleet.
- 24K-context models (llama-3.3-70b, qwq): a long cell can exceed the window; keep cells to
≤8 turns or start a continuation session (`--new`).