docs/design.md
# emergence — system model
emergence is the live-events sensor net of fact.ngo: a second-order heap of news reports,
machine-distilled into a structured, browsable representation of what is happening in the
world right now — without claiming verified status for any of it. Fact-checking is a
separate, currently inactive lane. The only standing cost is distillation, budget-capped
at $1/day; observed cost is ~$0.02–0.05/day at current volumes.
## Why "second-order"
The net never observes the world; it observes *reports about* the world. That indirection
is the honest core of the design:
- Every record is distilled from a named newsroom's output, so it inherits real
journalistic signal (some validity).
- No human or council has checked any of it, so it never gets verified status.
- The distiller (glm-5.3-flash) is cheap, fixed-prompt, and neutral-framing — an
interpreter, not a judge.
The site must therefore always present emergence records as labeled **"not fact-checked · summarized from news reports"** (internal `status: "unverified"`). The label language is deliberately plain: "AI-distilled" read as if the AI originated the news; the summaries come from real newsroom reports.
That label is part of the data contract, not a styling choice.
## Lanes
| lane | status | writes | consumer |
| --- | --- | --- | --- |
| unverified | active | scrape + distill + aggregate + gis | site `/emergence/` |
| verified | built, inactive | council.js only, `--activate` gated | site verified placeholder |
A record's `status` is set by the mechanism at birth: `"unverified"` from the distiller,
`"verified"`/`"refuted"`/`"unclear"` only from a council verdict. Nothing upgrades in place.
## Pipeline stages
```
sources/sources.yaml
|
v
scrape.js --feed fetch (RSS/Atom, teaser depth)--> raw/<date>/articles.jsonl
| (content-hash dedupe across all days)
v
distill.js --glm-5.3-flash, fixed prompt------> distilled/<date>/essence.jsonl
| (+ ledger.jsonl, budget-guarded)
v
aggregate.js --fractal topic rollup----------> aggregates/<date>/topics.json
| aggregates/<date>/day.json
v
gis.js --country centroids (vendored)--------> aggregates/<date>/gis.json
|
v
pipeline.js orchestrates all four; cron on the self-hosted remote runs it every 2h UTC and commits.
```
- **Days are UTC.** `today`/`yesterday` lenses derive from partition dates.
- **Idempotent per day.** Every stage can re-run; append-only with content-addressed ids
(sha256 of source+title for articles, of raw-id for essences).
- **Failure is data.** A source that fails is recorded in `day.json` as failed with its
HTTP status; a distillation that fails to parse is recorded as a `parse_failure` record.
## The event layer (dedupe + facets + ranking signal)
Reports are not events: three newsrooms covering one happening used to count three times.
`cluster.js` (between distill and aggregate) partitions the day's essences into **events**,
with two phases:
- **Phase A (free, deterministic)**: union-find over similarity signals — word-set
Jaccard ≥ 0.30, or same subtopic + shared country + Jaccard ≥ 0.18 — produces
candidate clusters.
- **Phase B (AI, budget-guarded)**: one glm-5.3-flash call per candidate cluster resolves
true events and their **facets** — the distinct angles/nuances the reports contribute
("initial strike", "official response", "corroboration"). Parse failures and budget
exhaustion degrade to the deterministic cluster as a facet-less event; nothing is ever
silently dropped, and every ok essence lands in exactly one event (enforced).
Weight feeds the day portrait — corroboration and complexity must carry upward:
```
corroboration = 1 / 1.6 / 2.0 for 1 / 2 / 3+ distinct sources
complexity = 1 + 0.25 x distinct facets
event.weight = corroboration x complexity
topic.score = sum of event weights <- treemap area
```
Counts stay honest alongside: tiles show `N events · M reports`; scores are exposed, not
hidden. A single uncorroborated report is weight 1.0 — present in the portrait, never
inflated. Cost: one call per candidate cluster, ~$0.005/day at current volume.
Events are recomputed each cycle (phase A is deterministic; phase B output is cached per
day in `aggregates/<date>/events.json`). Event ids are seeded from their first member,
stable while membership grows.
**Cross-day continuity (phase C, event threads).** After the day's events are formed, they
are matched against the previous day's (fixed, committed) events: deterministic candidate
pairs (word-set Jaccard, topic/country/subtopic gates), then ONE batched AI call
adjudicates up to 24 shortlisted pairs. A confirmed match chains the event onto
yesterday's thread (`continues`, `thread_id`, `thread_days`; threads inherit identity from
their first day). Matching is one-to-one and greedy by similarity. Without an AI verdict
(budget exhausted, parse failure) only pairs with Jaccard ≥ 0.30 link — conservative,
because a false thread hides novelty while a missed link is cheap to see side by side.
Sustention feeds the ranking: `weight = corroboration x complexity x (1 + 0.1 x
min(thread_days - 1, 5))` — an event in its second day carries 10% more weight than a
fresh one, capped at +50%. Cross-day links never rewrite yesterday's data; yesterday's
event ids are stable once committed.
## Fractal topic model
`ontology/topics.yaml` fixes a two-level topic tree (~9 topics × 4–7 subtopics, plus
`other`). The distiller maps each report to one primary topic (+ optional secondary).
Aggregates roll up per day: topic → subtopic → representative essences. The site can then
show any zoom level: all news, one topic, one subtopic, one day, a rolling window.
Adding topics changes future interviews only (same as xray's ontology rule) — past
records keep their mapping; re-aggregation is always allowed, re-distillation is not.
## GIS lens
The distiller emits `locations: [{name, country_code}]` (ISO-3166 alpha-2). Two
deterministic joins place them — no external geocoding API, zero AI cost, and never
model-emitted coordinates:
- **Country level** (stated limit, fallback): `pipeline/geo.js` — Google's canonical
countries centroids, public data.
- **Named places**: `pipeline/gazetteer.js` — Natural Earth 10m populated places
(public domain, ~7,600 names, regenerated by `pipeline/gen-gazetteer.mjs`). The
gis stage resolves event location names within their stated country; suffix words
like "village" are tolerated. Unresolved names fall back to country placement,
honestly: the map caption separates named-place pins from country-level mentions.
Places fuse **events** (not raw reports), so a pin's size is event weight — the same
corroboration x complexity signal as the day portrait. The site renders pins over the
Equal Earth basemap with the shared zoom camera; the map and its boundary policy
(Natural Earth v5.1.2 lineage, disputed-boundary notes) are preserved unchanged.
## Budget and ledger
- Every Workers AI call appends a line to `emergence-data/ledger.jsonl`:
`{date, script, model, tokens_in, tokens_out, cost_usd}`.
- `lib.js` sums today's spend before each distill call and stops the distill stage when
the next call would cross the cap in `scripts/config.json` (`daily_budget_usd`, default
1.00). `max_articles_per_run` (default 120) bounds backfill spikes.
- Expected steady state: ~40–80 new articles/day → ~60–130K tokens in, ~16–32K out on
glm-5.3-flash ($0.15/M in, $0.50/M out) → **$0.02–0.07/day**.
## The council (inactive)
`council.js --claim "..." --activate` runs the fact-checking deliberation: 3 seats of
`@cf/zai-org/glm-5.3`, 2 rounds (independent first takes → cross-examination → verdict),
producing a verdict record (`verified | refuted | unclear`, confidence, reasoning, what
evidence a human still needs). Sessions are archived; verdicts would land in
`verified/<date>/verdicts.jsonl` and the site would render them in the verified lane.
Cost per claim ≈ $0.02–0.05 — irrelevant while inactive, and priced so activation never
threatens the $1/day envelope at moderate volumes.
## Data layout (emergence-data)
```
raw/<YYYY-MM-DD>/articles.jsonl scraped reports (id, source, url, title, teaser)
distilled/<YYYY-MM-DD>/essence.jsonl unverified distillation records
aggregates/<YYYY-MM-DD>/day.json cycle stats, source health, spend
aggregates/<YYYY-MM-DD>/topics.json fractal rollup
aggregates/<YYYY-MM-DD>/gis.json country points for the map lens
verified/ empty until the council is activated
ledger.jsonl append-only cost accounting
manifest.json mechanism versions, models, ontology revision
```
The site build picks the newest date partition present. No mutable pointers.
## Runtime topology (2026-10-02)
The cycle runs on Cloudflare, not on any single machine:
- **Worker `emergence-net`** (cron `7 */2 * * *`) runs the same pipeline cores as the
local scripts (`pipeline/*` — one brain, two runtimes), against the Workers AI
binding (no token) and an R2 bucket. Each invocation: scrape -> distill (capped
per run; deferred articles carry to the next cycle) -> cluster/threads ->
aggregate -> gis -> `latest.json` snapshot + `status.json`.
- **R2 is the live primary**: fresh data every cycle regardless of any machine's
status. **Git (`emergence-data`) is the durable archive**, synced by
`scripts/sync-from-r2.sh` (the self-hosted remote cron or by hand); the sync is the only writer
of archive commits from live state.
- **The site reads both**: the deployed static snapshot is the no-JS floor;
the JS layer additionally fetches the Worker's `latest.json` and adopts it when
newer — rail meters refresh, records newer than the page build render into a
"live since this page build" group. Manual `POST /run` (Bearer secret) holds
the HTTP connection; cron needs no client and gets the full window.
- **Budget guard unchanged**: the same ledger logic, kept in R2; per-run AI calls
capped (35) so a cycle always completes and free-plan subrequest limits hold.
Workers Paid raises headroom, not semantics.
Local Node scripts remain first-class: same cores, fs storage, useful for
backfills, tests, and the eventual council activation.
## Provenance and licensing
- All live sources are taken at the depth their public feeds offer (titles and
summaries), attributed per record, with links to the original articles. Reuters is
registered but deferred: public RSS retired, site bot-gated, third-party relays'
terms forbid non-personal use. The basis for each source is recorded in the registry
and rendered publicly in the site's About pane.
- Code and curation: AGPL-3.0. Distillate is a mechanical transformation of newsroom
output used at teaser depth with attribution and links; records keep `source_id`.
No license statement ever claims the newsroom content itself.
- The registry (`sources/sources.yaml`) carries, per source, what we take and the basis
on which we use it; `scrape.js` ships that metadata through `sources.json` into the
site's About pane, so the public can audit the source list and its legality from the
live data rather than trusting prose.
## Roadmap: local granularity (not built)
Today the net aggregates global newsrooms. The planned extension is country-level lanes:
per-country or per-region source sets built from reputable local outlets, ingested at
the same teaser depth under the same rules, aggregated into local day portraits and
event threads. This gives finer local resolution without a new mechanism — a lane is a
source registry subset plus a filter in the aggregation stage. Design questions to settle
first: which outlets qualify as reputable per country, how local lanes rank against the
global lane (never just more volume), and how country selection handles contested or
under-reported regions. Not built yet; nothing here should imply it exists.
## Non-goals (v1)
- No full-text scraping, no paywall or robots bypass.
- No same-event cross-source dedupe (topics lens clusters near-dupes naturally; revisit
with a similarity pass if the heap feels repetitive).
- No historical backfill beyond what feeds offer naturally.
- No automated posting of any emergence record as fact anywhere. Ever.