source.fact.ngo a coherence.ngo project

docs/cost-estimate.md

raw ↗ · AGPL-3.0

# Cost estimate — Cloudflare Workers AI, Sep 2026 All figures from `node scripts/cost.js` (pricing per 1M tokens from `scripts/models.js`, sourced from Cloudflare's published Workers AI pricing). Workers AI bills at $0.011/1,000 Neurons; the free allocation is 10,000 Neurons/day (~$0.11/day) — negligible at these scales. Frontier models (glm-5.x, kimi, deepseek-v4) require a paid billing method or AI Gateway credits. ## Interview budget (standard depth) ~73 cells × ~6 turns ≈ **1.0M input / 0.4M base output tokens per model**; reasoning models inflate output ×2–3 (chain-of-thought is billed output). ## Pilot (5 models: llama-3.1-8b-fp8, qwen3-30b-a3b-fp8, gemma-4-26b, gpt-oss-120b, glm-5.3) ``` @cf/meta/llama-3.1-8b-instruct-fp8 $0.267 @cf/qwen/qwen3-30b-a3b-fp8 $0.319 @cf/google/gemma-4-26b-a4b-it $0.220 @cf/openai/gpt-oss-120b $0.950 @cf/zai-org/glm-5.3 $3.16 [paid billing] PILOT TOTAL $4.92 ``` glm-5.3 dominates the pilot: 1.4/4.4 $/M in/out. Dropping it to glm-5.3-flash (0.15/0.5) cuts the pilot to ~$1.75; using prompt caching (cached input 0.26 vs 1.4 $/M) reduces glm-5.3 further if sessions hit the cache. ## Full catalogue sweep (26 text-gen models, standard depth) **Total ≈ $37.78** (run `node scripts/cost.js --all` for the live per-model table). - Cheap tier ($0.02–0.35/model): granite micro, llama 1B/3B/8B-fp8, qwen3-30b-a3b, gemma-4-26b, glm-4.7-flash - Mid tier ($0.35–1.20/model): gpt-oss-20b/120b, llama-4-scout, mistral-small, sea-lion, nemotron, qwen2.5-coder, llama-3.3-70b, glm-5.3-flash, deepseek-v4-flash - High tier ($1.30–5.40/model): qwen3.8-27b, qwq-32b, deepseek-r1-distill (worst value: 0.497/4.881 $/M with ×3 reasoning output ≈ $5.38), kimi-k2.6/k2.7-code, deepseek-v4-pro, glm-5.2/5.3 ## Sensitivity | Scenario | Cost | | --- | --- | | Minimal depth (8 domains × 2 lenses + self-model) | ~$11 total sweep | | Standard depth (baseline) | ~$38 total sweep | | Deep depth (all 6 lenses × all 24 domains, per cell count ×2.5) | ~$95 total sweep | | Standard depth, 2× verbosity headroom | ~$75 total sweep | Practical ceiling for the entire catalogue at any sane depth: **well under $100**. Re-eliciting the whole catalogue annually (to measure fleet drift): same numbers again.