Status: DROPPED 2026-07-15 (ran once) · Source: audit · Version: 1
· Model: openai/gpt-5.6-luna (OpenAI budget tier, via Hermes + OpenRouter)
Verdict after one run: at default reasoning effort, Luna spent ~80 seconds and $0.05, did no web research, and anchored to the market price on all 12 candidates ("quoted market probability retained") → zero edge, zero picks. The cheap-model hypothesis got a fast answer: budget-tier judgment on this pipeline is market-echo, not forecasting. Jobs and schedulers deleted; the 12 shadow rows stay as the comparison record. Its siblings [[sextant]] and [[ballast]] continue at xhigh reasoning effort.
The moon drives the tides. Keel's methodology with GPT-5.6 Luna as the probability judge — the budget tier at roughly a fifth of Sol's token price. Pipeline, design choices, shadow scoping, and feature parity are identical to [[sextant]]; only the model differs.
The question it answers
The cheap-model hypothesis: if Luna calibrates anywhere near Sol/Terra/Fable once selection and sizing are house code, then judgment is commoditized and the audit load (and any future higher-frequency scanning) can run at ~$2/run instead of ~$12+. If it's clearly worse, that's the measured price of frontier judgment.
Schedule & tracks
- Politics $1k — Mon 11:00 PT (Keel ran 9:00, Sextant 10:00, Ballast 10:30)
- AI $1k — Tue 10:00 PT (Keel ran 8:00, Sextant 9:00, Ballast 9:30)
Staggered 30 min from its siblings so per-run OpenRouter key-usage deltas (the cost telemetry) don't overlap.
Graduation / kill criteria
Same bar as [[sextant]]: ~15–20 resolved picks per category, Brier + CLV +
% returned vs Fable 5 Keel and its siblings on overlapping markets. Even a
negative result is useful — it prices what the frontier premium is worth on
this workload and feeds MODEL_WEIGHTS for [[chorus]].