Status: shadow (not auto-copy) · Source: audit · Current version: 1
· Model: openai/gpt-5.6-terra (OpenAI mid-tier, via Hermes + OpenRouter)
Keel's methodology with GPT-5.6 Terra as the probability judge — the
mid-tier of OpenAI's Sol/Terra/Luna family at half Sol's token price. Pipeline,
design choices, shadow scoping, feature parity, and xhigh reasoning effort
(recorded as model openai/gpt-5.6-terra@xhigh, proven by reasoning_tokens
in the run logs) are identical to [[sextant]]; only the model differs.
The question it answers
Where on the price/capability curve does Kalshi-judgment quality saturate? Terra vs Sol isolates whether the flagship premium buys better calibration, or whether mid-tier reasoning is already enough once the candidate filter and sizing policy are house code.
Schedule & tracks
- Politics $1k — Mon 10:30 PT (Keel ran 9:00, Sextant 10:00)
- AI $1k — Tue 9:30 PT (Keel ran 8:00, Sextant 9:00)
- Economics $1k — Wed 10:30 PT (Keel ran 9:00, Sextant 10:00)
- Entertainment $1k — Thu 9:30 PT (Sextant 9:00; no Keel counterpart)
Staggered 30 min from its siblings so per-run OpenRouter key-usage deltas (the cost telemetry) don't overlap.
Graduation / kill criteria
Same bar as [[sextant]]: ~15–20 resolved picks per category, then Brier + CLV + % returned vs Fable 5 Keel and vs its siblings on overlapping markets. The interesting outcome isn't only "does it beat Keel" — it's the shape of the price/calibration curve across Sol/Terra/Luna.