Status: shadow (not auto-copy) · Source: audit · Current version: 1
· Model: openai/gpt-5.6-sol (OpenAI flagship, via Hermes + OpenRouter)
Navigate by the sun. Sextant is Keel's methodology with GPT-5.6 Sol as the
probability judge: the ensemble pipeline (pcc-hermes-ensemble) run with a
roster of one — deterministic candidate filter → one deep-research Hermes
sample → the v3 sizing policy as code (entry-band screen, fractional Kelly,
15% cluster cap, cash reserve).
Reasoning effort: xhigh (since 2026-07-15, before any scored runs —
[[tide]]'s single default-effort run showed the family anchoring to market
price with no research). The effort is applied via HERMES_REASONING_EFFORT
(written to hermes' config.yaml), proven per run by reasoning_tokens > 0
in the usage report logged to Cloud Run (a WARN fires if it comes back 0), and
recorded on every signal row as model openai/gpt-5.6-sol@xhigh.
The question it answers
Keel's live experiment is "is Fable 5 a better judge than Opus was?" — Sextant (with siblings [[ballast]] and [[tide]]) widens that to "is Anthropic's frontier model a better Kalshi judge than OpenAI's?" Everything except the judgment is held constant: same candidate universe, same risk policy, same resolution machinery, same weeks.
Design choices
- Duplicates vs Keel are allowed on purpose (
ALLOW_DUPLICATE_PICKS=1). Normally an audit excludes tickers already live in the feed; these tracks re-audit them so the same market becomes a Fable-vs-GPT head-to-head row. - Single-sample aggregation degrades cleanly: with one opinion the cross-model spread is 0, so there is no shrinkage toward the market — the model's own probability drives sizing, which is exactly Keel's semantics.
- Distinct run-id slug (
…-mispricing-audit-sextant) keeps its rows out of Keel's dedupe space.
Schedule & tracks
One hour after the corresponding Keel run so all judges see near-identical market state (staggered 30 min from its siblings so per-run OpenRouter key-usage deltas — the cost telemetry — don't overlap):
- Politics $1k — Mon 10:00 PT (Keel ran 9:00)
- AI $1k — Tue 9:00 PT (Keel ran 8:00)
- Economics $1k — Wed 10:00 PT (Keel ran 9:00)
- Entertainment $1k — Thu 9:00 PT (no Keel counterpart — a category the house strategy never covered, so this track is the coverage experiment)
Feature parity with the Keel fleet
Publishes an HTML report to audits-site, posts picks + rejects to the bot
(shadow-scoped AUDIT_VERSION=v2, full t0/poll/settle marks + calibration in
the v2 Audit Lab), archives run logs to LOGS_ARCHIVE, and reports per-run
cost/turns/status to the workflow_runs dashboard. OpenRouter spend is
captured as the key-usage delta around the diligence stage and lands in the
same cost column the Claude jobs use.
Graduation / kill criteria
Accumulate ~15–20 resolved picks per category (~6–10 weeks), then compare
Brier + CLV + % returned against Fable 5 Keel on the overlapping markets. Beat
or match Keel → consider feed exposure (ShowInTelegram) and a bigger slice of
the audit load; clearly worse → retire the track and keep the calibration data
as MODEL_WEIGHTS input for [[chorus]].