kalmari.app / Strategies
Strategy

Sextant — methodology

concluded experiment · runs Paused — was Politics Mon 10:00 PT +3 more (never fired on cron; paused since it was created 2026-07-15)

Status: shadow (not auto-copy) · Source: audit · Current version: 1 · Model: openai/gpt-5.6-sol (OpenAI flagship, via Hermes + OpenRouter)

Navigate by the sun. Sextant is Keel's methodology with GPT-5.6 Sol as the probability judge: the ensemble pipeline (pcc-hermes-ensemble) run with a roster of one — deterministic candidate filter → one deep-research Hermes sample → the v3 sizing policy as code (entry-band screen, fractional Kelly, 15% cluster cap, cash reserve).

Reasoning effort: xhigh (since 2026-07-15, before any scored runs — [[tide]]'s single default-effort run showed the family anchoring to market price with no research). The effort is applied via HERMES_REASONING_EFFORT (written to hermes' config.yaml), proven per run by reasoning_tokens > 0 in the usage report logged to Cloud Run (a WARN fires if it comes back 0), and recorded on every signal row as model openai/gpt-5.6-sol@xhigh.

The question it answers

Keel's live experiment is "is Fable 5 a better judge than Opus was?" — Sextant (with siblings [[ballast]] and [[tide]]) widens that to "is Anthropic's frontier model a better Kalshi judge than OpenAI's?" Everything except the judgment is held constant: same candidate universe, same risk policy, same resolution machinery, same weeks.

Design choices

Schedule & tracks

One hour after the corresponding Keel run so all judges see near-identical market state (staggered 30 min from its siblings so per-run OpenRouter key-usage deltas — the cost telemetry — don't overlap):

Feature parity with the Keel fleet

Publishes an HTML report to audits-site, posts picks + rejects to the bot (shadow-scoped AUDIT_VERSION=v2, full t0/poll/settle marks + calibration in the v2 Audit Lab), archives run logs to LOGS_ARCHIVE, and reports per-run cost/turns/status to the workflow_runs dashboard. OpenRouter spend is captured as the key-usage delta around the diligence stage and lands in the same cost column the Claude jobs use.

Graduation / kill criteria

Accumulate ~15–20 resolved picks per category (~6–10 weeks), then compare Brier + CLV + % returned against Fable 5 Keel on the overlapping markets. Beat or match Keel → consider feed exposure (ShowInTelegram) and a bigger slice of the audit load; clearly worse → retire the track and keep the calibration data as MODEL_WEIGHTS input for [[chorus]].

See this strategy's live track record

Every pick this strategy has ever published — resolved results, P&L, and calibration vs the market — is public and auditable.

Open in the Strategy Explorer →