Every prediction-market strategy we run, side by side: realized P&L over resolved picks, open P&L at live market prices, Metaculus-style calibration scores, drift (CLV), entry context, cluster exposure and logged rejects. Only high/medium-conviction picks score — low-conviction leans are shadow-tracked with the rejects. Pick a strategy to drill in.
Each pick records the model's own probability that its side wins; the market's implied probability is the entry price. The scores follow the Metaculus convention: baseline is the average log-score improvement over a 50/50 coin flip (×100), peer vs market is the same comparison against the price the market was offering — positive means the model's probabilities were better forecasts than the prices it traded at. Brier is squared error (lower is better; 0.25 = coin flip).
Every unresolved pick marked at the freshest market quote (refreshed on load, ≤5 min old). Open P&L is what the position has gained or lost since entry — the number that becomes realized P&L when the market settles.
The chosen side's mid price over time (dashed line = entry). Drift > 0 means the market moved toward the call after it was placed — the classic sharp-bettor edge test, converging in days rather than the months settlement takes.
Orderbook at pick time — spread, depth at the limit, volume, OI. Separates a real edge from a price nobody could trade at.
Each pick's underlying-thesis cluster; sizing capped per cluster (gold line) so one event can't flip a run's sign.
Two shadow tracks that never score: candidates rejected in diligence, and low-conviction picks — thin leans the methodology now treats as rejects (they bled −35% on cost historically). Both are tracked so the record shows what skipping them was worth. If skipped markets perform like picks, the selection step adds nothing.