Terminal experiments.
Four projections. Two direction checks. Two live controls.
Forecasts & UP share price
White line: UP bid–ask midpoint. Shading: bid–ask spread. Market quotes are recorded each second; forecasts every five seconds. Gaps remain visible. Quotes are before fees; the midpoint is not a fill price. Forecast probability is a settlement estimate, not a promise that shares will trade at that price.
Sources and timing
UP quotes come from the same Rust Polymarket order-book recorder used by the paper portfolios. Only fresh, verified quotes for this round's UP token are recorded. Price history starts with this chart update and continues independently of model availability.
The left axis shows probability; the right shows cents per share on the same 0–100 scale. Tap the chart or move the time slider to inspect recorded values. A uses its Binance opening reference; B uses the exact official TWAP opening reference.
A vs TWAP vs share price
Finding the recorded leaders.
Blue: selected A forecast. Green: selected TWAP forecast. White: actual UP midpoint, with the bid–ask spread shaded. The same timestamps and 0–100 scale let you compare all three directly.
How the defaults are chosen
Each family's default has the lowest recorded historical Brier error before the final 30 seconds, including the existing terminal and projection variants. This is an exploratory settlement-forecast ranking, not proof of better share-price prediction. The defaults do not use this round's future prices or outcome.
Choose any variant above. Your selections stay fixed during refreshes; use “Use recorded leaders” to restore the defaults. A smaller gap from today's quote can mean a model follows the market rather than predicts its next move. Quotes and timing use the same verified source as the family chart above.
Compare the probabilities
Metrics and experiment rules
Brier is the average squared difference between UP probability and the official outcome (UP = 1, DOWN = 0). Lower is better. Log loss also penalizes confident mistakes. Direction accuracy is the fraction of predictions on the correct side of 50%, averaged equally across rounds; it is not trade win rate.
Each comparison uses the same observations and qualified rounds for all eight variants. A round needs at least 42 of 60 five-second buckets and coverage in every minute. Within each round, timestamps receive equal weight; each round then receives equal weight.
Held move: retain the champion's 30-second median move and widen its uncertainty with square-root time. Fading continuation: allow further directional movement, fading over 60 seconds and capped below twice the initial forecast move. Both use the price champion throughout the round, with interpolated trained marginals inside 30 seconds. The direction-removed checks subtract each forecast median while preserving its centered distribution. They test how much value comes from the predicted move versus the learned uncertainty.
Probability ranges are quantile rank bounds, not confidence intervals. The score uses the bound nearest 50% when a tail is unidentified. The table's 95% interval resamples whole rounds for the candidate-minus-control Brier difference; it is exploratory, assumes rounds are independent, and does not adjust for multiple comparisons.
All scores use the official Polymarket TWAP60 outcome. A's different Binance reference remains a proxy. Historical forecasts use the champion versions available then; B's price forecasts are reconstructed from stored causal features. Its terminal may embed an older price champion, so this compares complete serving choices, including model recency. The historical replay uses a conservative one-second receipt bound. It does not establish profitable execution or an untouched validation result.
Round by round
Brier before the final 30 seconds. Latest 30 qualified rounds; lower is better.