BTC / LAB

Model B calibration

TWAP champion projection and terminal probability.

ConnectingWaiting for data
Two frozen calibration experiments observe live inputs alongside the raw models. Existing trading signals remain active. Calibration does not yet control orders.
ROUND
TIME LEFT
UP · BID / ASK

Does calibration help?

Lower Brier and log loss are better. Direction accuracy counts correct sides of 50%; it is not trade win rate. Raw and calibrated scores use the same timestamps within each comparison, with equal weight per round.

Native TWAP price distribution

The price champion already learns residual quantiles. An additional rank calibration was tested at +3, +5, +10 and +30 seconds. Its held-out CRPS worsened, so this adjustment is not applied to live price forecasts or either probability experiment.

Method, sources and limits

TWAP champion projection: retains the trained 30-second median move and widens uncertainty with square-root time for longer horizons. Calibration then maps that projected probability to observed official five-minute results. This remains an extrapolation, not a newly trained 300-second price path.

B terminal: the current expiry-aware terminal model receives a separate positive-slope logistic correction that varies with remaining time. It stays tied to the exact terminal version used in training. A different terminal champion pauses this experiment until it is refitted.

Historical test: whole rounds are split chronologically: 50% fit, 25% method selection, 25% test. Training labels must have arrived before the next partition. Price forecasts are reconstructed from the champion available then and reconciled to its recorded live score. The historical analysis is exploratory; confidence intervals resample independent rounds and do not correct for multiple experiments.

Live forward: selected methods are refitted on all labels available at the recorded cutoff, then frozen. Only subsequently issued predictions count. A round needs at least 42 five-second observations and six in each minute. Partial or interrupted rounds are counted separately. Scores use up to 14 days of newly recorded predictions.

Estimated edge: probability minus the currently displayed ask and estimated taker fee, per share held to settlement. The midpoint chart is indicative. Neither these edges nor settlement scores simulate available depth, the 500 ms execution delay, slippage, or resale; they are not P/L.

Sources: B's live TWAP forecast and terminal APIs; official outcomes from the local TWAP label ledger; Polymarket bid/ask quotes from the existing Rust book feed. Missing or stale inputs produce gaps. A new price champion is flagged as unseen if absent from the calibration fit.

Inspect current evidence JSON ↗