Across 402 graded contracts since 2026-07-03, we stated an average 58.8% chance and 50.2% of them happened.
This page is the record of how good Kanon’s probabilities are. It is not a record of profit, and there is no claim of one here: over the same contracts the book is -1.5 pts per dollar, which is a loss, and it is published on the terminal beside this. Accuracy and profit are different questions and this page answers only the first.
Across 402 graded contracts since 2026-07-03, we stated an average 58.8% chance and 50.2% of them happened. That difference is the single most important number on this page and it points the wrong way for us, so it goes first.
Expected calibration error is the average distance between what we said and what happened, weighted by how many contracts sat in each band. Lower is better and zero is perfect. Each cell shows how many contracts were graded, the average probability we stated, and the share that actually happened. A band with fewer than five graded contracts says too thin instead of a number, because a rate off three coin flips is not a measurement.
| Lead time | ECE | graded | 0-20% | 20-40% | 40-60% | 60-80% | 80-100% |
|---|---|---|---|---|---|---|---|
| 2 days or less | 0.0739 | 146 | 2·18.9% too thin | 26·31.5% 30.8% | 52·50.2% 44.2% | 46·68.8% 58.7% | 20·87.4% 75.0% |
| up to 2 weeks | 0.1038 | 231 | 3·18.3% too thin | 45·29.6% 22.2% | 68·49.6% 44.1% | 71·70.9% 53.5% | 44·86.5% 77.3% |
| up to 3 months | 0.1779 | 22 | 2·26.6% too thin | 6·52.4% 33.3% | 9·69.2% 88.9% | 5·89.4% 80.0% | - |
| beyond 3 months | 0.2816 | 3 | 1·49.7% too thin | 2·82.9% too thin | - | - | - |
The table above is calibration over the contracts we chose to grade — and choosing is where the error concentrates. Averaged over all of them we stated 58.8% and 50.2% happened, a gap of +8.5 pts (95% interval +2.8 pts to +14.3 pts, and it does not cross zero). The interval comes from resampling whole settlement days rather than individual rows, because contracts that settled together are not independent draws and an interval that pretended otherwise would read as a stronger result than we have.
One aggregate gap can be one bad category, so the same comparison runs inside every slice we can cut: 13 of 13 slices with at least five graded contracts lean the same way. Slices thinner than that are excluded from both sides of that fraction. The widest are these.
| Slice | Cut | graded | we said | it happened | gap |
|---|---|---|---|---|---|
| barrier | structure | 27 | 60.3% | 48.1% | +12.2 pts |
| Crypto | category | 169 | 60.2% | 48.5% | +11.6 pts |
| threshold | structure | 145 | 60.8% | 49.7% | +11.1 pts |
| short | horizon | 231 | 58.9% | 48.5% | +10.4 pts |
| 0.40-0.70 | cost band | 188 | 64.3% | 54.3% | +10.0 pts |
| global | venue | 382 | 58.2% | 49.5% | +8.7 pts |
The same rows predicted a gain of 22.1 cents per dollar and the book returned -1.5. Those numbers sit on opposite sides of zero, so there is no ratio to quote: the honest statement is that we predicted a profit and the record lost money, by 23.6 pts.
What this does not say. It does not show that selecting makes calibration worse. Proving that would need the same measurement over every probability we have ever published, and while Kanon does keep a permanent daily record of its own forecasts, that record carries no outcomes yet and nothing has scored it. Until it does, this is a statement about the contracts we picked and nothing wider.
A probability can be well calibrated and still be late. Closing-line value compares the price when a contract was first read against the last clean price before the outcome could be known — the cheapest honest instrument there is, because it needs weeks of contracts rather than years.
| Closing-line value | Value | Contracts |
|---|---|---|
| Net of half the spread, per contract | +0.0 pts | 83 |
| Share that beat the closing price | 42.2% | 83 of 462 contracts · 18.0% had a clean close |
Only contracts with a price captured strictly before the outcome could be known are in this table, and the coverage figure is the honest denominator. It is low. It can only rise going forward, never retroactively: the closing price of a market that has already settled is gone, because neither venue serves a closed book. A low coverage number here is a limit on what we can prove, not a result.
Every contract is logged at first sighting, before its outcome is known, and graded against the venue’s own resolution. Nothing is back-tested and nothing is re-scored later. The record starts at 2026-07-03 because the engine changed materially then; the start date is shown on every figure and no claim is made about anything before it. One position per underlying, so two legs of the same match cannot count twice. Contracts whose entry price came from a stuck feed are excluded and the terminal states what the numbers were before that exclusion, because quietly improving a published figure is indistinguishable from inventing one.