Six synthetic prediction datasets, each built to teach one lesson about testing a trading signal or a forecast. Each comes with the full Kvantix report it produced.
The data is generated, so the true answer is known. You can check every verdict in reports/ against how the data was actually made in generate.py. Nothing here comes from a real provider, model or trade.
| # | File | Rows | Horizon · cost | Result | The truth | Lesson |
|---|---|---|---|---|---|---|
| 01 | 01_model_scores_multi_asset.csv |
45,952 | 4h · 0.06 % | 6/7 | A small real edge | Statistically real, but not tradeable. 45,952 rows are worth only ~544 independent observations: outcomes overlap 16 times, and eight coins move together. The edge is real, and smaller than the trading cost. This is the reference format. |
| 02 | 02_directional_calls.csv |
420 | 1d · 0.20 % | 1/6 | Random entries | A 79 % win rate from a coin flip. Take-profit 1.2 % away and stop-loss 4.5 % away make random entries "win" 79 % of the time. Direction over 24 hours: ρ = −0.04. The format of a Telegram-style call log. |
| 03 | 03_daily_probability_forecast.csv |
730 | 1d · 0.20 % | 2/6 | No edge at all | Pure noise can pass secondary tests. Two of six pass by chance. The main test (ρ = 0.04, below the noise floor of 0.07) does not. The column names are non-standard, so a column mapping is needed. |
| 04 | 04_weekly_demand_forecast.csv |
1,248 | 1w · 0 | 6/6 | A strong real edge | Not only crypto. Twelve stores instead of coins, cost 0. Extra columns are ignored. The report still flags ρ = 0.79 as unusually strong and asks for look-ahead to be ruled out, which is normal for demand and would be suspicious for a market. |
| 05 | 05_common_mistakes_excel_dk.csv |
2,880 | 4h · 0.06 % | 1/7 | No edge, 707 duplicates | Formatting can be read, but mistakes can't be fixed. A Danish Excel export (;, decimal comma, BOM, Danish headers, DD-MM-YYYY) is read as it is. Duplicate rows and 53 missing outcomes are flagged, not repaired. |
| 06 | 06_misaligned_outcome.csv |
12,912 | 4h · 0.06 % | 7/7, INVALID | Look-ahead bug | Seven passes that mean nothing. The "forward" return column is the past 4-hour move, and the signal is built from that same move, so every test passes with ρ = 0.94. Statistics on the numbers alone can't see the bug. The alignment check against the price column can, and it marks the report invalid. |
The worst data errors produce excellent results. That is why the data is checked before any statistic is computed.
- See what a report looks like. Open any PDF in
reports/. Each is the full report, synthetic-data banner included. - Try the free Quick Check on kvantix.tech with files 01, 02, 04, 05 or 06. It runs the same seven tests, with no login, and the file is not kept. File 03 is rejected with missing columns on purpose. That is what a provider sees before renaming columns. File 06 shows 7/7 marked INVALID DATA: since 29 September 2026 the Quick Check compares the outcome with the price column, the same alignment check as in the full report.
- Test your own tools. A good validation method should reject 02, 03 and 05, accept 04, find 01 real but too small to trade on, and refuse 06.
- Format your own data. Start from
DATA-GUIDE.md.
python3 generate.py --out /tmp/check && (cd /tmp/check && sha256sum *.csv)
sha256sum -c SHA256SUMSgenerate.py uses only the standard library and a fixed seed. CI regenerates all six files on every push and checks them byte for byte against SHA256SUMS, together with the reports. Each report also prints the SHA-256 of its data file in the header ("Data fingerprint").
A file can be tested when every row has:
- timestamp: when the prediction was made, in UTC.
- prediction: a number. It can be a score, a probability, an expected return, or +1/−1.
- outcome: what happened over one fixed horizon after the timestamp. This is the market's move, not the provider's profit.
- optionally a group (coin, ticker, store) and a price series. The price series enables the regime split and the alignment check.
The provider also gives:
- the horizon
- the real round-trip cost
- every prediction in the period, with nothing removed and nothing added afterwards.
The full guide is in DATA-GUIDE.md.
- weather-forecast-test: the same method run in public on DMI, MET Norway, OpenWeatherMap and the pilots' TAF.
- lock-your-prediction: seal a prediction before the outcome.
- kvantix-reports: the reports on our own engine. We tested ourselves first, and it failed.
- Data, reports and guide: CC BY 4.0. Credit "Kvantix validation examples".
generate.py: MIT.
© 2026 Kvantix (CVR 46296036), Hjørring, Denmark · validation@kvantix.tech · Statistics, not advice.