72.67 Rekordbox 7measured by us, same 566 tracks
held-out, shipping appThe boardmeasured 2026-07-16
The numbers, and the means to check them.
Mixterio is the only DJ product that publishes reproducible, held-out analysis accuracy. Every row below states the dataset, the split, and the comparison figures: published ones where they exist, and our own same-dataset measurements, labelled as ours, where a tool publishes none. Where a stronger number exists, it is on the chart next to ours.
89.1 Beat This!Foscarin, ISMIR 2024
held-out, never trained on78.3 Beat This!Foscarin, ISMIR 2024
held-out, never trained onno published comparisonnobody publishes a figure for this metric, so we draw none
raw grid, general material0.472 Foote (2000)measured by us, same corpus
held-out, coarse HR3FEvery figure, with its dataset and its split.
Held-out numbers first. Where we are level or behind rather than ahead, the row says so.
Key detection
MIREX weighted - GiantSteps, 566 tracksabove the published figures74.80 held-out (73.78 full set, 66.08 exact), measured through the shipping analysis path at the 44.1 kHz the application decodes at, and agreeing with the reference 22.05 kHz call on 566 of 566 tracks. On tracks it was never tuned on, the shipping application scores above the strongest published academic figures. Rekordbox and Serato publish no number for this benchmark, so we ran both ourselves on the same 566 tracks with key tags stripped and kept the per-track results; our directly comparable full-set score is 73.78 and also leads them. Standalone key-tagging utilities are a different category of product and are not on this chart.
results files: key_product_conditions_full566_2026-09-10.json - key_product_conditions.py - key_giantsteps_shipped_2026-07-16.json - competitor_key_rekordbox_2026-07-16.json - competitor_key_serato_2026-07-16.json - competitor_key_bench.py
Beat tracking
F-measure - GTZAN, 992 trackslevel with the published figuresWe reproduce the published state of the art on a set our tracker never saw. Strong where DJs live (hip-hop 97.5), weakest on classical (65.9), which is the hard case for every system. We deliberately do not report our Ballroom score, because Ballroom sits in the reference model's training data and would measure memorisation.
results files: beats_gtzan_beatthis_2026-07-16.json - beats_extract_gtzan.py - beats_eval_gtzan.py
Downbeat tracking
F-measure - GTZAN, 992 trackslevel with the published figuresDownbeats are the ones that matter for phrasing, and they are harder than beats. On top of the model we add a figure no benchmark asks for: how tightly the grid sits on the actual kick, a median of 8.5 ms across our library.
results files: beats_gtzan_beatthis_2026-07-16.json - beats_eval_gtzan.py
Tempo
ACC1, strict - GTZAN, 992 trackslevel with the published figuresRight 81.85 percent of the time strictly, 93.48 percent if you allow double or half time. The 11.63 percent gap between those two is the octave problem, which is open for every published system, not just ours. Nobody publishes a comparable figure for this metric, so we draw no comparison.
results files: tempo_gtzan_2026-07-16.json - tempo_eval_gtzan.py
Structure
boundary HR3F - SALAMI IA, 153 held-out of a 361-track corpusbehind the published figuresThis was our weakest pillar at 0.221, published rather than hidden. Measured at the operating point and decode rate the application runs: coarse HR.5F 0.1966 / HR3F 0.4025 held-out (n=153, 9.49 boundaries per track). We are level with our own Foote reproduction at the strict tolerance and 0.070 behind it at the lenient one. The denser research operating point scores 0.2341 / 0.4823 on the same tracks and is an ablation, not the product. We do not claim the higher published figures beaten: those come from different, easier slices of the benchmark, so we measured the field's own algorithms on our corpus instead of comparing rulers. Human annotators only agree with each other near 0.6 to 0.7 on this task, which is the honest ceiling.
results files: structure_product_conditions_held_2026-09-10.json - structure_product_conditions.py - structure_compare_held_2026-08-04.json - structure_eval.py
Measured 2026-07-16. Bars are drawn to each benchmark's own scale, so no gap is stretched. Every comparison is either a figure its makers published, or our own measurement on the same tracks, labelled as ours and backed by the per-track results.
How a third party checks this
Reproducible without our source.
Checking these numbers means running the app you install over a fixed public dataset and recomputing the table. How the analyzer works stays ours; the claim does not depend on that being open, because anyone with the same audio can run the same app and get the same numbers.
- 01
Pinned dataset: exact source, version, and audio checksum.
- 02
Declared split: what tuned, what was held out. The headline is always the held-out number.
- 03
Written protocol, published alongside the harness.
- 04
Per-track results: one row per track, prediction and score, regenerated whenever the analyzer changes.
- 05
Verification drives the shipped product over public audio, so our source and models can stay closed while the claim stays checkable.
- The headline is always the held-out number, never the tuning half.
- We do not report scores on data our models were trained on, even when they look excellent.
- Comparisons prefer figures their authors published. Where a widely used tool publishes none, we measure it ourselves on the same pinned dataset, say so on the chart, and keep the per-track results.
- Every figure on this site traces back to the benchmark document behind it. If it is not in there, it does not go on the site.
Get the results files when they go public.
The per-track files and the harness publish with the release. What the analyzer measures.