same utterance→serial numberSame checkpoint ≠ same executable system
SURE-EVAL declares every metric as a versioned pipeline of nodes. Same route, same inputs, same locked environment — the same score, with every step on the record.
An evaluation is not a black-box call but a fully declared pipeline: describe the route first, then run it into a report. Scoring nodes wrap the community standards — sclite, sacrebleu, meeteval, WeNet — so the numbers reconcile with the toolchain you already trust. Take Chinese ASR character error rate (CER) —
Each colored ribbon is an auditable evaluation declaration: a task passes through frontend, transcription, normalization and scoring nodes before landing in a report. Click a route to open its Route JSON; enter from a task card below and the atlas keeps only that task’s real paths, with no unrelated shadows.
Choose what you evaluate, then inspect the nodes it actually runs, the contracts it pins and the JSON it produces. The catalog spans recognition, translation, synthesis, conversion, enhancement, target-speaker extraction and speaker analysis.
SURE Harness plugs this framework into a Pi-style terminal agent: you state intent through slash commands, the agent plans and executes, and artifact gates keep every step reproducible and auditable — four commands from model discovery to final report, with no silent execution fallbacks.
Same checkpoint is not enough. Same predictions are not enough.
same utterance→serial numberSame checkpoint ≠ same executable system
Model B ranks first
A scoring contract can reverse the apparent ranking