sft_v5 lineage (v3base → 4k/17k/31k) + probe campaign. Qwen3-ASR
round-trip scoring; probe verdicts are word-diff + Gemini-adjudicated. Audio is a
bounded subset (all rows appear in tables; failures prioritized for audio).
Metric updated: verdicts are Gemini-adjudicated (romanization/number-format
ASR artifacts removed). Raw CER retained for reference — it is inflated by
romanization artifacts and should not be read as an error rate.
References
run
text / ASR
verdict / raw
audio
failure classes (native)
class/lang
asked → heard
cer
audio
Adversarial battery (probeADV31k_hi, 4 seeds each,
ref nat_hi). Control tokens are swallowed; empty/punct inputs produce phantom
speech; native numerals break.
short_v1 — 1–10-word ladder gates
adjudicated-bad % by word count
median seconds by word count
per-language drilldown — bad% by word count
run
text / ASR
verdict
audio
Verdict
config × lang summary
category × config
same text, same sampling noise, 4 configs
Each row is one text that failed under the control config; the same
seed is shown for every config, so differences are the temperature, not the
noise.