Maya FireRedTTS2 — eval dashboard

sft_v5 lineage (v3base → 4k/17k/31k) + probe campaign. Qwen3-ASR round-trip scoring; probe verdicts are word-diff + Gemini-adjudicated. Audio is a bounded subset (all rows appear in tables; failures prioritized for audio).

Metric updated: verdicts are Gemini-adjudicated (romanization/number-format ASR artifacts removed). Raw CER retained for reference — it is inflated by romanization artifacts and should not be read as an error rate.

References

runtext / ASRverdict / rawaudio

T1: 12000 rows, 885 audio · T2: 76 exemplars, 76 audio · T3: 22 adv items + 4 mismatch conditions, 22+ audio · T4: 2×2 temp sweep, 28 comparison texts × 4 configs, 112 audio · T5: short_v1 ladder, 3960 clips, 3960 audio