Round 06 — DE-ESSER A/B (pure post-processing on round-05 audio; no model changes).
Diagnosis: shigure's SOURCE has hot sibilants (S-harshness 2.92 vs ~1.5 for others, hiss floor ~9x) — the anchor taught the model hot S. On top, GAN vocoders render S bursts as spectrally messy splatter (texture, not level — invisible to metrics, obvious to ears).
Rows: each winner raw vs deesser-mild (i=0.4) vs deesser-strong (i=0.75). Measured S-harshness on 'short': raw 1.34 / mild 0.47 / strong 0.97.
Listen for: (1) does mild fix the S noise? (2) does strong dull the voice? (3) is the residual overall noise acceptable after de-essing?
If a de-ess setting wins, it becomes part of the BAKE: the master corpus gets de-essed, so the student model learns clean sibilants natively.
voice
title進め、僕の無敵マシン!!!the title line (JP marketing title; the shouted VO line becomes its conlang translation)
shortえっ嘘でしょ。ITA EMOTION100_001 — short exclamation