●builderYou can implement this lightweight guidance predictor to improve transcription accuracy in noisy, multi-speaker environments without retraining the base model.
●researcherThe use of asymmetric conditioning branches provides a novel way to calibrate speaker-conditioned inference-time decoding.