●builderYou can add speech capabilities to existing VLM deployments without expensive retraining or architectural modifications.
●researcherThis provides a way to study multimodal integration without the confounding variable of backbone training changes.