[arXiv]score: 0.12
Streaming Translation and Transcription Through Speech-to-Text Causal Alignment
September 11, 2026
Hikari implements end-to-end simultaneous speech-to-text translation and transcription using a policy-free architecture. By utilizing Decoder Time Dilation and supervised fine-tuning for delay recovery, the model achieves competitive quality-latency trade-offs on en-ja, en-de, and en-ru benchmarks, outperforming models up to 38x its size.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy