Open-Vocabulary Sign Recognition via Video Captioning and Retrieval
September 4, 2026
This method replaces closed-set gloss classification with an open-vocabulary system that captions sign articulation and retrieves descriptions via a multilingual encoder. By leveraging vision-language models for procedural descriptions, the system avoids the need for gloss-annotated lexicons and generalizes to unseen signs.
HOW THIS AFFECTS YOU
●
builderYou can build sign language tools that support arbitrary vocabularies without retraining classifiers.
●
designerThis enables more natural, descriptive interactions for users with diverse signing styles.