LAST Framework for Query-Dependent Visual Token Pruning
July 31, 2026
LAST reduces cloud-side MLLM inference costs by using a compact edge-side VLM as a guidance proxy for visual token pruning. This training-free framework allows for query-dependent pruning, determining token relevance before transmitting data from edge to cloud.
HOW THIS AFFECTS YOU
●
builderYou can optimize edge-cloud deployment costs by reducing the volume of visual tokens sent to cloud-based multimodal models.
●
researcherThis offers a training-free approach to query-guided token importance signaling in collaborative inference.