This survey analyzes mechanisms for reducing computation and memory costs in Video Large Language Models. It categorizes efficiency gains across frame sampling, modality encoding, token reduction, and LLM prefilling to address the high costs of long-context video understanding.
HOW THIS AFFECTS YOU
●
builderUse these categorized techniques to optimize the latency and cost of your video-based AI products.
●
researcherThis provides a roadmap of current bottlenecks in audiovisual model deployment.