Projector-Only Training Achieves Competitive 3D MLLM Performance
August 21, 2026
Training only the modality projector instead of the full language model backbone yields comparable 3D classification and captioning performance while doubling training sample throughput. This method prevents language model capability drift and significantly reduces computational overhead during multimodal adaptation.
HOW THIS AFFECTS YOU
●
builderYou can achieve faster multimodal adaptation and lower compute costs by freezing the LLM backbone.
●
researcherThis suggests joint training may introduce unnecessary linguistic drift in multimodal architectures.