The KBMR framework uses MLLM autoregressive capabilities to map images into a semantic space, improving retrieval for long-tail entities in Knowledge-Based Visual Question Answering.
HOW THIS AFFECTS YOU
●
builderThis method can improve the accuracy of multimodal RAG pipelines involving complex visual entities.