TraceCLIP: Training-Free Local Semantic Recovery for CLIP
July 30, 2026
TraceCLIP is a training-free framework that recovers latent patch-level semantic evidence by isolating patch-specific terms within the CLIP CLS attention mechanism. It enables dense vision-language tasks like object localization and segmentation without requiring additional supervision or fine-tuning.
HOW THIS AFFECTS YOU
●
builderYou can achieve dense semantic segmentation and localization using existing CLIP weights without extra training costs.
●
researcherThis provides a method to probe local vision-language correspondence directly from global embeddings.