RelateAnything 53M Model Enables Open-Vocabulary Relation Prediction
September 10, 2026
RelateAnything uses a 53M-parameter architecture to perform real-time relation prediction from arbitrary image regions and text inputs. Unlike standard scene-graph models limited to fixed predicate sets, this method decouples relation heads from specific object labels to support open-vocabulary inference.
HOW THIS AFFECTS YOU
●
builderYou can integrate arbitrary object detection outputs into relation-aware vision pipelines.
●
researcherYou can move beyond fixed-taxonomy scene graphs using this decoupled architecture.