The EASEL benchmark evaluates multimodal agents on fine-grained, closed-loop parameterized visual actions using a reference-guided visual reconstruction task. It includes a 440k-sample curriculum dataset to measure how precisely models map visual evidence to execution parameters.
HOW THIS AFFECTS YOU
●
builderYou can leverage the EASEL-Data curriculum to train agents for high-precision visual tool manipulation.
●
researcherYou can use this to evaluate the precision of multimodal agents in parameterized action spaces.