The SEE benchmark evaluates multimodal models on chemistry, biology, and materials science, revealing that top models achieve only 48.7% accuracy. Findings show general-purpose models often outperform specialized ones, and tool use provides only marginal accuracy gains to 52.7%.
HOW THIS AFFECTS YOU
●
researcherYou can use this benchmark to evaluate reasoning reliability in scientific multimodal tasks.
●
healthThis highlights the current limitations of using general-purpose MLLMs for autonomous laboratory science.