Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery
Evaluation of 19 multimodal large language models shows that current MLLMs still cannot reliably make justified and evidence-bounded inferences from experimental results, which is an essential capability in real scientific discovery.