[arXiv]score: 0.24
MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios
July 29, 2026
MyoCardBench evaluates LLM performance across the cardiovascular care continuum using 2,263 items from 13 de-identified clinical and examination datasets. The benchmark assesses longitudinal, multimodal workflows and specialist tasks through physician-annotated references. In zero-shot testing, GPT-5.4 achieved the highest macro-average score of 62.55.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy