ObGynLongBench Benchmark Reveals LLM Performance Drop in Longitudinal EHR Analysis
September 9, 2026
ObGynLongBench introduces a dataset of 1,500 clinical decision points derived from 976 real pregnancy EHR histories. Testing 17 LLMs shows significant accuracy degradation when models must extract evidence from longitudinal histories compared to being provided direct evidence.
HOW THIS AFFECTS YOU
●
researcherYou should account for the evidence-extraction gap when evaluating long-context medical models.
●
healthThis highlights the difficulty of deploying LLMs for reliable longitudinal obstetric decision support.