PetQA Benchmark for Veterinary Knowledge and Clinical Reasoning
September 7, 2026
PetQA introduces a Korean benchmark comprising 10,076 text and 8,751 multimodal QA pairs for evaluating veterinary reasoning. The benchmark tests eighteen models across zero-shot, RAG, and SFT settings to assess clinical reliability.
HOW THIS AFFECTS YOU
●
builderUse this to benchmark RAG and SFT performance in specialized, high-stakes domains.
●
healthThis offers a specific metric for evaluating the clinical safety of models in veterinary medicine.