Automated Generation of Ontology-Grounded Reasoning Benchmarks
October 2, 2026
A new pipeline generates multiple-choice benchmarks from OWL 2 ontologies by using formal reasoners to verify correct answers. Distractors are created by perturbing class definition axioms, ensuring that all reasoning tasks are grounded in explicit, verifiable background knowledge.
HOW THIS AFFECTS YOU
●
researcherYou can use this to create high-fidelity, verifiable evaluation datasets for scientific AI without manual labeling.