LLM vs Rule-Based Annotation Reliability in Turkish Narrative Datasets
September 15, 2026
Comparative studies evaluate inter-rater reliability between rule-based detectors and LLMs like Gemini 2.5 Flash, Grok, Claude Fable 5, and ChatGPT 5.5 on the Objective Projection Turkish corpus. Results focus on how well machine-generated labels for craft features align with human annotations.
HOW THIS AFFECTS YOU
●
researcherYou can use these findings to assess the validity of using LLM-based labels for narrative feature annotation.