LLMs Achieve Human-Level F1 in German Parliamentary Sentiment Analysis
August 28, 2026
Using GPT-5 and gpt-oss-120B, researchers demonstrated that LLMs can annotate 150 years of German political discourse regarding migration with macro-F1 scores comparable to human agreement. The study highlights how systematic model errors can bias downstream inference despite high accuracy.
HOW THIS AFFECTS YOU
●
researcherYou should account for model-specific systematic biases when using LLMs for large-scale historical discourse annotation.