LLM Annotation Variance Driven by Task Design Rather Than Sampling
September 30, 2026
Changing task design for LLM annotation causes label variance for hate speech to increase by a factor of 110.6 compared to sampling variance. While models show high intra-design agreement (median Fleiss' kappa = 0.91), small prompt variations significantly impact prevalence estimates.
HOW THIS AFFECTS YOU
●
builderSmall changes to your prompt engineering can drastically alter the ground truth data used for model training.
●
researcherYou cannot rely on confidence scores to mitigate design-induced bias in large-scale labeling.