●builderYou must strip historical metadata from evaluation prompts to ensure independent model scoring.
●researcherThis highlights a critical flaw in how LLMs perform comparative or iterative evaluation tasks.
●policyReliability concerns in automated content moderation and auditing arise from these scoring dependencies.