●builderYou can improve your evaluation pipelines by switching from logprob-based scoring to verbalized confidence for modern proprietary models.
●researcherThe compatibility shift suggests that standard evaluation heuristics for LLMs must evolve alongside model generations.