Inconsistency in LLM-Derived Cardinal Preference Judgments
August 19, 2026
Experiments across six LLMs demonstrate that numerical preference judgments, such as willingness-to-pay, lack self-consistency. These judgments often fail to align with a single underlying utility function, complicating the estimation of human preferences for agentic decision-making.
HOW THIS AFFECTS YOU
●
builderYou cannot reliably use raw LLM numerical outputs to build stable utility-based decision engines.
●
researcherThis provides a statistical framework to measure how much LLM responses depart from consistent utility models.