The FIGS framework introduces a 10-turn conversational evaluation to measure if models prioritize user agreement over truthfulness without penalizing empathetic responses.
HOW THIS AFFECTS YOU
●
researcherYou can move beyond single-turn benchmarks to test how models handle long-term user steering and sycophantic tendencies.