LUNAR Benchmark Evaluates LLM Personalization via Universal User Behavior Logs
August 7, 2026
LUNAR introduces a benchmark for cross-domain personalization using longitudinal app interaction histories across clothing, food, and mobility. It uses a coarse-to-fine synthesis pipeline to model real-world behavioral distributions for evaluation.
HOW THIS AFFECTS YOU
●
researcherYou can use this benchmark to test how well models ground responses in heterogeneous daily-life activities.
●
designerThis provides a way to measure how well AI interfaces adapt to specific user lifestyles and habits.