The TriviaRoomQA benchmark evaluates 30 open-weight models (7B to 70B parameters) across 288 topics and six European languages. While models excel in history and mathematics, they demonstrate significant performance degradation on everyday popular culture and niche, culturally grounded knowledge.
HOW THIS AFFECTS YOU
●
researcherYou can use this benchmark to identify specific cultural and 'long-tail' knowledge weaknesses in multilingual models.