The Endless Exam uses parameterized mathematical construction problems to provide verifiable, continuous quality scores. Unlike existing benchmarks, it allows for tracking progress beyond published mathematical frontiers without capping scores at 1.
HOW THIS AFFECTS YOU
●
researcherYou can use this to measure model capability on evolving, open-ended mathematical problems.
●
investorThis benchmark provides a more granular way to track the actual intelligence gains of frontier models.