●builderCurrent results suggest that increasing search agent complexity or RAG over math papers may not be the primary lever for improving theorem proving.
●researcherThis offers a secure, open-source method to evaluate formal mathematical reasoning in language models.