●builderYou can use this benchmark to evaluate how your agent's tool access affects financial reasoning accuracy.
●researcherThis offers a more rigorous evaluation framework for long-form, open-ended financial queries than existing extraction benchmarks.