A new benchmark developed with equity analysts evaluates agents on their ability to replicate expert human judgment regarding stock-moving information. Top-performing agents matched expert labels in only 52.4% of cases, with false-positive rates varying significantly between models.
HOW THIS AFFECTS YOU
●
builderYou can use this benchmark to evaluate the specialized reasoning capabilities of agents in financial contexts.
●
investorThis highlights the current limitations of AI agents in replicating high-stakes professional financial analysis.