●builderWhen building data analysis tools, you must account for the high failure rate of LLMs in maintaining analysis consistency across table perturbations.
●researcherUse this benchmark to evaluate the robustness of LLMs specifically in data-driven reasoning tasks.