ORCA Benchmark Evaluates LLMs on Data Science Code Translation
September 28, 2026
ORCA introduces a benchmark for Data Science Code Translation (DSCT) consisting of 1,600 grounding-level tasks and 200 full-project translation tasks. It evaluates performance across data querying, manipulation, and deep learning domains to ensure functional equivalence.
HOW THIS AFFECTS YOU
●
builderThis highlights the current limitations of LLMs in maintaining functional equivalence during library migrations.
●
researcherYou can use this benchmark to measure how well models handle complex cross-library code translations.