[HN]score: 0.23
Big Pickle on SWE Atlas – Codebase QnA
August 15, 2026
The big-pickle model achieved a 50.8% task resolve rate on the Scale AI SWE Atlas Codebase QnA benchmark using the mini-swe-agent scaffold. Within the same scaffold class, it outperformed GLM 5.2 and GPT-5.6-Sol, trailing only Claude-based models running on the native Claude Code scaffold.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy