UkisAI released the Swift 1.5 27B model family, utilizing GSPO and OPD to reduce thinking tokens by 58.5% while increasing accuracy. The 27B model outperforms the base model on Terminal Bench 2.1 by avoiding overthinking error loops and offering improved agentic performance.
HOW THIS AFFECTS YOU
●
builderYou can achieve significantly lower latency and token costs for reasoning tasks using these optimized weights.
●
researcherWorth watching because it demonstrates effective penalization of pathological overthinking patterns in RL-tuned models.