SSO Enables Skill Optimization from Unlabeled Data
August 3, 2026
Self-Supervised Skill Optimization (SSO) allows LLM agents to learn reusable procedural skills from unlabeled task instances. It uses an LLM judge to compare executions and a separate behavior extractor to aggregate evidence for and against observed behaviors without ground-truth labels.
HOW THIS AFFECTS YOU
●
builderYou can now improve agentic workflows and skill sets using your existing unlabeled interaction data.
●
researcherThis provides a framework for optimizing agent behavior in environments where reward signals are unavailable.