REMORY Uses Residual Soft Memory Tokens for Context Compaction
October 9, 2026
REMORY augments textual summaries with a bounded sequence of soft memory tokens to help frozen LLMs approximate full-context performance. On the SummHay benchmark, it achieves near-full-context scores using only 5.2% of input positions, reducing tool errors in Qwen-27B and GLM-5.3-Flash.
HOW THIS AFFECTS YOU
●
builderYou can significantly reduce context costs while maintaining long-horizon agent reasoning capabilities.
●
researcherThis provides a new way to think about sequence-dimension residual connections in long-context models.