DCO: Achieving 1.80x Speedup via Dynamic Cache Orchestration
September 11, 2026
DCO implements application-aware cache management for multi-core AI accelerators to reduce programming complexity compared to scratchpad memory. By using software-stack dataflow information for dead-block prediction and bypass decisions, the system achieves up to 1.80x speedup in cycle-accurate simulations.
HOW THIS AFFECTS YOU
●
builderThis architecture offers a path to higher LLM throughput without the high development cost of managing asynchronous scratchpad memories.