Latency-Minimizing Request Scheduling for Edge LLM Inference
September 16, 2026
An online scheduling framework for edge LLM inference jointly minimizes end-to-end latency and balances workloads across heterogeneous distributed servers. It addresses the challenges of multi-stage LLM execution dynamics and the delayed observation of scheduling consequences in agentic AI services.
HOW THIS AFFECTS YOU
●
builderYou can reduce latency in agentic workflows by implementing more sophisticated, state-aware edge scheduling.