●builderYou can achieve higher throughput and lower latency in production inference clusters by implementing measured request routing.
●researcherThis study provides a framework for evaluating routing policies in decoupled prefill and decode GPU environments.