●builderYou can optimize inference latency and accuracy by applying different quantization strategies to prefill and decode stages.
●researcherThis provides a method to break the trade-off between memory traffic and arithmetic precision in LLM inference.