●builderYou can reduce inference latency and memory overhead in small LLMs by quantizing output heads without losing prediction accuracy.
●researcherThe method provides a way to handle nonlinear logit paths during quantization using rank-one corrections.