Asymptotic Closeness of Alignment Methods for Markovian Language Model Outputs
October 2, 2026
This work extends the proof of asymptotic closeness between KL-constrained RL and best-of-n alignment to Markovian token sequences. It addresses the limitations of previous proofs that relied on unrealistic i.i.d. assumptions for language model outputs.
HOW THIS AFFECTS YOU
●
researcherYou can apply these findings to better understand the mathematical relationship between different alignment strategies in non-i.i.d. settings.