The study shows that low-rank recovery in factored models depends on an optimizer's gauge equivariance. While gradient descent and Adam (with shared scalars) are equivariant, coordinate-wise methods like standard Adam and RMSProp are not.
HOW THIS AFFECTS YOU
●
researcherThis clarifies why certain optimizers succeed in finding low-rank solutions while others fail.