Perturbation Method for Efficient Representation Learning Analysis
September 11, 2026
Perturbation identifies linguistic representations by fine-tuning a language model on a single adversarial example and measuring the resulting influence on other inputs. This method avoids geometric assumptions and more accurately reveals structured transfer across linguistic scales than linear alignment approaches.
HOW THIS AFFECTS YOU
●
researcherYou can use this to probe model internals without the pitfalls of unconstrained or strictly linear alignment methods.