Influence-Guided Response Rewriting for Training Data Attribution
September 1, 2026
Instead of simple reweighting, this method uses influence functions to identify key training examples and then rewrites their responses to align with desired behaviors. Experiments across four open-weight LLMs show this improves the intervention effectiveness of data attribution.
HOW THIS AFFECTS YOU
●
researcherYou can move beyond infinitesimal reweighting to achieve more significant behavioral shifts during model fine-tuning.