BiasReducer is a lightweight framework that edits the linear reward head of a model using sparse autoencoders to mitigate biases like length or confidence. It allows for dataset-specific bias correction without requiring full model retraining or pre-specified target biases.
HOW THIS AFFECTS YOU
●
builderYou can adapt reward models to your specific dataset's biases with minimal compute.
●
policyThis provides tools for more granular control over model preference alignment.