DSO: Direct Steering Optimization for Bias Mitigation
Key point
DSO uses RL to optimize steering transformations to reduce bias in VLMs and LLMs.
Details
Generative models can produce biased outputs influenced by demographic attributes when helping users make decisions. For example, in a task of finding a doctor in a room for a visually impaired user, errors can occur such as missing a female doctor.
Existing activation steering has been used for safety control in LLMs, but it was not sufficient for bias mitigation, which requires matching outcomes across groups with equal probability. A method was needed that could reduce bias without degrading performance, while also allowing the strength of mitigation to be adjusted depending on the situation.
The proposed Direct Steering Optimization (DSO) directly finds a linear transformation to apply to activations using reinforcement learning. As a result, it showed a better balance between fairness and performance in both VLMs and LLMs, and allowed users to adjust the trade-off between bias mitigation and model capability to their desired level at inference time.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.