DPO

Direct Preference Optimization

Literally upweight the preferred response while downweight the unpreferred response, which is a really simple mechanism.

It can learn human preferences without RL, using significantly less memory compared to the

PPO architecture in RLHF.

The model directly incorporates user's pairwise preferences into training using a set of preference pairs and logits.


{'prompt': '<|im_start|>system\nYou are an AI assistant. You will be given a task. You must generate a detailed and long answer.<|im_end|>\n<|im_start|>user\nGenerate an approximately fifteen-word sentence that describes all this data: Midsummer House eatType restaurant; Midsummer House food Chinese; Midsummer House priceRange moderate; Midsummer House customer rating 3 out of 5; Midsummer House near All Bar One<|im_end|>\n<|im_start|>assistant\n',
'chosen': 'Midsummer House is a moderately priced Chinese restaurant with a 3/5 customer rating, located near All Bar One.<|im_end|>\n',
'rejected': ' Sure! Here\'s a sentence that describes all the data you provided:\n\n"Midsummer House is a moderately priced Chinese restaurant with a customer rating of 3 out of 5, located near All Bar One, offering a variety of delicious dishes."<|im_end|>\n'}

Toxicity reduction interpretation

DPO reduces toxicity not through a few neurons, but via distributed activation shifts across all MLP neurons. DPO operates through balanced action of four neuron groups:

TP↓: Toxicity-aligned + positive activation → decrease

TN↓: Toxicity-aligned + negative activation → decrease

AP↓: Anti-toxicity aligned + positive activation → increase (anti-toxicity reinforcement)

AN↓: Anti-toxicity aligned + negative activation → increase (anti-toxicity reinforcement)

Patching activations of all four groups to post-DPO values reproduces or exceeds DPO effects. In contrast, patching only toxic neurons has minimal effect.

arxiv.org

https://arxiv.org/pdf/2411.06424

suppresses a few toxic neurons

arxiv.org

https://arxiv.org/pdf/2401.01967

DPO

Direct Preference Optimization

Implementation

Self-Rewarding Language Models

DPO Datasets

sDPO from upstage

Toxicity reduction interpretation

Recommendations