Verifiable Reward

Creator
Creator
Seonglae ChoSeonglae Cho
Created
Created
2025 Mar 21 11:54
Editor
Edited
Edited
2026 Feb 18 14:28

RLVR

Programmatic Reward, Verifiable Feedback

Externally provided feedback that can be objectively validated against specific criteria.
Humans do not consider all paths in parallel, nor do they give bonus points to every step of the process just because the answer was correct. They should regret intermediate steps that were strange and give extra points to those that were helpful. In other words, a reflection process that synthesizes multiple trials is also necessary.
  • ORM (Outcome Reward Model): A method that gives rewards based only on the final output.

For Reasoning Data with
GRPO

Rule-based Verifiable Reward. LLM itself is policy network.
  • Accuracy rewards: Checking if the model’s final answer is correct
  • Format rewards: Incentivizing a structured chain-of-thought (enclose CoT tokens by <think> and </think>)
Verifiable Rewards
 
 
 
 

Deepseek R1

arxiv.org
From Zero to Reasoning Hero: How DeepSeek-R1 Leverages Reinforcement Learning to Master Complex Reasoning
A Blog post by Yihua Zhang on Hugging Face
From Zero to Reasoning Hero: How DeepSeek-R1 Leverages Reinforcement Learning to Master Complex Reasoning
 
 
 

Recommendations