jh

Home | Projects | Blog | Quotes | Connect

Let's Verify Step by Step

Introduction

Some tasks (e.g. math questions) have an unambiguous correct answer. Generative models, such as GPT, can arrive at the correct answer but often make mistakes along the way and fail. How to make them more reliable?

One way is to train a model which can distinguish between correct and incorrect answers. Then we can generate multiple solutions with the generative, and choose only the ones which this model judges as the most correct.

The focus of Let's Verify Step by Step paper is on training these models.

Setup

There are two models: a generator and a reward model. The generator generates answers to a question, and the reward model judges how likely the answer is to be correct.

During evaluation, the generator is sampled multiple times, and the final answer is selected as the one with the highest score from the reward model. The process of selecting this final answer is referred to as “search” (there is no multi-step search as in Tree of Thoughts paper).

The generator is a base GPT-4 model fine-tuned to produce answers in the desired format and is not trained further.

“Although finetuning the generator with RL is a natural next step, it is intentionally not the focus of this work.”

The reward models are initialized from base GPT-4 and further trained with outcome or process supervision.

Training reward models

In outcome supervision, the reward model is only trained to output the probability of the correct solution. Obtaining training data is straightforward: Generate multiple solutions with the generator, check if the final answer is correct, and use correctness as a label.

On the other hand, the PRM has to judge the correctness of each individual reasoning step. The data are provided by human labellers, who judge each step as positive/neutral/negative, stopping at the first incorrect step1.

During the evaluation, the final score is computed as a product of the correctness probabilities of each step.

To use labellers’ time more effectively, there is a selection process to avoid labelling easy, non-ambiguous steps and solutions. Instead, the solutions which are convincing (look good to the current best RM) and wrong (incorrect answers) solutions are surfaced for labeling. The best RM is periodically re-trained. (This method, in which the model actively chooses the next examples for labelling, is generally known as active learning).

Results

The main result is the PRM outperforming ORM. That doesn’t look surprising, but in the previous work (Ueasato et al., 2022) they achieved similar performance. The authors cite these differences from that work:

“we use a more capable base model, we use significantly more human feedback, and we train and test on the more challenging MATH dataset.”

Additionally, the difference scales with a number of sampled solutions, and the PRM score seems to keep improving while the ORM levels off.

(Majority voting just simply takes the most common answer from the N-generated answers from the generator.)

  1. To keep parity between both supervision regimes. If the solution is correct, both models receive the same information (all steps are correct). If the solution is incorrect, both models receive information that the solution is incorrect, with the only additional information in process supervision being the location of the mistake.

    Additionally, this makes labelling costs similar in both regimes. ↩