Human-in-the-Loop through Chain-of-Thought

← All research

Reasoning & Agents

Preprint 2023 · PDF

Use human attention on the reasoning steps that need repair. The Manual Correction System (MCS) samples several chain-of-thought rationales, ranks questions by Diversity Entropy and selects uncertain cases for human review. A person can modify, add or delete erroneous sub-steps, and the model then produces an answer from the corrected rationale.

MCS samples reasoning paths, filters uncertain questions, receives human corrections and generates the final answer.
Figure 1: MCS combines rationale sampling, diversity-based filtering, human correction and a final model answer. Click to enlarge.

The paper also introduces CAMLOP, a cost–utility model that balances improvements in accuracy and user satisfaction against money and time. It treats the amount of human intervention as a design choice and studies when correcting a short reasoning step is worthwhile.

Experiments use GPT-3 text-davinci-002 on 12 arithmetic, commonsense and symbolic reasoning datasets. With the reported intervention threshold, MCS reaches 61.56% accuracy on GSM8K, compared with 56.48% for chain-of-thought prompting. This is a result for a system with human correction, and the benefit must be assessed together with the measured correction effort and the paper’s cost assumptions.