Human-in-the-Loop through Chain-of-Thought
Preprint 2023 · PDF
Use human attention on the reasoning steps that need repair. The Manual Correction System (MCS) samples several chain-of-thought rationales, ranks questions by Diversity Entropy and selects uncertain cases for human review. A person can modify, add or delete erroneous sub-steps, and the model then produces an answer from the corrected rationale.
The paper also introduces CAMLOP, a cost–utility model that balances improvements in accuracy and user satisfaction against money and time. It treats the amount of human intervention as a design choice and studies when correcting a short reasoning step is worthwhile.
Experiments use GPT-3 text-davinci-002 on 12 arithmetic, commonsense and symbolic reasoning datasets. With the reported intervention threshold, MCS reaches 61.56% accuracy on GSM8K, compared with 56.48% for chain-of-thought prompting. This is a result for a system with human correction, and the benefit must be assessed together with the measured correction effort and the paper’s cost assumptions.
