DialogVCS: Robust Natural Language Understanding in Dialogue System Upgrade

← All research

Language Understanding

NAACL 2024 (Main) — Long Paper · PDF · Code · DialogVCS

Keep intent understanding robust as a dialogue system evolves. When a chatbot adds new capabilities, its new intent labels can be narrower or broader than existing ones. A request may then match several intents, even though the accumulated training data gives it only one label. DialogVCS makes this mismatch an explicit benchmark.

Dialogue system updates introduce intent labels that overlap with existing labels through narrower or broader meanings.
Figure 1: New intents can be subsets or supersets of existing ones, creating version conflicts and merge friction during system upgrades. Click to enlarge.

The paper constructs four version-control datasets from ATIS, SNIPS, MultiWOZ and CrossWOZ. It formulates the upgrade problem as multi-label classification with positive but unlabeled intents and evaluates baselines that handle false negative labels, class imbalance and overlapping semantics.

Across these simulated upgrades, methods that account for partial labels substantially improve recognition of all valid intents. For example, the paper’s LS Focal loss baseline reaches 92.48 F1 on CrossWOZ-VCS, compared with 38.35 for its basic classifier, using BERT-base and median scores over five runs. The benchmark isolates semantic overlap in system updates; it does not cover every form of production dialogue drift.