DialogVCS: Robust Natural Language Understanding in Dialogue System Upgrade
NAACL 2024 (Main) — Long Paper · PDF · Code · DialogVCS
Keep intent understanding robust as a dialogue system evolves. When a chatbot adds new capabilities, its new intent labels can be narrower or broader than existing ones. A request may then match several intents, even though the accumulated training data gives it only one label. DialogVCS makes this mismatch an explicit benchmark.
The paper constructs four version-control datasets from ATIS, SNIPS, MultiWOZ and CrossWOZ. It formulates the upgrade problem as multi-label classification with positive but unlabeled intents and evaluates baselines that handle false negative labels, class imbalance and overlapping semantics.
Across these simulated upgrades, methods that account for partial labels substantially improve recognition of all valid intents. For example, the paper’s LS Focal loss baseline reaches 92.48 F1 on CrossWOZ-VCS, compared with 38.35 for its basic classifier, using BERT-base and median scores over five runs. The benchmark isolates semantic overlap in system updates; it does not cover every form of production dialogue drift.
