Dialogue & Conversational Systems
Benchmarking Dialogue Management and Planning Beyond Predefined Action Spaces and Static Plans
Generated automatically from the limitations stated in 6 papers (EMNLP, ACL), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Existing dialogue management and conversational planning frameworks have predominantly been evaluated within restricted operational scopes, such as predefined sets of dialogue acts, single-turn updates, static initial plans, and two-party interactions. Consequently, it remains unknown how robustly current dialogue planning architectures generalise when moving to open-ended token generation, mid-dialogue plan revision, and multi-party settings without pre-annotated action schemas. Without systematic empirical evaluation across these untested conditions, dialogue system designers cannot determine whether structured planning methods degrade or transfer to unconstrained, dynamic conversations.
Why it matters
Establishes clear performance baselines and failure modes for conversational planning methods under unconstrained action spaces and dynamic conversation trajectories, identifying which planning components survive generalization to real-world dialogue conditions.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Unified Robustness Evaluation Suite: Construct a benchmark suite aggregating diverse dialogue settings (open-ended goal-oriented dialogues, dynamic multi-intent shifts requiring mid-dialogue replanning, and multi-party conversations) and measure the task completion rate and plan consistency of baseline planning models across both discrete-action and open-generation variants.
- 2
Dynamic Replanning Stress-Testing: Implement test harnesses where user intents diverge mid-conversation from initial plans, evaluating whether state-of-the-art conversational agents update trajectory plans dynamically versus failing to adapt, measured via dialogue goal accuracy and turn-level coherence.
- 3
Open-Ended Action Generation Assessment: Evaluate modular dialogue agents by replacing fixed dialogue act classifiers with token-level action generators, measuring the pragmatic validity, repetition rate, and conversational efficiency compared to discrete-action baselines.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
Modern large language models equipped with basic zero-shot prompting might already handle mid-dialogue replanning and open-ended generation without structured planning mechanisms, rendering the evaluation of specialized dialogue managers obsolete.
Evidence
Each paper's own statement of the limitation, verbatim.
- YIELD: A Large-Scale Dataset and Evaluation Framework for Information Elicitation AgentsACL 2026
The approach models single-turn generation and lacks conversation management mechanisms needed to sustain structured, long-horizon dialogue plans.
- Value of Information: A Framework for Human–Agent CommunicationACL 2026
Actions and clarifying questions are drawn from predefined sets; the framework does not generate open-ended questions or actions, so applicability to free-form dialogue is untested
- Inter-sentence Context Modeling and Structure-aware Representation Enhancement for Conversational Sentiment Quadruple ExtractionEMNLP 2025
Multi-party dialogue scenarios with more complex interaction patterns are not adequately addressed
- Collaborative Rational Speech Act: Pragmatic Reasoning for Multi-Turn DialogEMNLP 2025
Requires a fixed, predefined set of possible utterances at each turn, restricting use to non-open-ended dialogs; scaling to token-level generation is left to future work and may degrade pragmatic reasoning
- Thoughts to Target: Enhance Planning for Target-driven ConversationEMNLP 2024
Plans are only generated once at conversation start; no mid-dialogue plan revision, limiting real-time adaptability.
- Reinforced Target-driven Conversational PromotionEMNLP 2023
Relies on a pre-defined set of dialogue actions (goal/topic tuples), which restricts applicability to real-world scenarios without such annotation schemes
Nearest existing work
- Deep Reinforcement Learning for Dialogue GenerationEMNLP 2016
- FASTDial: Abstracting Dialogue Policies for Fast Development of Task Oriented AgentsACL 2019
- Dialogue is the Plan: From Interface to Joint Action in Agentic AIACL 2026
- Offline Reinforcement Learning for Mixture-of-Expert Dialogue ManagementNeurIPS 2023
- ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue AgentsACL 2026
- Controllable Mixed-Initiative Dialogue Generation through PromptingACL 2023
- CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level InteractionsACL 2025
- Modeling Multi-Action Policy for Task-Oriented DialoguesEMNLP 2019
- Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy PlanningEMNLP 2023
- ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue AgentsACL 2025
- OpenDial: A Toolkit for Developing Spoken Dialogue Systems with Probabilistic RulesACL 2016
- Open Grounded Planning: Challenges and Benchmark ConstructionACL 2024
- Unsupervised Extraction of Dialogue Policies from ConversationsEMNLP 2024
- MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World ConversationACL 2025
- Deep Learning for Dialogue SystemsACL 2017