Open Problems

Dialogue & Conversational Systems

Benchmarking Dialogue Management and Planning Beyond Predefined Action Spaces and Static Plans

Scope to testOpen
Possible candidate · 2/5 runs6 papers report this67% from 2025+

Generated automatically from the limitations stated in 6 papers (EMNLP, ACL), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Existing dialogue management and conversational planning frameworks have predominantly been evaluated within restricted operational scopes, such as predefined sets of dialogue acts, single-turn updates, static initial plans, and two-party interactions. Consequently, it remains unknown how robustly current dialogue planning architectures generalise when moving to open-ended token generation, mid-dialogue plan revision, and multi-party settings without pre-annotated action schemas. Without systematic empirical evaluation across these untested conditions, dialogue system designers cannot determine whether structured planning methods degrade or transfer to unconstrained, dynamic conversations.

Why it matters

Establishes clear performance baselines and failure modes for conversational planning methods under unconstrained action spaces and dynamic conversation trajectories, identifying which planning components survive generalization to real-world dialogue conditions.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Unified Robustness Evaluation Suite: Construct a benchmark suite aggregating diverse dialogue settings (open-ended goal-oriented dialogues, dynamic multi-intent shifts requiring mid-dialogue replanning, and multi-party conversations) and measure the task completion rate and plan consistency of baseline planning models across both discrete-action and open-generation variants.

  2. 2

    Dynamic Replanning Stress-Testing: Implement test harnesses where user intents diverge mid-conversation from initial plans, evaluating whether state-of-the-art conversational agents update trajectory plans dynamically versus failing to adapt, measured via dialogue goal accuracy and turn-level coherence.

  3. 3

    Open-Ended Action Generation Assessment: Evaluate modular dialogue agents by replacing fixed dialogue act classifiers with token-level action generators, measuring the pragmatic validity, repetition rate, and conversational efficiency compared to discrete-action baselines.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

Modern large language models equipped with basic zero-shot prompting might already handle mid-dialogue replanning and open-ended generation without structured planning mechanisms, rendering the evaluation of specialized dialogue managers obsolete.

Evidence

Each paper's own statement of the limitation, verbatim.

Nearest existing work

Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.