AivexaNewsSearch
AI news for builders and product teamsChecked every hour

What's Missing From LLM Chatbots: A Sense of Purpose

Collected Oct 1, 2026

An article published by The Gradient examines what its author, Kenneth Li, describes as a missing element in LLM-based chatbots: a sense of purpose in multi-round dialogue. The piece argues that as benchmarks such as MMLU, HumanEval and MATH become saturated, non-interactive measurement of dialogue systems may be insufficient for a future of human-AI collaboration.

The article defines purposeful dialogue as a multi-round user-chatbot conversation centered on a goal or intention. It reviews how dialogue systems are built — pretraining, the introduction of dialogue formatting (such as Huggingface's tokenizer.apply_chat_template), and RLHF — noting that RLHF data is small relative to pretraining and that RLHF formulates reward maximization as a one-step bandit problem, limiting multi-turn planning.

Citing work by Li and collaborators, the article reports that two system-prompted LM agents were made to chat for extended rounds to stress-test instruction following. Aggregated results on LLaMA2-chat-70B and gpt-3.5-turbo-16k were described as alarming: models became distracted after only about 1.6k tokens in dialogue, despite long-context windows of up to 100k tokens. The authors proposed split-softmax to mitigate this. The article also cites evidence that LLMs can be brittle in following instructions under adversarial conditions, and links lost instruction stability to jailbreaking susceptibility and hallucinations.

Li and collaborators propose Dialogue Action Tokens (DAT), in which the last token embedding of the dialogue history is the state input to a planner (actor) that predicts prefix tokens (actions) to control generation. The planner is trained with the RL algorithm TD3+BC. The article states DAT showed significant improvement over baselines on Sotopia, surpassing GPT-4's social capability scores. It adds that a multi-round red-teaming experiment was conducted, and recommends further research into multi-round dialogue as a potential attack surface.

Read at The Gradient

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

LLM-based chatbots’ capabilities have been advancing every month. These improvements are mostly measured by benchmarks like MMLU, HumanEval, and MATH (e.g. sonnet 3.5, gpt-4o). However, as these measures get more and more saturated, is user experience increasing in proportion to these scores? If we envision a future