A research team from HKUST and Tencent introduced Intent-Eval in an arXiv preprint posted this month. The benchmark asks whether an AI system can track which instructions remain active after a user clarifies a task, accepts a change, or rejects one.
The researchers adapted 414 tasks covering tool actions, code, database queries, and math. Each task appeared in matched single-turn and multi-turn forms, producing 3,312 evaluation instances per model. Eight models were included in the main comparison, with ten used in the error analysis.
The first loss appeared before anyone changed the plan. Splitting the same task across several turns reduced mean accuracy by 36.30 percentage points compared with placing the complete task in one message. A neutral clarification produced another 4.43-point loss relative to the uninterrupted multi-turn condition.
Changes created a more specific problem. When a proposed revision was rejected, performance averaged 8.06 points below the uninterrupted multi-turn condition. When the revision was accepted, performance averaged 5.55 points lower. In the authors' error analysis, inactive content appeared in 446 of 748 new errors after rejected proposals and 201 of 649 new errors after accepted revisions.
The paper describes the pattern this way: "conversational content is treated as active requirements even after it has been rejected or replaced."
The chat remembers more than the current plan
This finding concerns active intent. It differs from the familiar complaint that an assistant forgot something earlier in a conversation. Here, the old detail remains available and influences the answer after it should have stopped mattering.
Imagine planning a purchase with four items, briefly considering two, then rejecting that change. The active number is still four. Yet the conversation now contains both numbers and a record of the proposed change. Intent-Eval tests whether the model follows the final decision instead of letting the discarded number leak into its work.
The authors call this "mentioned-as-in-effect confusion." The phrase is useful because it avoids pretending that a model holds a stubborn belief. The system receives a sequence containing active instructions, old instructions, questions, rejected ideas, and its own earlier replies. It can generate an answer shaped by material that is still present in the text even though it is no longer part of the task.
Independent research points in the same general direction. The 2025 preprint LLMs Get Lost In Multi-Turn Conversation analyzed more than 200,000 simulated conversations across 15 models and six generation tasks. It reported an average 39% performance drop when information arrived over multiple turns rather than in one complete prompt.
A separate paper published at ACL 2026 introduced EvolIF, a benchmark that changes constraints and topics as a conversation continues. Its authors reported weaknesses in failure recovery and fine-grained instruction following as conversations became deeper. These studies use different tasks and metrics, so their numbers should not be compared as if they measured one phenomenon. Together, they support a narrower judgment: long, changing conversations add reliability problems that a strong answer to a clean one-shot prompt may not reveal.
The intended rule is clear
OpenAI's December 2025 Model Spec gives a direct statement of the desired behavior: "An instruction is superseded if an instruction in a later message at the same level either contradicts it, overrides it, or otherwise makes it irrelevant."
That is a vendor specification, not an independent performance result. The same document also acknowledges that production models do not yet fully reflect the specification. The contrast matters. A product may be designed to follow the latest applicable instruction while still failing to do so consistently when the conversation contains several versions of the plan.
The Intent-Eval authors also tested a training method that shows the problem is not fixed by a slogan. Their approach improved mean accuracy by 10.81 points over the corresponding base models across four models and four domains. Performance after rejected changes remained close to standard supervised fine-tuning, and the method has not been established as a deployed product feature.
What the evidence cannot establish
Intent-Eval is a version-one preprint and has not been independently replicated. Its tasks use one atomic, unambiguous, verifiable change. Natural conversations may contain several partial revisions, unstated assumptions, and vague decisions. That makes everyday use messier, but it also means the benchmark cannot supply a personal failure rate for your assistant or workflow.
The studies do not show that every long chat will fail, that one provider is always safer, or that a short recap guarantees a correct result. They also do not measure the exact consumer procedure proposed below. It is a practical inference from the evidence, rather than a tested intervention.
Put the current plan where you can inspect it
Before an AI system moves money, changes an account, submits a form, books something, makes a purchase, executes code, or edits important records after several rounds of revision, stop the conversation before execution.
Paste a compact block labeled Current plan. Include the requirements that remain active, the choices you cancelled, and the boundary around the next action. Ask the system to restate the active plan and identify any conflict it sees in the conversation. Read that restatement yourself. If it matches your intent, authorize one bounded step rather than the entire remaining chain.
A recap can still be wrong. Its value is that it turns a hidden state problem into a visible object you can check. The closing question is simple: if this action matters, can you point to one clean statement of the plan that is in force now?
