ModelsArchitectures & capability
Microsoft research finds LLMs lose reliability over multi-turn chats and long document-editing workflows
Microsoft Research's Jennifer Neville argues standard benchmarks miss how LLMs fail in real work. Her team's studies find sharp drops when tasks unfold over many turns and silent document corruption across long delegated editing workflows, favoring human oversight over full handoff.

On the DELEGATE-52 benchmark, 19 LLMs made chained edits across 52 professional domains. Frontier models, including Gemini 3.1 Pro, Claude 4.6 Opus and GPT 5.4, corrupted about a quarter of document content on average by the end of long workflows. The errors were sparse but severe, and they built up over time. They got worse with larger documents, longer interactions and distractor files. Basic agent tool use did not help, and press coverage reported that it lowered scores. [3] [6] [7]
When fully specified single-turn tasks were split into pieces and revealed over several turns, performance fell by 39% on average across more than 200,000 simulated conversations. The authors attribute most of the loss to unreliability, such as early assumptions and premature answers, rather than to lost ability. A separate team reproduced a similar drop but blames a mismatch between user intent and model interpretation. That team reports that rewriting ambiguous requests into explicit instructions recovers about 20% of performance. [4] [5]
Neville's practical guidance is to treat current systems as collaborators under oversight, not as tools that can be fully delegated to. Users should verify outputs and retry or rephrase when an answer looks wrong. If a conversation goes off track, they should restart with one consolidated prompt. When something fails, describing exactly what went wrong is more useful than a thumbs-down. She notes that failures surprise users because tasks that are easy for humans, or that a model got right once, may not succeed reliably. [1] [2]
Microsoft-authored preprints show frontier models losing reliability over long interactions, with short tests failing to predict later damage.
Read the full assessment
Implication: pilots of chat or agent workflows should run long enough to expose compounding errors and keep human verification in place.
Executive brief
The most striking evidence behind this Microsoft Research podcast episode comes from a paper by Jennifer Neville's team. On their DELEGATE-52 benchmark, even frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) corrupted an average of 25% of document content by the end of long editing workflows (arXiv). In the episode, Neville, a partner research manager at Microsoft Research and a Purdue professor, argues that standard benchmarks are too simple. She says the useful failures show up in multi-turn, long-horizon and collaborative use (video, 10:58). Her practical advice: check outputs, ask again or rephrase, and keep humans overseeing the work rather than handing it off completely (video, 25:27).
What changed and event timeline
Multi-turn degradation paper
Laban, Hayashi, Zhou and Neville report an average 39% performance drop when single-turn tasks are spread across multiple turns, based on more than 200,000 simulated conversations ().
Competing explanation published
Liu et al. argue the drop comes from a gap between what users mean and how models read vague requests, rather than from a lack of capability. They propose a "mediator" step that rewrites ambiguous input into explicit instructions ().
DELEGATE-52 released
Laban, Schnabel and Neville test 19 LLMs on chained document edits across 52 professional domains. Microsoft released the code (;).
Press coverage
The Register reports that only one domain, Python, counted as "ready," and that adding agent tools lowered scores by an average of 6% ().
Podcast published
Neville and Chad Atalla discuss both papers, why some failures are "surprising," and what users should do ().
Capabilities and access
Read the full section
- Models in DELEGATE-52: 19, including Gemini 3.1 Pro, Claude 4.6 Opus and GPT 5.4 (arXiv).
- Code: the benchmark code is public on GitHub.
- Multi-turn study: covered "top open- and closed-weight LLMs" on six generation tasks (arXiv).
- New products: the episode announces no model or product. Neville says the team is working on reinforcement learning methods to fix multi-turn behavior (video, 14:36).
Technical analysis for researchers and developers
- Multi-turn method ("sharding"): fully specified prompts from public single-turn benchmarks are split into pieces.
- What goes wrong: most of the loss comes from a sharp rise in unreliability, with only a small loss in raw ability.
- DELEGATE-52 method: each "round trip" applies an edit and then the inverse edit.
Read the full section
- Multi-turn method ("sharding"): fully specified prompts from public single-turn benchmarks are split into pieces. A simulated user reveals the pieces one turn at a time (video, 13:32).
- What goes wrong: most of the loss comes from a sharp rise in unreliability, with only a small loss in raw ability. Models make assumptions early, attempt a final answer too soon, and then fail to recover (arXiv).
- DELEGATE-52 method: each "round trip" applies an edit and then the inverse edit. Many round trips are chained, and the recovered document is compared with the original (GitHub).
- Results: errors are sparse but severe, and they compound. Damage gets worse with larger documents, longer interactions and distractor files (arXiv). Scores after two interactions did not reliably predict scores after twenty (GitHub).
Claims and evidence
- Average 39% multi-turn drop
- Frontier models corrupt ~25% of content
- Restarting with one fully specified prompt helps
Read the full section
| Claim | Status |
| Average 39% multi-turn drop | Reported by the authors in a Microsoft/Salesforce preprint (arXiv). A separate team reproduced a similar ~60% relative drop but disputes the cause (arXiv 2602.07338). |
| Frontier models corrupt ~25% of content | Microsoft-authored preprint (arXiv). The Register repeats it but did not test it (The Register). The reviewed sources contain no independent replication. |
| Restarting with one fully specified prompt helps | Neville's practical advice, citing the paper (video, 14:48). |
| Consumer logs are analyzed only in privacy-preserving ways and are not used for fine-tuning | Microsoft's own statement, not independently checked (video, 18:30). |
Context and prior work
- Where her agenda comes from: Neville's work began in statistical relational learning (video, 4:36).
- Transformer limits: she says the open question is how much of the transformer's known limits can be handled by agent wrappers and reasoning, and how much needs a new architecture (video, 28:42).
- Forecasting caution: she recalls that experts in 2000 expected computer Go to take 50 to 100 years to solve (video, 30:35).
Read the full section
- Where her agenda comes from: Neville's work began in statistical relational learning (video, 4:36). She now leads Microsoft's AI Interaction and Learning team, which uses large-scale analysis of consumer logs to find common failure patterns (video, 16:35).
- Transformer limits: she says the open question is how much of the transformer's known limits can be handled by agent wrappers and reasoning, and how much needs a new architecture (video, 28:42).
- Forecasting caution: she recalls that experts in 2000 expected computer Go to take 50 to 100 years to solve (video, 30:35).
Limitations, safety and contested findings
- The cause of the multi-turn drop is disputed. Liu et al. attribute it to intent misalignment, not unreliability.
- Both benchmarks are artificial. Both rely on simulation: simulated users in one, synthetic edit-and-undo round trips in the other (GitHub).
- Why failures surprise people: Neville says these errors are subtle losses of meaning, not obvious mistakes like hallucinations.
Read the full section
- The cause of the multi-turn drop is disputed. Liu et al. attribute it to intent misalignment, not unreliability. Their mediator step recovers about 20% of performance, while memory retrieval alone adds only about 3% (arXiv 2602.07338).
- Both benchmarks are artificial. Both rely on simulation: simulated users in one, synthetic edit-and-undo round trips in the other (GitHub).
- Why failures surprise people: Neville says these errors are subtle losses of meaning, not obvious mistakes like hallucinations. Users also wrongly assume that tasks easy for humans, or tasks a model got right once, will keep working (video, 23:09).
- Progress is real but partial. The Register reports that GPT-family scores rose from 14.7% to 71.5% over 16 months (The Register).
Business and practitioner implications
- Use models within workflows that include oversight and verification (video, 26:17). Short pilots can hide failures that only appear after many interactions (GitHub).
- Don't assume agent tooling fixes this. Basic tool use did not help on DELEGATE-52 (arXiv).
- Recover from confused sessions by restarting. Start a new session with a consolidated prompt (video, 14:48).
Read the full section
- Don't delegate fully. Use models within workflows that include oversight and verification (video, 26:17). Short pilots can hide failures that only appear after many interactions (GitHub).
- Don't assume agent tooling fixes this. Basic tool use did not help on DELEGATE-52 (arXiv).
- Recover from confused sessions by restarting. Start a new session with a consolidated prompt (video, 14:48).
- Give detailed feedback. Neville says describing exactly what went wrong is more useful for improving models than a thumbs-down (video, 27:08).
- Look at the data. When a metric stalls, inspect the data before concluding the task is too hard (video, 35:37).
Sources
Read the full section
- What AI gets wrong and what failure teaches us (Microsoft Research Podcast)
- Podcast video (YouTube)
- LLMs Get Lost In Multi-Turn Conversation (arXiv 2505.06120)
- LLMs Corrupt Your Documents When You Delegate (arXiv 2604.15597)
- microsoft/DELEGATE52 (GitHub)
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation (arXiv 2602.07338)
- The Register: Microsoft researchers find AI models and agents can't handle long-running tasks
The source trail.
Sources (7)
What AI gets wrong and what failure teaches us
Article text retrieved; extracted text may omit tables or interactive elements. A transcript of the video linked from this page is supplied (linked video manual captions; https://www.youtube.com/watch?v=Zs-SPS8OQb8); automatic text may contain errors.
microsoft.comLinked video: What AI gets wrong and what failure teaches us
linked video manual captions
www.youtube.com