Oct 7 edition/Reporting & analysis
ModelsAgentsResearchBusiness

ModelsArchitectures & capability

Microsoft research finds LLMs lose reliability over multi-turn chats and long document-editing workflows

Microsoft Research's Jennifer Neville argues standard benchmarks miss how LLMs fail in real work. Her team's studies find sharp drops when tasks unfold over many turns and silent document corruption across long delegated editing workflows, favoring human oversight over full handoff.

Illustration from Microsoft Research: Microsoft research finds LLMs lose reliability over multi-turn chats and long document-editing workflows
Image: Microsoft Research — Original article ↗
THE CORE IDEAS3 TAKEAWAYS
01

On the DELEGATE-52 benchmark, 19 LLMs made chained edits across 52 professional domains. Frontier models, including Gemini 3.1 Pro, Claude 4.6 Opus and GPT 5.4, corrupted about a quarter of document content on average by the end of long workflows. The errors were sparse but severe, and they built up over time. They got worse with larger documents, longer interactions and distractor files. Basic agent tool use did not help, and press coverage reported that it lowered scores. [3] [6] [7]

02

When fully specified single-turn tasks were split into pieces and revealed over several turns, performance fell by 39% on average across more than 200,000 simulated conversations. The authors attribute most of the loss to unreliability, such as early assumptions and premature answers, rather than to lost ability. A separate team reproduced a similar drop but blames a mismatch between user intent and model interpretation. That team reports that rewriting ambiguous requests into explicit instructions recovers about 20% of performance. [4] [5]

03

Neville's practical guidance is to treat current systems as collaborators under oversight, not as tools that can be fully delegated to. Users should verify outputs and retry or rephrase when an answer looks wrong. If a conversation goes off track, they should restart with one consolidated prompt. When something fails, describing exactly what went wrong is more useful than a thumbs-down. She notes that failures surprise users because tasks that are easy for humans, or that a model got right once, may not succeed reliably. [1] [2]

WHY IT MATTERS

Microsoft-authored preprints show frontier models losing reliability over long interactions, with short tests failing to predict later damage.

Read the full assessment

Implication: pilots of chat or agent workflows should run long enough to expose compounding errors and keep human verification in place.

Executive brief

The most striking evidence behind this Microsoft Research podcast episode comes from a paper by Jennifer Neville's team. On their DELEGATE-52 benchmark, even frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) corrupted an average of 25% of document content by the end of long editing workflows (arXiv). In the episode, Neville, a partner research manager at Microsoft Research and a Purdue professor, argues that standard benchmarks are too simple. She says the useful failures show up in multi-turn, long-horizon and collaborative use (video, 10:58). Her practical advice: check outputs, ask again or rephrase, and keep humans overseeing the work rather than handing it off completely (video, 25:27).

What changed and event timeline

  1. Multi-turn degradation paper

    Laban, Hayashi, Zhou and Neville report an average 39% performance drop when single-turn tasks are spread across multiple turns, based on more than 200,000 simulated conversations ().

  2. Competing explanation published

    Liu et al. argue the drop comes from a gap between what users mean and how models read vague requests, rather than from a lack of capability. They propose a "mediator" step that rewrites ambiguous input into explicit instructions ().

  3. DELEGATE-52 released

    Laban, Schnabel and Neville test 19 LLMs on chained document edits across 52 professional domains. Microsoft released the code (;).

  4. Press coverage

    The Register reports that only one domain, Python, counted as "ready," and that adding agent tools lowered scores by an average of 6% ().

  5. Podcast published

    Neville and Chad Atalla discuss both papers, why some failures are "surprising," and what users should do ().

Capabilities and access

  • Models in DELEGATE-52: 19, including Gemini 3.1 Pro, Claude 4.6 Opus and GPT 5.4 (arXiv).
  • Code: the benchmark code is public on GitHub.
  • Multi-turn study: covered "top open- and closed-weight LLMs" on six generation tasks (arXiv).
Read the full section
  • Models in DELEGATE-52: 19, including Gemini 3.1 Pro, Claude 4.6 Opus and GPT 5.4 (arXiv).
  • Code: the benchmark code is public on GitHub.
  • Multi-turn study: covered "top open- and closed-weight LLMs" on six generation tasks (arXiv).
  • New products: the episode announces no model or product. Neville says the team is working on reinforcement learning methods to fix multi-turn behavior (video, 14:36).

Technical analysis for researchers and developers

  • Multi-turn method ("sharding"): fully specified prompts from public single-turn benchmarks are split into pieces.
  • What goes wrong: most of the loss comes from a sharp rise in unreliability, with only a small loss in raw ability.
  • DELEGATE-52 method: each "round trip" applies an edit and then the inverse edit.
Read the full section
  • Multi-turn method ("sharding"): fully specified prompts from public single-turn benchmarks are split into pieces. A simulated user reveals the pieces one turn at a time (video, 13:32).
  • What goes wrong: most of the loss comes from a sharp rise in unreliability, with only a small loss in raw ability. Models make assumptions early, attempt a final answer too soon, and then fail to recover (arXiv).
  • DELEGATE-52 method: each "round trip" applies an edit and then the inverse edit. Many round trips are chained, and the recovered document is compared with the original (GitHub).
  • Results: errors are sparse but severe, and they compound. Damage gets worse with larger documents, longer interactions and distractor files (arXiv). Scores after two interactions did not reliably predict scores after twenty (GitHub).

Claims and evidence

  • Average 39% multi-turn drop
  • Frontier models corrupt ~25% of content
  • Restarting with one fully specified prompt helps
Read the full section
ClaimStatus
Average 39% multi-turn dropReported by the authors in a Microsoft/Salesforce preprint (arXiv). A separate team reproduced a similar ~60% relative drop but disputes the cause (arXiv 2602.07338).
Frontier models corrupt ~25% of contentMicrosoft-authored preprint (arXiv). The Register repeats it but did not test it (The Register). The reviewed sources contain no independent replication.
Restarting with one fully specified prompt helpsNeville's practical advice, citing the paper (video, 14:48).
Consumer logs are analyzed only in privacy-preserving ways and are not used for fine-tuningMicrosoft's own statement, not independently checked (video, 18:30).

Context and prior work

  • Where her agenda comes from: Neville's work began in statistical relational learning (video, 4:36).
  • Transformer limits: she says the open question is how much of the transformer's known limits can be handled by agent wrappers and reasoning, and how much needs a new architecture (video, 28:42).
  • Forecasting caution: she recalls that experts in 2000 expected computer Go to take 50 to 100 years to solve (video, 30:35).
Read the full section
  • Where her agenda comes from: Neville's work began in statistical relational learning (video, 4:36). She now leads Microsoft's AI Interaction and Learning team, which uses large-scale analysis of consumer logs to find common failure patterns (video, 16:35).
  • Transformer limits: she says the open question is how much of the transformer's known limits can be handled by agent wrappers and reasoning, and how much needs a new architecture (video, 28:42).
  • Forecasting caution: she recalls that experts in 2000 expected computer Go to take 50 to 100 years to solve (video, 30:35).

Limitations, safety and contested findings

  • The cause of the multi-turn drop is disputed. Liu et al. attribute it to intent misalignment, not unreliability.
  • Both benchmarks are artificial. Both rely on simulation: simulated users in one, synthetic edit-and-undo round trips in the other (GitHub).
  • Why failures surprise people: Neville says these errors are subtle losses of meaning, not obvious mistakes like hallucinations.
Read the full section
  • The cause of the multi-turn drop is disputed. Liu et al. attribute it to intent misalignment, not unreliability. Their mediator step recovers about 20% of performance, while memory retrieval alone adds only about 3% (arXiv 2602.07338).
  • Both benchmarks are artificial. Both rely on simulation: simulated users in one, synthetic edit-and-undo round trips in the other (GitHub).
  • Why failures surprise people: Neville says these errors are subtle losses of meaning, not obvious mistakes like hallucinations. Users also wrongly assume that tasks easy for humans, or tasks a model got right once, will keep working (video, 23:09).
  • Progress is real but partial. The Register reports that GPT-family scores rose from 14.7% to 71.5% over 16 months (The Register).

Business and practitioner implications

  • Use models within workflows that include oversight and verification (video, 26:17). Short pilots can hide failures that only appear after many interactions (GitHub).
  • Don't assume agent tooling fixes this. Basic tool use did not help on DELEGATE-52 (arXiv).
  • Recover from confused sessions by restarting. Start a new session with a consolidated prompt (video, 14:48).
Read the full section
  • Don't delegate fully. Use models within workflows that include oversight and verification (video, 26:17). Short pilots can hide failures that only appear after many interactions (GitHub).
  • Don't assume agent tooling fixes this. Basic tool use did not help on DELEGATE-52 (arXiv).
  • Recover from confused sessions by restarting. Start a new session with a consolidated prompt (video, 14:48).
  • Give detailed feedback. Neville says describing exactly what went wrong is more useful for improving models than a thumbs-down (video, 27:08).
  • Look at the data. When a metric stalls, inspect the data before concluding the task is too hard (video, 35:37).

Sources

Read the full section
FOLLOW THE EVIDENCE

The source trail.

Sources (7)
A LITTLE LESS NOISE. A LOT MORE CONTEXT.

Stay curious.
Follow the evidence.

Independent perspectives, the original sources, and room for the questions that don't have easy answers.

How we build the brief