ModelsArchitectures & capability
LLM writing workflow emphasizes critique over ghostwriting to preserve author voice
A practitioner debate around Thomas Ptacek’s “How To Write With An LLM,” amplified by Simon Willison, frames LLMs as copyeditors rather than ghostwriters. The rule: let models identify weaknesses, but keep human authors responsible for wording, verification, and voice.
Ptacek’s workflow separates drafting from critique: humans write the prose, LLMs flag problems, and humans decide how to revise without adopting model-generated phrasing. [1] [10]
Willison reports a similar boundary for his own writing: LLMs can help with proofreading, grammar, repetition, weak arguments, links, and possible factual errors, but not with first-person prose in his voice. [8] [9]
Evidence in the reviewed research shows a practitioner norm forming around LLMs as diagnostic writing tools, while empirical studies suggest model text and suggestions can have stylistic attractors.
Read the full assessment
Those findings do not prove Ptacek’s method improves quality. The implication for organizations is governance-focused: separate critique, generation, verification, and authorship so teams can use AI assistance without blurring accountability, flattening brand voice, or treating model output as checked fact.
Executive brief
On September 17, 2026, Simon Willison amplified Thomas Ptacek’s essay “How To Write With An LLM”, framing it as a practical rule set for using language models as copyeditors, not ghostwriters. Willison endorses the same boundary for his own blog: he says he does not let LLMs write content in his voice, but does use them for proofreading, fact-checking, spelling, grammar, and occasional thesaurus-like help. How To Write With An LLM — A Final Ward OpenAI’s own help guidance still warns that ChatGPT can produce inaccurate information and recommends checking sources directly when accuracy matters.
Read the full section
On September 17, 2026, Simon Willison amplified Thomas Ptacek’s essay “How To Write With An LLM”, framing it as a practical rule set for using language models as copyeditors, not ghostwriters. Ptacek’s central prescription is strict: draft the prose yourself, ask the model to identify problems, but do not adopt the model’s phrasing. Willison endorses the same boundary for his own blog: he says he does not let LLMs write content in his voice, but does use them for proofreading, fact-checking, spelling, grammar, and occasional thesaurus-like help. How To Write With An LLM — A Final Ward
This is not a model launch or benchmark story. It is a practitioner workflow story, but it lands in a live research debate: empirical work has found that LLM text often has recognizable stylistic regularities, can homogenize creative output, and may pull writing toward dominant cultural norms. Those findings do not prove Ptacek’s workflow improves writing quality, but they support the risk model behind it: AI writing assistance can change not just efficiency, but style, ownership, and diversity of expression. Do LLMs write like humans? Variation in grammatical and rhetorical styles | PNAS
For business leaders and developers, the practical takeaway is governance-oriented: separate generation, critique, verification, and authorship into distinct workflow stages. If a company wants AI-assisted writing without brand flattening, factual drift, or unclear accountability, the safest pattern is not “let the model rewrite this”; it is “let the model mark suspected weaknesses, then require a human author to make the editorial judgment and write the replacement.” OpenAI’s own help guidance still warns that ChatGPT can produce inaccurate information and recommends checking sources directly when accuracy matters. Does ChatGPT tell the truth? | OpenAI Help Center
What changed and event timeline
Thomas Ptacek published “How To Write With An LLM” on his personal site, A Final Ward
The essay sets out two rules: first, do not use any words or turns of phrase suggested by the LLM; second, avoid model encouragement because praise can cause authors to preserve weak first-draft decisions.
More detail
Ptacek’s suggested workflow is: write the draft, have a model flag defects, rewrite the relevant passage yourself, and optionally compare variants in a way that hides which version is the new one.
Simon Willison posted a link-blog entry summarizing Ptacek’s argument and connecting it to Willison’s own practice. Willison wrote that he does not let LLMs write blog content, but uses them for fact-checking, spelling, grammar, and occasional thesaurus support; he also linked to his published proofreading prompt.
The post generated Hacker News discussion
That discussion is useful as community reaction, not independent verification. Willison also commented that AI fact-checking should not be blindly trusted; users should check what the model points out.
More detail
Comments include support for “second eyes” rather than “first hands,” skepticism about LLMs as judges of prose quality, and debate over whether even model critique can influence a writer’s voice.
Bruce R. Lewis published a response agreeing with the main “copyeditor not ghostwriter” thesis but disputing details. Lewis argues that a model can be useful for finding a single word and that total avoidance of encouragement may be overbroad.
More detail
This is independent commentary, not empirical validation.
Capabilities and access
The source story does not document a specific deployed model, benchmark, architecture, release channel, or API version. Separately, Google describes Antigravity CLI as bringing the Antigravity agent harness into a terminal, but Ptacek’s post does not depend on Google’s product claims or document a specific Antigravity configuration.
Read the full section
The source story does not document a specific deployed model, benchmark, architecture, release channel, or API version. Ptacek mentions using “a good model,” later names GPT5 in an anecdotal proofreading example, and suggests a local workshopping tool could run prompts through Codex, Claude, or Antigravity CLIs. No exact model snapshot, provider settings, temperature, system prompt, context length, or evaluation harness is specified. How To Write With An LLM — A Final Ward
The only concrete implementation stack in the essay is for a proposed writing-workshopping application: Python, HTMX, SQLite, and Tailwind, with a Notion-style editor, highlighted passages, sidebar commentary, revision tracking, and navigation through suggestions. That is a product/workflow sketch, not a reproducible software release. How To Write With An LLM — A Final Ward
Separately, Google describes Antigravity CLI as bringing the Antigravity agent harness into a terminal, but Ptacek’s post does not depend on Google’s product claims or document a specific Antigravity configuration. GitHub - google-antigravity/antigravity-cli: Antigravity CLI brings the reasoning, execution, and orchestration capabilities of Antigravity agent harness directly into your terminal. · GitHub
Technical analysis for researchers and developers
The documented architecture is a human-authored draft plus LLM critique loop: This is closer to a diagnostic assistant than a generative writing system. Any claim that this workflow is “better” should therefore be read as practitioner judgment, not independently measured performance.
Read the full section
Architecture: what is documented
The documented architecture is a human-authored draft plus LLM critique loop:
- Human writes the original draft.
- LLM flags problems: repetition, passive voice, nominalizations, weak flow, movable paragraphs, filler adverbs, grammar, typos, logic, or factual issues.
- Human rewrites the relevant passage in their own words.
- Optionally, an LLM compares original vs. revised passages, ideally without being told which one the author just produced.
- The human accepts or rejects the advice.
This is closer to a diagnostic assistant than a generative writing system. The key design choice is that model output is treated as an annotation layer, not a source of final prose. Ptacek’s own tool concept uses highlighted spans and sidebar comments, which mirrors code-review UI more than conventional chat. How To Write With An LLM — A Final Ward
Evaluation methodology: what is missing
There is no controlled evaluation in the source article. It does not measure time saved, edit quality, reader preference, factual accuracy, stylistic preservation, or brand lift. It also does not compare models. Any claim that this workflow is “better” should therefore be read as practitioner judgment, not independently measured performance. How To Write With An LLM — A Final Ward
A reproducible evaluation would need at least: paired drafts, blind human raters, task categories, author satisfaction measures, voice-preservation measures, factual-error audits, and a condition comparing ghostwriting, copyediting, grammar-only tools, and human-only editing. The PNAS study on LLM style shows that corpus-linguistic features and source-classification methods can detect systematic differences between human and model prose, suggesting a possible measurement approach for “voice drift.” Do LLMs write like humans? Variation in grammatical and rhetorical styles | PNAS
Implementation implications
For product teams building AI writing tools, the workflow implies a different UI contract:
- Prefer comments, highlights, and checklists over full rewritten paragraphs.
- Separate diagnostic suggestions from replacement prose.
- Provide a “no praise / no encouragement” mode.
- Log which passages were model-flagged and which were human-rewritten.
- Support blind A/B review, but do not treat model preference as ground truth.
- For fact-checking, require citations or source links and human verification.
Willison’s proofreading prompt is a concrete lightweight version: it asks the model to identify spelling, grammar, repetition, logical errors, factual mistakes, weak arguments, and placeholder links, while his policy says opinionated first-person writing must be his own. Prompts I use - Agentic Engineering Patterns - Simon Willison's Weblog
Claims and evidence
- Ptacek recommends LLMs as copyeditors rather than ghostwriters.
- Ptacek’s Rule One is not to use model-suggested wording.
- Willison uses LLMs for proofreading and fact-checking, not for writing blog content in his voice.
Read the full section
| Material claim | Evidence status |
| Ptacek recommends LLMs as copyeditors rather than ghostwriters. | Author-reported commentary. Directly documented in Ptacek’s essay and Willison’s link post. How To Write With An LLM — A Final Ward |
| Ptacek’s Rule One is not to use model-suggested wording. | Author-reported commentary. Directly stated in the essay and quoted by Willison. How To Write With An LLM — A Final Ward |
| Willison uses LLMs for proofreading and fact-checking, not for writing blog content in his voice. | Author-reported practice. Documented in Willison’s link post and prompt guide. Archive for Thursday, 17th September 2026 |
| LLM-generated text can have systematic stylistic differences from human text. | Independent research support. PNAS study using GPT-4o, GPT-4o Mini, and Llama 3 variants found systematic grammatical, lexical, and stylistic differences and reported reproducibility materials. Do LLMs write like humans? Variation in grammatical and rhetorical styles | PNAS |
| LLM assistance can homogenize ideas or writing. | Independent research support, task-limited. A C&C 2024 study found ChatGPT users produced less semantically distinct ideas than users of another creativity-support tool; a CHI 2025 paper found GPT-4o suggestions shifted Indian participants’ writing toward Western styles. 2402.01536 Homogenization Effects of Large Language Models on Human Creative Ideation |
| LLMs are reliable fact-checkers. | Contested / not proven here. Willison reports strong practical results with search-enabled models, but OpenAI’s help center still warns ChatGPT can be factually wrong and advises checking sources directly. How to Write with an LLM | Hacker News |
Context and prior work
Ptacek’s argument fits a broader move from “AI writes for you” toward AI as critic, reviewer, or Socratic counterpart. The empirical backdrop is mixed. Research on culturally grounded writing found that GPT-4o autocomplete suggestions changed how Indian participants wrote, pushing style toward Western norms.
Read the full section
Ptacek’s argument fits a broader move from “AI writes for you” toward AI as critic, reviewer, or Socratic counterpart. In education and writing research, some scholars argue that writing with LLMs should be taught as a dialogical practice rather than simple text outsourcing; the key question becomes who retains rhetorical agency and responsibility. Writing with ChatGPT
The empirical backdrop is mixed. Research on creative ideation suggests LLMs can make individuals feel more productive or produce more detailed outputs, while reducing distinctiveness across users. 2402.01536 Homogenization Effects of Large Language Models on Human Creative Ideation Research on culturally grounded writing found that GPT-4o autocomplete suggestions changed how Indian participants wrote, pushing style toward Western norms. 2409.11360 AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances And the PNAS work on rhetorical style found that instruction-tuned models can produce a dense, noun-heavy style and struggle to match genre variation even when prompted to imitate human text. Do LLMs write like humans? Variation in grammatical and rhetorical styles | PNAS
These studies do not validate Ptacek’s exact rules. But they make his concern plausible: if models have default stylistic attractors, then accepting model phrasings can create a convergence pressure on voice.
Limitations, safety, and contested findings
The strongest limitation is that the source is commentary, not a controlled experiment. Bruce R. Lewis agrees with using LLMs as copyeditors but argues that banning even a single word is too strong, because a model can serve as a more flexible thesaurus than conventional synonym lookup. LLM as copyeditor - Writing by Bruce R. Lewis (brlewis) Fact-checking is the highest-risk part of the workflow.
Read the full section
The strongest limitation is that the source is commentary, not a controlled experiment. Ptacek’s rules are disciplined craft advice from an experienced practitioner; they are not a measured safety standard. How To Write With An LLM — A Final Ward
There is also disagreement about strictness. Bruce R. Lewis agrees with using LLMs as copyeditors but argues that banning even a single word is too strong, because a model can serve as a more flexible thesaurus than conventional synonym lookup. LLM as copyeditor - Writing by Bruce R. Lewis (brlewis) Hacker News commenters similarly split between “don’t use it for writing,” “proofreading is enough,” and “models can be useful as reader-role simulators or fact-checking prompts.” How to Write with an LLM | Hacker News
Fact-checking is the highest-risk part of the workflow. A model can identify possible factual issues, but it can also hallucinate. OpenAI’s help center explicitly states that ChatGPT may produce factually inaccurate responses and recommends using tools like search or deep research and checking sources directly when accuracy matters. Does ChatGPT tell the truth? | OpenAI Help Center
Business and practitioner implications
For executives, this story is about workflow design and accountability, not writing-tool procurement. Teams that publish thought leadership, developer relations content, policy material, legal-adjacent guidance, or executive communications should define which uses are allowed: For developers, the opportunity is to build writing tools that act like linters for prose: flag issues, localize evidence, avoid praise, and leave the final patch to the author.
Read the full section
For executives, this story is about workflow design and accountability, not writing-tool procurement. Teams that publish thought leadership, developer relations content, policy material, legal-adjacent guidance, or executive communications should define which uses are allowed:
- Allowed: typo detection, grammar review, repetition detection, argument stress-testing, source checking, empty-link detection.
- Restricted: wholesale rewrites, undisclosed ghostwriting, unsupported factual claims, model-generated executive voice.
- Auditable: final human owner, source list, model role, and whether any model-generated sentence entered final copy.
For developers, the opportunity is to build writing tools that act like linters for prose: flag issues, localize evidence, avoid praise, and leave the final patch to the author. That design aligns better with preserving authorship than chatbots that return polished replacement paragraphs.
Sources
- Thomas Ptacek, “How To Write With An LLM”, September 17, 2026. How To Write With An LLM — A Final Ward
- Simon Willison, “How To Write With An LLM” link post and archive entry, September 17, 2026. Archive for Thursday, 17th September 2026
- Simon Willison, “Prompts I use”, including proofreader prompt. Prompts I use - Agentic Engineering Patterns - Simon Willison's Weblog
Read the full section
- Thomas Ptacek, “How To Write With An LLM”, September 17, 2026. How To Write With An LLM — A Final Ward
- Simon Willison, “How To Write With An LLM” link post and archive entry, September 17, 2026. Archive for Thursday, 17th September 2026
- Simon Willison, “Prompts I use”, including proofreader prompt. Prompts I use - Agentic Engineering Patterns - Simon Willison's Weblog
- Bruce R. Lewis, “LLM as copyeditor”, September 17, 2026. LLM as copyeditor - Writing by Bruce R. Lewis (brlewis)
- Hacker News discussion, “How to Write with an LLM.” How to Write with an LLM | Hacker News
- Reinhart et al., “Do LLMs write like humans? Variation in grammatical and rhetorical styles,” PNAS, 2025. Do LLMs write like humans? Variation in grammatical and rhetorical styles | PNAS
- Anderson, Shah, and Kreminski, “Homogenization Effects of Large Language Models on Human Creative Ideation,” C&C 2024. 2402.01536 Homogenization Effects of Large Language Models on Human Creative Ideation
- Agarwal, Naaman, and Vashistha, “AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances,” accepted at CHI 2025. 2409.11360 AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances
- OpenAI Help Center, “Does ChatGPT tell the truth?” Does ChatGPT tell the truth? | OpenAI Help Center
The source trail.
Sources (12)
How To Write With An LLM
Article text retrieved; extracted text may omit tables or interactive elements.
simonwillison.net