ModelsArchitectures & capability
Willison note frames LLMs as too consequential for computer science leaders to dismiss
Simon Willison’s September 18 note is best read as a cultural signal: LLMs have become too central to software, evaluation, AI R&D and governance for technical leaders to ignore, even while current evidence remains uneven, benchmark-dependent and often vendor-reported.
Willison’s post is commentary, not a model release, benchmark, incident report or technical experiment; its significance is rhetorical rather than empirical. [1] [13]
Independent survey and index sources support the broader point that LLMs now span pre-training, post-training, prompting, agents, evaluation, software engineering and multiple applied domains. [9] [10] [12]
the original note shows a prominent practitioner arguing that LLMs are now impossible for computer science to ignore, while surveys, AI Index coverage and software-engineering literature show the technology spreading across research and development workflows.
Read the full assessment
Anthropic’s R&D automation claims add a concrete but vendor-reported example. Implication: leaders should neither dismiss LLMs as mere chatbots nor accept capability claims uncritically; adoption needs instrumentation, independent evaluation and clear human review.
Executive brief
On September 18, 2026, Simon Willison published a short commentary note arguing that computer scientists who currently refuse to find LLMs interesting resemble geneticists ignoring a newly opened “Jurassic Park”: even if the science looks inelegant or the deployment reckless, the phenomenon is too consequential to dismiss. Stanford HAI’s 2026 AI Index frames AI performance in 2025 across language, reasoning, robotics, agentic systems, scientific domains, medicine, education, and economic impact, while recent LLM surveys describe rapid technical expansion across pre-training, post-training, prompting, agents, evaluation, and software-engineering workflows. hai.stanford.edu A Survey of Large Language Models | Frontiers of Computer Science | Springer Nature Link A survey on large language models for software engineering | Science China Information Sciences | Springer Nature Link The cautionary side of the analogy is also evidence-backed.
Read the full section
On September 18, 2026, Simon Willison published a short commentary note arguing that computer scientists who currently refuse to find LLMs interesting resemble geneticists ignoring a newly opened “Jurassic Park”: even if the science looks inelegant or the deployment reckless, the phenomenon is too consequential to dismiss. The post is not a model release, benchmark paper, incident report, or product announcement; it is a commentary artifact capturing an attitude shift: LLMs are no longer only NLP systems but general-purpose research, software, security, evaluation, and organizational infrastructure. Willison’s original page contains only the analogy and posting metadata, and his tag feed repeats the same note with the fuller “frog DNA” punchline. Note on 18th September 2026 Simon Willison on ai
The strongest evidence supporting the substance of Willison’s claim comes not from the note itself, but from the surrounding research and reporting environment. Stanford HAI’s 2026 AI Index frames AI performance in 2025 across language, reasoning, robotics, agentic systems, scientific domains, medicine, education, and economic impact, while recent LLM surveys describe rapid technical expansion across pre-training, post-training, prompting, agents, evaluation, and software-engineering workflows. hai.stanford.edu A Survey of Large Language Models | Frontiers of Computer Science | Springer Nature Link A survey on large language models for software engineering | Science China Information Sciences | Springer Nature Link
The cautionary side of the analogy is also evidence-backed. Anthropic now reports that Claude “leads” 26% of its AI R&D work under its internal automation scale, but also says Claude is not fully autonomous for any measured AI R&D subset; Anthropic’s own methodology depends partly on internal data and Claude-based judging, so it is vendor-reported evidence, not independent verification. Measurements for understanding the pace of AI development inside frontier labs \ Anthropic Independent and semi-independent sources also warn that existing capability benchmarks and safety evaluations can be incomplete, marketing-shaped, or hard to compare across labs. What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks Inside the scramble for trusted AI cops
Bottom line for AI practitioners and leaders: the note is best read as a cultural signal, not a factual claim requiring a benchmark table. Its defensible interpretation is: LLMs have become too entangled with computer science, software engineering, AI R&D, and safety governance to ignore, even if one remains skeptical of hype, benchmark claims, or deployment practices.
What changed and event timeline
Willison posted the note on his weblog
The page labels it as a note, tags it under AI, generative AI, and LLMs, and contains no technical experiment, model name, benchmark, or source links beyond site metadata.
Same-day feed context
Willison’s AI tag feed places the note among posts about AI security, Claude Code, OpenAI alignment reporting, ChatGPT Work, GPT-6 Astra, and agentic coding.
More detail
That context matters: the “Jurassic Park” analogy appears amid a broader stream of posts about LLMs moving from chat into coding agents, security auditing, agent monitoring, and AI R&D infrastructure.
Surrounding news context
AP reported that Anthropic says Claude is helping build future versions of itself, and separately reported on the broader debate over recursive self-improvement. The AP account attributes the 26% AI-led R&D figure to Anthropic and notes the systems remain under human supervision.
- Corroboration status
The original Willison page and his own feed containing the same note were found in the reviewed sources
No independent reporting specifically about Willison’s note was found in the reviewed sources. That means the “event” is corroborated as a published commentary item, but its importance must be assessed through adjacent evidence, not through independent coverage of the note itself.
Capabilities and access
No specific model, version, API, architecture, or access tier is named in Willison’s September 18 note. Willison’s tag feed says Claude Code added AGENTS.md support in version 2.1.277, and separately references GPT-6 Astra and ChatGPT Work in nearby posts; those are contextual, not part of the September 18 analogy.
Read the full section
No specific model, version, API, architecture, or access tier is named in Willison’s September 18 note. The post refers to LLMs in general, not to GPT-6 Astra, Claude, Gemini, or another named system. Note on 18th September 2026
Relevant adjacent systems mentioned in nearby material include Anthropic’s Claude, Claude Code, GPT-6 Astra, and ChatGPT Work, but these are not the subject of the note itself. Willison’s tag feed says Claude Code added AGENTS.md support in version 2.1.277, and separately references GPT-6 Astra and ChatGPT Work in nearby posts; those are contextual, not part of the September 18 analogy. Simon Willison on ai
Anthropic’s September 2026 measurement post uses the generic “Claude” brand and reports internal automation levels for Anthropic R&D. It states that as of August 2026, Claude led 26% of Anthropic’s AI R&D work, while no measured subset was fully autonomous. This is vendor-reported and methodology-dependent. Measurements for understanding the pace of AI development inside frontier labs \ Anthropic
Technical analysis for researchers and developers
The Willison note contains no architecture claims. For software engineering, a 2026 survey in Science China Information Sciences reports that LLM-based techniques now cover a broad set of software-engineering tasks. A September 2026 arXiv metaresearch paper mapped 14,767 arXiv papers introducing or updating LLM evaluation resources from January 2022 through August 2026.
Read the full section
Architecture: documented only at the field level
The Willison note contains no architecture claims. For the broader LLM field, a 2026 survey in Frontiers of Computer Science describes LLM development across four documented axes: pre-training, post-training, utilization strategies, and evaluation. Pre-training establishes broad capabilities through large-scale self-supervised learning, architecture choices, and data curation; post-training adapts systems through supervised fine-tuning and reinforcement learning; utilization covers prompting, in-context learning, and agentic workflows; evaluation spans language, reasoning, safety, and other benchmarked abilities. A Survey of Large Language Models | Frontiers of Computer Science | Springer Nature Link
For software engineering, a 2026 survey in Science China Information Sciences reports that LLM-based techniques now cover a broad set of software-engineering tasks. The paper summarizes 62 representative code LLMs, 15 pre-training objectives, 16 downstream task categories, and 926 studies across 112 code-related tasks. These numbers are from the review authors’ survey methodology, not independent evidence that every task is production-ready. A survey on large language models for software engineering | Science China Information Sciences | Springer Nature Link
Evaluation methodology: benchmarks are central but unstable
A September 2026 arXiv metaresearch paper mapped 14,767 arXiv papers introducing or updating LLM evaluation resources from January 2022 through August 2026. Its key finding for practitioners is not “models are good” or “models are bad,” but that evaluation expectations themselves are changing: benchmarks increasingly test action, interaction, and professional applications, while LLMs are also becoming involved in scoring and evaluation design. The authors explicitly raise the risk that expanding evaluation may reproduce the preferences and blind spots of the models participating in evaluation. What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks
This directly affects reproducibility. A leaderboard score or vendor-reported task success rate is not enough unless the benchmark’s data provenance, contamination controls, scoring method, model/tool scaffold, number of attempts, latency/cost constraints, and human baseline are documented.
AI R&D automation as a concrete implementation frontier
Anthropic’s internal measurement post is one of the most relevant adjacent sources because it operationalizes what “LLMs are interesting to computer science” now means: not just code completion, but AI participation in the production of frontier AI systems. Anthropic says it built an R&D Automation Index by cataloging AI R&D tasks, rating automation levels from no involvement to full autonomy, and aggregating them by estimated person-time. It reports that Claude “leads” 26% of Anthropic’s AI R&D work and collaborates or leads in more than 90%, but also acknowledges obstacles to cross-lab comparison, including lack of common methodology and the use of its own models to evaluate its systems. Measurements for understanding the pace of AI development inside frontier labs \ Anthropic
A separate March 2026 paper on measuring AI R&D automation argues that existing capability benchmarks may not reflect real-world automation or capture broader consequences such as whether AI accelerates capabilities faster than safety, whether oversight keeps pace, or whether AI-introduced errors escape human review. The paper recommends operational metrics such as researcher time allocation, AI use in high-stakes decisions, and subversion incidents. Measuring AI R&D Automation
For developers, the practical implication is that LLM adoption should be instrumented like a distributed systems rollout: log agent actions, capture provenance, evaluate failure modes, measure review latency, track human override rates, and separate coding throughput from verified correctness.
Claims and evidence
- Willison published the “Jurassic Park” LLM analogy on September 18, 2026. — Direct source: Willison’s page and feed.
- The note is commentary, not a model release or technical report.
- LLMs are now technically relevant across software engineering, agentic use, evaluation, and safety.
Read the full section
| Material claim | Evidence status |
| Willison published the “Jurassic Park” LLM analogy on September 18, 2026. | Direct source: Willison’s page and feed. Note on 18th September 2026 Simon Willison on ai |
| The note is commentary, not a model release or technical report. | Direct source: the page contains a short note and metadata only. Note on 18th September 2026 |
| LLMs are now technically relevant across software engineering, agentic use, evaluation, and safety. | Supported by 2026 surveys and AI Index context; independent of Willison’s note. A Survey of Large Language Models | Frontiers of Computer Science | Springer Nature Link A survey on large language models for software engineering | Science China Information Sciences | Springer Nature Link hai.stanford.edu |
| Claude leads 26% of Anthropic AI R&D work. | Vendor-reported by Anthropic; repeated in AP coverage. Not independently verified from raw internal data. Measurements for understanding the pace of AI development inside frontier labs \ Anthropic Anthropic says its model Claude is helping to build the next version of itself |
| Claude is fully autonomously building successors. | Not supported. Anthropic explicitly says no measured subset is fully autonomous. Measurements for understanding the pace of AI development inside frontier labs \ Anthropic |
| AI evaluation is mature and independently reliable. | Contested. Benchmark metaresearch and Axios reporting both identify evaluation-design and independence concerns. What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks Inside the scramble for trusted AI cops |
Context and prior work
The analogy’s force comes from a mismatch between some computer scientists’ skepticism and the breadth of LLM diffusion. The AI Index widens that lens beyond software: its 2026 report includes chapters on technical performance, responsible AI, economic impact, science, medicine, education, policy, and public opinion. The most contested prior-work thread is AI R&D automation.
Read the full section
The analogy’s force comes from a mismatch between some computer scientists’ skepticism and the breadth of LLM diffusion. Software engineering is the clearest technical domain: code generation, automated repair, testing, fuzzing, commit messages, refactoring, compiler optimization, vulnerability detection, and agentic development now appear throughout the survey literature. A survey on large language models for software engineering | Science China Information Sciences | Springer Nature Link
The AI Index widens that lens beyond software: its 2026 report includes chapters on technical performance, responsible AI, economic impact, science, medicine, education, policy, and public opinion. That breadth supports the claim that LLMs and adjacent foundation models have become a general research object, not merely a chatbot category. hai.stanford.edu
The most contested prior-work thread is AI R&D automation. Anthropic presents recursive self-improvement as both a potential source of scientific acceleration and a possible control risk. Its own analysis says systems capable of building successors would make securing, monitoring, and shaping model behavior more important. When AI builds itself \ Anthropic The independent measurement literature similarly treats AI R&D automation as uncertain but important enough to require explicit metrics and governance. Measuring AI R&D Automation
Limitations, safety, and contested findings
First, Willison’s note is rhetorically strong but empirically thin. It does not demonstrate that any specific model has crossed a technical threshold. Axios separately reports debate over who can credibly serve as third-party AI evaluators, with concerns about close ties among evaluators, labs, and the broader AI safety community.
Read the full section
First, Willison’s note is rhetorically strong but empirically thin. It does not demonstrate that any specific model has crossed a technical threshold.
Second, much of the most dramatic current evidence comes from frontier labs’ internal reporting. Anthropic’s automation metrics are useful, but they are based on Anthropic’s internal workflows, a chosen automation rubric, person-time weighting, and Claude-assisted classification. Anthropic itself acknowledges that cross-lab comparison is blocked by methodological differences and that judge models can share failure modes with the systems being judged. Measurements for understanding the pace of AI development inside frontier labs \ Anthropic
Third, safety governance remains unsettled. The Future of Life Institute’s Summer 2026 AI Safety Index says leading companies still lack elements such as genuinely independent audits, quantitative thresholds, and clear decision authority in some frameworks; its methodology uses public materials, company survey responses, and expert panel grading, which is informative but still judgment-based. AI Safety Index — Summer 2026 | Future of Life Institute Axios separately reports debate over who can credibly serve as third-party AI evaluators, with concerns about close ties among evaluators, labs, and the broader AI safety community. Inside the scramble for trusted AI cops
Fourth, benchmark inflation and benchmark dependence remain live problems. The benchmark-design paper’s warning is especially important: as models help construct, perform, and judge evaluations, “independent evidence” can become harder to establish. What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks
Business and practitioner implications
- Do not ignore LLMs because the paradigm feels inelegant.
- Generated code, agentic fixes, and AI-authored analyses should be evaluated through tests, review, provenance, monitoring, and post-deployment telemetry—not accepted because they were produced quickly.
- Track which tasks are AI-assisted, AI-led, or human-led. And maintain audit trails for agent actions.
Read the full section
- Do not ignore LLMs because the paradigm feels inelegant. Even skeptical teams should maintain internal literacy around model behavior, eval design, prompt-injection risk, agent scaffolds, and AI-assisted development.
- Separate productivity from assurance. Generated code, agentic fixes, and AI-authored analyses should be evaluated through tests, review, provenance, monitoring, and post-deployment telemetry—not accepted because they were produced quickly.
- Instrument AI work. Track which tasks are AI-assisted, AI-led, or human-led; measure defect escape rates; record review latency; and maintain audit trails for agent actions.
- Demand evidence granularity from vendors. Ask for model/version identifiers, scaffolding details, eval prompts, pass/fail criteria, contamination controls, human baselines, and whether results were independently reproduced.
- Plan for organizational bottlenecks. If AI reduces implementation cost, bottlenecks move to specification, validation, security review, product judgment, compliance, and incident response.
Sources
Key sources used: Simon Willison’s original note and AI tag feed; Stanford HAI 2026 AI Index; 2026 LLM and software-engineering surveys; Anthropic’s AI R&D automation measurement posts; AP reporting on Claude and recursive self-improvement; Future of Life Institute’s Summer 2026 AI Safety Index; Axios reporting on evaluator independence; and March/September 2026 research papers on AI R&D automation and benchmark design.
The source trail.
Sources (14)
Note on 18th September 2026
simonwillison.netBeing a computer scientist who refuses to find anything about LLMs interesting right now is a bit like being a geneticist who refuses to find anything interesting about t
Related coverage; assess separately
x.com