Sep 14 edition/Reporting & analysis
AgentsCodingBusinessSafety

AgentsAutonomy & tool use

Voss argues AI coding agents could make product judgment the new software bottleneck

Simon Willison’s quote of Laurie Voss’s essay frames AI coding agents as more than productivity tooling: if code generation keeps getting cheaper, the scarce work may move to product discovery, specification, review, operations and customer context. Evidence shows agentic coding activity, not full lifecycle automation.

THE CORE IDEAS4 TAKEAWAYS
01

Voss’s central claim is strategic rather than product-specific: as coding agents reduce the cost of implementation, durable value may shift toward deciding what should be built, specifying it well, and judging whether it serves users. [1] [6]

02

Repository evidence supports rising agentic coding activity, including measurable agent-authored pull requests, but the reviewed studies do not establish that agents can reliably replace human review, product reasoning, or system-level accountability. [2] [11] [14]

03

Cheaper software generation can create downstream governance pressure: review gaps, maintainer overload, repository controls, and low-quality security reports all point to triage and verification becoming more important. [4] [7] [9]

04

The labor implication is unsettled but consequential: descriptive data show pressure on young workers in AI-exposed roles, while Voss’s thesis suggests teams may need more deliberate pathways into product, domain, and operational judgment. [3] [6]

WHY IT MATTERS

Evidence in the reviewed research shows real agent-authored code appearing in repositories, observable PR datasets, and reported stress on review channels; it does not show that agents can autonomously own product intent, production operations, or long-term quality.

Read the full assessment

The implication for practitioners is to treat AI coding as an organizational design issue, not just a tooling upgrade: requirements quality, test oracles, review records, deployment governance, customer context, and junior training may become more decisive than raw code throughput.

Executive brief

On September 14, 2026, Simon Willison published a short “quotation collected” post highlighting Laurie Voss’s essay “We are all Product Engineers now.” The underlying event is not a model launch or product release; it is an argument about how AI coding agents may shift the scarce work in software from writing code toward product discovery, specification, taste, review governance, deployment, and operation. Willison’s post quotes Voss’s thesis that code-writing costs have “collapsed,” that review/fix/operate costs may follow, and that the remaining durable work is finding out what users actually want and making software pleasant to use.

Read the full section

On September 14, 2026, Simon Willison published a short “quotation collected” post highlighting Laurie Voss’s essay “We are all Product Engineers now.” The underlying event is not a model launch or product release; it is an argument about how AI coding agents may shift the scarce work in software from writing code toward product discovery, specification, taste, review governance, deployment, and operation. Willison’s post quotes Voss’s thesis that code-writing costs have “collapsed,” that review/fix/operate costs may follow, and that the remaining durable work is finding out what users actually want and making software pleasant to use. A quote from Laurie Voss

For practitioners and executives, the dossier-level takeaway is: do not treat AI coding as merely an engineering-productivity tool. If Voss’s forecast is even partially right, the bottleneck moves from code production to requirements quality, human context extraction, design judgment, verification, and change management. Available evidence supports parts of the trend—coding-agent activity is rising, GitHub activity has grown, agent PRs are measurable at scale, and young workers in AI-exposed occupations are under pressure—but the evidence does not yet prove that the entire software-development lifecycle will be automated or that “product engineer” will become the dominant future role. Octoverse: A new developer joins GitHub every second as AI leads TypeScript to #1 - The GitHub Blog

What changed and event timeline

  1. Laurie Voss published “We are all Product Engineers now,” a long-form commentary essay on Seldo.com

    Voss frames the essay as a forecast over roughly a decade, explicitly built on two large assumptions: first, that agents will eventually handle the software-development lifecycle beyond code generation; second, that demand for software has no practical upper bound over that horizon.

  2. Simon Willison reposted a short excerpt as a quote item on his weblog, classifying it under tags including AI, LLMs, generative AI, careers, and agentic engineering. This is secondary commentary, not independent verification of Voss’s claims.

  3. Also

    Prior context cited by Voss

    The essay points to adjacent claims about junior developer labor-market weakness, growing coding-agent use, agent-authored pull requests, coding benchmarks, and forward-deployed engineering roles. Live search found relevant support for several of these background trends, but not independent coverage of Willison’s specific quote post beyond repost/aggregation-style pages.

Capabilities and access

There is no single model, model version, API, product access tier, or system card associated with this story. The source item is commentary quoting an essay. For business readers: this means the story should not be evaluated like a vendor benchmark.

Read the full section

There is no single model, model version, API, product access tier, or system card associated with this story. The source item is commentary quoting an essay. Where the essay references AI systems, it does so generically as “agents,” with examples elsewhere in the evidence base including Claude Code, OpenAI Codex, GitHub Copilot coding agent, Devin, Cursor, OpenHands, Aider, Gemini CLI, and Google Big Sleep. Exact underlying foundation-model versions are usually not documented in the cited labor and repository-mining studies; the Claude Code PR study identifies Claude Code as the agentic coding tool but does not, in the excerpt retrieved, pin a model snapshot. On the Use of Agentic Coding: An Empirical Study of Pull Requests on GitHub

For business readers: this means the story should not be evaluated like a vendor benchmark. It is better read as a strategic thesis about how the software labor stack may reorganize if agentic coding, agentic review, and agentic operations continue improving. The access question is therefore organizational, not technical: which teams have coding agents, how agent output enters repositories, how review is recorded, and who owns product definition. We are all Product Engineers now | Seldo.com

Technical analysis for researchers and developers

Voss’s technical argument decomposes software work into categories: code writing already “collapsed”; review and maintenance are “going soon”; deployment and scaling are “next”; product discovery, definition of “good,” and design delight are comparatively durable. The agentic PR evidence is mixed. The repository-mining evidence shows measurement is itself hard.

Read the full section

Voss’s technical argument decomposes software work into categories: code writing already “collapsed”; review and maintenance are “going soon”; deployment and scaling are “next”; product discovery, definition of “good,” and design delight are comparatively durable. This is a forecast, not a measured causal study. The strongest technical evidence nearby is narrower: agents can now produce repository-level patches, and researchers can observe those patches in real open-source workflows. We are all Product Engineers now | Seldo.com

The agentic PR evidence is mixed. One empirical study of 567 Claude Code pull requests across 157 open-source projects reports that many were accepted and that over half of merged PRs needed no further modification, but it also says remaining PRs benefited from human revision, especially for bug fixes, documentation, and project-specific standards. That supports the view that agent code can be useful; it does not prove agents understand product intent or maintain system-level quality unaided. On the Use of Agentic Coding: An Empirical Study of Pull Requests on GitHub

The repository-mining evidence shows measurement is itself hard. The AIDev dataset reports 932,791 agentic PRs across 116,211 repositories, plus a curated subset of 33,596 PRs from 2,807 repositories with richer review metadata. A separate census of coding-agent commits argues that single-signal detection undercounts activity, because bot-account lookup captured only a small fraction of Claude Code commits in one snapshot. For reproducibility, this matters: benchmark, PR, and commit studies can disagree because they observe different channels of agent activity. AIDev: Studying AI Coding Agents on GitHub

Review remains a bottleneck. A practitioner analysis summarizing an EASE 2026 study of 33,596 agent-authored PRs says most had no recorded review, and that among reviewed agent PRs, many were reviewed only by other agents; the author explicitly warns that “no recorded review” does not prove no human looked silently. This caveat is central: GitHub artifacts measure recorded workflow, not cognition. Still, for developers building agent workflows, the implication is concrete: require durable reasoning artifacts, issue-to-test traceability, CI evidence, and human sign-off policies for high-risk changes. 61% of AI-authored pull requests are never reviewed — Aaditya Kushwaha

Benchmark evidence also needs caution. SWE-bench’s official site describes variants such as Verified, Multilingual, Multimodal, Lite, and Full, and says the metric is percentage of task instances resolved. But new work on SWE-Bench Pro Verified argues that prior Pro evaluation can be undermined by reward hacking and task-quality issues, potentially inflating reported capability. That is directly relevant to Voss’s “review/fix/operate will follow” assumption: benchmark saturation is not the same as production readiness. SWE-bench Leaderboards

Security work shows both upside and downside. Google Project Zero and DeepMind’s Big Sleep reported a real-world SQLite memory-safety vulnerability discovered by an LLM-assisted agent, and TechCrunch later reported Google’s claim that Big Sleep found and reported 20 vulnerabilities in open-source projects such as FFmpeg and ImageMagick, with human experts still in the reporting loop. Conversely, curl’s bug bounty shutdown illustrates how cheap AI-generated reports can overload maintainer attention. From Naptime to Big Sleep: Using Large Language Models To Catch Vulnerabilities In Real-World Code - Project Zero

Claims and evidence

  • Willison quoted Voss on September 14, 2026. — Supported by primary/secondary page evidence.
  • “Writing code cost collapsed.”
  • “Review, fixing, and operating will follow.” — Forecast / inference, not established.
Read the full section
Material claimStatusEvidence
Willison quoted Voss on September 14, 2026.Supported by primary/secondary page evidence.Willison quote page and Seldo essay. A quote from Laurie Voss
“Writing code cost collapsed.”Voss commentary; directionally supported by agent adoption, not independently proven as an economic universal.GitHub reports rapid AI-linked developer activity growth; PR studies show agent-authored code in real projects. Octoverse: A new developer joins GitHub every second as AI leads TypeScript to #1 - The GitHub Blog
“Review, fixing, and operating will follow.”Forecast / inference, not established.Some coding benchmarks and PR studies support progress; benchmark-reliability work and review-gap evidence caution against overclaiming. On the Use of Agentic Coding: An Empirical Study of Pull Requests on GitHub
Young workers in AI-exposed occupations are under pressure.Supported descriptively, not causally proven.Stanford says employment for ages 22–25 in highly AI-exposed occupations is about 19% below a peer-tracking counterfactual and explicitly says the data are descriptive, not causal. No Widespread Displacement, but the AI Employment Gap for Young Workers Has Widened to 19% - Stanford Digital Economy Lab
Forward-deployed / customer-embedded engineering roles are growing.Plausible, but data quality varies by source.Salesforce describes FDEs as in-demand and customer-embedded; Plank reports a first-party census and third-party posting growth but should not be treated as neutral labor-market proof. Today’s Hottest Role: Forward Deployed Engineer - Salesforce
Product discovery and design taste remain hard to automate.Reasoned argument, not directly measured.Voss’s essay; supported indirectly by tacit-vs-codified labor evidence, but no direct benchmark validates “taste.” We are all Product Engineers now | Seldo.com

Context and prior work

The “product engineer” thesis revives older software-industry roles rather than inventing a wholly new category. Stanford’s Digital Economy Lab reports no broad economy-wide AI displacement in its dashboard update, while finding a widening gap for young workers in highly AI-exposed occupations and a distinction between codified and tacit knowledge. The open-source context is also important.

Read the full section

The “product engineer” thesis revives older software-industry roles rather than inventing a wholly new category. Voss argues that systems analysts and later product managers historically sat between business/user needs and programmers, and that AI may collapse parts of that separation back into a senior hybrid role. Whether or not one accepts the history as complete, the practical point is familiar to experienced teams: ambiguous requirements and weak product judgment generate more waste than slow typing. We are all Product Engineers now | Seldo.com

The labor-market context is unsettled. Stanford’s Digital Economy Lab reports no broad economy-wide AI displacement in its dashboard update, while finding a widening gap for young workers in highly AI-exposed occupations and a distinction between codified and tacit knowledge. That maps onto Voss’s claim that junior “well-specified ticket” work is most exposed, but Stanford warns that the patterns are descriptive and cannot yet isolate AI from other labor-market forces. No Widespread Displacement, but the AI Employment Gap for Young Workers Has Widened to 19% - Stanford Digital Economy Lab

The open-source context is also important. GitHub says 2025 saw major growth in developers, repositories, pull requests, and commits, and its July 2026 Innovation Graph update says rapid contribution growth has strained communities enough that GitHub shipped maintainer controls such as PR limits, repo-level PR/issue controls, pinned comments, noise-reduction banners, and temporary interaction limits. This supports the downstream-bottleneck framing: cheaper contribution creation increases the value of triage and governance. Octoverse: A new developer joins GitHub every second as AI leads TypeScript to #1 - The GitHub Blog

Limitations, safety, and contested findings

The largest limitation is that Voss’s essay is explicitly speculative. HackerOne told BleepingComputer that AI-helped report volume is increasing across security, while also arguing that managed triage and AI workflows can buffer enterprises better than small open-source teams. Benchmark contamination and task quality also remain contested.

Read the full section

The largest limitation is that Voss’s essay is explicitly speculative. It assumes agents will “eat” the SDLC, including review, testing, deployment, monitoring, and scaling. Current evidence is strongest for code generation and patch production, weaker for autonomous review quality, and thinner for production operations and scaling. It would be premature to restructure an engineering organization on the assumption that operations expertise is about to become optional. We are all Product Engineers now | Seldo.com

Safety concerns are not peripheral. The curl bug-bounty case shows that AI can create human-attention denial-of-service failure modes: reports may be cheap to generate but costly to validate. HackerOne told BleepingComputer that AI-helped report volume is increasing across security, while also arguing that managed triage and AI workflows can buffer enterprises better than small open-source teams. This is a contested operational point: AI can aid vulnerability discovery and flood vulnerability intake. Curl ending bug bounty program after flood of AI slop reports

Benchmark contamination and task quality also remain contested. SWE-Bench Pro Verified’s authors argue that leakage and flawed tasks can inflate model performance, while SWE-bench itself remains a widely used benchmark family. For researchers, the lesson is to report scaffold, tool access, model snapshot, cost budget, pass criteria, contamination controls, and human-intervention policy—not just a headline score. SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

Business and practitioner implications

For executives, the main implication is organizational design. For engineering managers, the risk is losing the apprenticeship ladder. Voss argues that junior developers historically learned judgment by implementing bounded tickets and receiving senior review; Stanford’s young-worker data and the rise of agent-authored PRs make that concern credible, though not conclusively causal.

Read the full section

For executives, the main implication is organizational design. If AI reduces code-production cost, competitive advantage shifts toward knowing what to build, validating that it is useful, operating it safely, and embedding it in customer workflows. That favors cross-functional engineers, forward-deployed teams, domain experts, designers, product managers with technical depth, and engineering leaders who can define rigorous human/agent handoffs. We are all Product Engineers now | Seldo.com

For engineering managers, the risk is losing the apprenticeship ladder. Voss argues that junior developers historically learned judgment by implementing bounded tickets and receiving senior review; Stanford’s young-worker data and the rise of agent-authored PRs make that concern credible, though not conclusively causal. Teams should create deliberate apprenticeships around product discovery, debugging, incident review, customer interviews, test design, and codebase stewardship rather than assuming juniors will learn by “just coding.” We are all Product Engineers now | Seldo.com

For developers, the durable skill bundle is likely to include domain modeling, requirements elicitation, test-oracle design, code review, observability, incident analysis, security triage, UX judgment, and agent orchestration. The tactical response is not to abandon coding craft, but to pair it with evidence production: every agent-generated change should answer what user problem it solves, what alternatives were rejected, what tests prove, what risks remain, and who accepted the tradeoff. All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code

Sources

Primary/commentary sources: Simon Willison’s quote post; Laurie Voss’s original Seldo essay. A quote from Laurie Voss Independent or research sources: Stanford Digital Economy Lab labor-market update; arXiv/ACM work on agent-authored PRs; AIDev dataset; SWE-Bench Pro Verified; SWE-bench official benchmark site.

Read the full section

Primary/commentary sources: Simon Willison’s quote post; Laurie Voss’s original Seldo essay. A quote from Laurie Voss

Independent or research sources: Stanford Digital Economy Lab labor-market update; arXiv/ACM work on agent-authored PRs; AIDev dataset; SWE-Bench Pro Verified; SWE-bench official benchmark site. No Widespread Displacement, but the AI Employment Gap for Young Workers Has Widened to 19% - Stanford Digital Economy Lab

Supporting ecosystem and vendor sources: GitHub Octoverse and Innovation Graph updates; Salesforce FDE materials; Plank FDE census; Google Project Zero/DeepMind Big Sleep; BleepingComputer reporting on curl’s bug-bounty shutdown. Octoverse: A new developer joins GitHub every second as AI leads TypeScript to #1 - The GitHub Blog

FOLLOW THE EVIDENCE

The source trail.

Sources (14)
A LITTLE LESS NOISE. A LOT MORE CONTEXT.

Stay curious.
Follow the evidence.

Independent perspectives, the original sources, and room for the questions that don't have easy answers.

How we build the brief