Sep 18 edition/Podcast
AgentsBusinessInfrastructureSafetyModels

AgentsAutonomy & tool use

Zapier CEO frames no-code automation as headless MCP infrastructure for AI agents

Wade Foster’s interview positions Zapier less as a visual workflow destination and more as infrastructure for agents working inside users’ preferred AI environments. The strongest substantiated takeaway is architectural: pair agent planning with deterministic, governed workflow execution rather than handing business processes fully to models.

Illustration from The Cognitive Revolution: Zapier CEO frames no-code automation as headless MCP infrastructure for AI agents
Image: The Cognitive Revolution — Original article ↗
THE CORE IDEAS4 TAKEAWAYS
01

Foster says knowledge workers are converging on a preferred AI workspace, so Zapier is moving its automation capabilities into those environments through MCP rather than relying only on its own workflow UI. [1] [6] [15]

02

The proposed operating model separates agent reasoning from execution: agents help interpret requests, build or modify workflows, and troubleshoot, while repeatable steps should be hardened into deterministic automation where possible. [1] [8] [9]

03

AutomationBench gives practitioners a more relevant signal than chat-style evaluations because it scores whether simulated business systems end in the correct state, but Zapier’s leaderboard uses a private held-out set. [8] [9] [16]

04

MCP expands agent reach across business apps and credentials, making governance controls, scoped authorization, auditability, and prompt-injection defenses central deployment requirements rather than optional add-ons. [3] [4] [10] [12]

WHY IT MATTERS

the reviewed materials document Zapier MCP access to many app actions, AutomationBench’s final-state workflow evaluation, and security guidance warning that MCP-style tool access creates risks around credentials, permissions, tool poisoning, and exfiltration.

Read the full assessment

Implication: for AI practitioners and business leaders, the near-term opportunity is not replacing operations with autonomous agents, but using agents to design, maintain, and monitor governed automations that remain testable, auditable, and cost-aware.

Executive brief

Zapier CEO Wade Foster’s September 17, 2026 Cognitive Revolution appearance is best read as a strategic update from an automation incumbent adapting to agentic AI, not as independently verified research. The core thesis: “no-code” is shifting from humans dragging boxes into workflows toward agents writing and maintaining deterministic workflows, exposed “headlessly” inside users’ preferred AI workspaces such as Cursor, Claude Code, ChatGPT, VS Code, or Copilot Studio. Zapier’s hosted leaderboard currently lists GPT-6 Astra (Max) at 41.4% task completion on its private held-out set; however, this is Zapier-reported leaderboard data, not an independent reproduction.

Read the full section

Zapier CEO Wade Foster’s September 17, 2026 Cognitive Revolution appearance is best read as a strategic update from an automation incumbent adapting to agentic AI, not as independently verified research. The core thesis: “no-code” is shifting from humans dragging boxes into workflows toward agents writing and maintaining deterministic workflows, exposed “headlessly” inside users’ preferred AI workspaces such as Cursor, Claude Code, ChatGPT, VS Code, or Copilot Studio. Foster argues that Zapier’s future is less about being the daily UI and more about giving agents governed access to business apps, credentials, actions, workflow state, and reusable automations through Zapier MCP. The episode page itself frames the discussion around “headless automation tools like Zapier MCP” and “deterministic code” for reliable systems. No Code Is Code: Zapier CEO Wade Foster on Headless Tools, Zapier MCP & Automation Bench

The most important evidence anchor is AutomationBench, Zapier’s benchmark for cross-application business workflow execution. Zapier’s hosted leaderboard currently lists GPT-6 Astra (Max) at 41.4% task completion on its private held-out set; however, this is Zapier-reported leaderboard data, not an independent reproduction. AutomationBench: AI Agent Benchmarks | Zapier Artificial Analysis independently benchmarks a related AutomationBench-AA variant, but it uses a different headline metric—average objectives completed without guardrail violations—so its scores should not be compared directly with Zapier’s strict pass/fail leaderboard. AutomationBench-AA: Agentic SaaS Workflow Benchmark | Artificial Analysis

The practitioner takeaway is sober: agents are improving, but business automation still needs scoped permissions, deterministic execution, audit logs, human verification, and cost governance. Foster’s repeated message is that AI should reason only where reasoning is needed; much of the rest should be ordinary code or workflow logic. That claim is plausible engineering guidance, but the episode’s numerical internal claims—such as “80%” of agent usage being better suited to deterministic code, “nearly 100%” employee AI usage, and $30,000/month token-spend outliers—remain speaker-reported and uncorroborated in the sources reviewed.

What changed and event timeline

  1. Zapier introduced AutomationBench as an open benchmark for business workflow automation

    Zapier described the benchmark as evaluating Sales, Marketing, Operations, Support, Finance, and HR tasks selected from common workflow patterns, with deterministic final-state scoring rather than LLM-as-judge scoring.

    More detail

    The companion paper describes the benchmark as cross-application orchestration via REST APIs, requiring endpoint discovery, layered policy-following, and filtering of irrelevant or misleading records.

  2. Zapier’s public GitHub repository and official leaderboard diverged in an important way: the GitHub public task set is for experimentation, while the official leaderboard uses a separate private held-out set.

    More detail

    The repository explicitly warns that private-set scores may not match local public-set scores and that private tasks may be changed to keep the benchmark challenging.

  3. OpenAI released GPT-6 Astra, according to OpenAI’s own launch page, with API availability as gpt-6-astra and rollout across ChatGPT plans, OpenAI API, Microsoft Azure, and AWS Bedrock.

    More detail

    Axios independently reported the launch and noted OpenAI’s unusually strong claims around AGI and cybersecurity capability, including that Astra was treated as reaching OpenAI’s “critical” cybersecurity threshold.

  4. The Cognitive Revolution episode was published

    In the transcript, Foster says users are consolidating around a single “daily driver” AI environment and that Zapier needs to work inside those environments rather than forcing users into Zapier-owned agent builders [00:06:37–00:08:46].

    More detail

    The published episode page includes the same framing in its show notes.

  5. As of today, Zapier’s MCP documentation says Zapier MCP connects MCP clients to Zapier accounts, exposing 9,000+ apps and 40,000+ actions, with Zapier handling connections, credentials, and rate limits.

    More detail

    Microsoft’s connector documentation, meanwhile, describes a Zapier MCP preview connector for Copilot Studio, Power Automate, Power Apps, and Logic Apps, but lists 8,000+ apps and 30,000+ tools—likely a versioning/scope discrepancy rather than independent confirmation of Zapier’s latest count.

Capabilities and access

Zapier’s official help page defines Zapier MCP as Zapier’s implementation of the Model Context Protocol, connecting an MCP client—such as Claude, ChatGPT, Cursor, or VS Code—to tools backed by Zapier app connections. OpenAI’s own launch page also reports AutomationBench 41.4% for GPT-6 Astra, apparently matching Zapier’s leaderboard, but that remains vendor-published performance reporting.

Read the full section

Zapier MCP. Zapier’s official help page defines Zapier MCP as Zapier’s implementation of the Model Context Protocol, connecting an MCP client—such as Claude, ChatGPT, Cursor, or VS Code—to tools backed by Zapier app connections. Zapier says each tool maps to a single action, such as sending a Slack message or creating a calendar event, and that successful MCP tool calls consume two Zapier tasks. What is Zapier MCP? – Zapier

Client coverage. Zapier’s quickstart lists popular clients including Claude, Claude Code, ChatGPT, Cursor, VS Code, Gemini CLI, Replit, Warp, Windsurf, Zed, and Microsoft Copilot Studio, provided the client supports MCP over Streamable HTTP. Zapier MCP quickstart: connect your MCP client and run your first tool call This supports Foster’s “headless” thesis: Zapier’s value can be delivered inside third-party AI workbenches instead of primarily through Zapier’s own visual editor.

AutomationBench. Zapier’s hosted leaderboard currently reports GPT-6 Astra variants at the top of its strict pass/fail private evaluation, with GPT-6 Astra (Max) at 41.4% and $1.77 cost per task. Those figures are Zapier-reported and should be treated as benchmark results under Zapier’s methodology, not general proof that Astra completes 41.4% of all real business work. AutomationBench: AI Agent Benchmarks | Zapier OpenAI’s own launch page also reports AutomationBench 41.4% for GPT-6 Astra, apparently matching Zapier’s leaderboard, but that remains vendor-published performance reporting. GPT-6 Astra: A new generation of intelligence | OpenAI

Technical analysis for researchers and developers

The technically significant idea in the episode is the separation of planning/building from execution. Foster’s claim is that the agent should understand the business request, synthesize or modify a workflow, write deterministic code where possible, and reserve LLM reasoning for steps that require interpretation, ambiguity resolution, or judgment [00:13:46–00:16:16; 00:18:02–00:18:30].

Read the full section

The technically significant idea in the episode is the separation of planning/building from execution. Foster’s claim is that the agent should understand the business request, synthesize or modify a workflow, write deterministic code where possible, and reserve LLM reasoning for steps that require interpretation, ambiguity resolution, or judgment [00:13:46–00:16:16; 00:18:02–00:18:30]. This maps to a practical architecture:

  1. Intent capture: user asks in a daily driver.
  2. Tool discovery: the agent uses Zapier MCP to discover enabled app actions.
  3. Workflow synthesis: the agent constructs either a one-off action chain or a reusable Zap/workflow.
  4. Deterministic hardening: stable steps become code or explicit workflow nodes.
  5. AI islands: LLM calls remain only where classification, summarization, fuzzy matching, or policy interpretation is required.
  6. Human-readable visualization: visual workflows remain valuable for verification and documentation, even if they are no longer the primary authoring interface [00:18:02].

AutomationBench is relevant because it evaluates the failure mode practitioners care about: not “did the model give a plausible answer?” but “did the external state end up correct?” Zapier says each task boots a simulated company, lets the agent interact through API-like tools, and grades the final data state with fixed assertions. AutomationBench: AI Agent Benchmarks | Zapier The benchmark’s public repository describes trigger data, initial state, domain-specific tools, and assertion-based rubrics, with strict pass/fail only when every assertion passes. AutomationBench/README.md at main · zapier/AutomationBench · GitHub

For reproducibility, the public AutomationBench task set is useful but insufficient to reproduce Zapier’s official leaderboard because the leaderboard uses a private held-out set. AutomationBench/README.md at main · zapier/AutomationBench · GitHub Artificial Analysis helps by independently running AutomationBench-AA, but it changes the metric to objective share net of guardrail violations, so it is better viewed as a complementary evaluation rather than direct verification of Zapier’s leaderboard. AutomationBench-AA: Agentic SaaS Workflow Benchmark | Artificial Analysis

Claims and evidence

  • Users are consolidating around one “daily driver” AI tool, and Zapier is moving headless into those tools.
  • Zapier MCP exposes thousands of apps/actions to MCP clients.
  • GPT-6 Astra is top on Zapier’s AutomationBench private leaderboard at roughly 40%.
Read the full section
Material claimEvidence status
Users are consolidating around one “daily driver” AI tool, and Zapier is moving headless into those tools.Speaker-reported in episode [00:06:37–00:08:46]; supported directionally by Zapier MCP client docs. Zapier MCP quickstart: connect your MCP client and run your first tool call
Zapier MCP exposes thousands of apps/actions to MCP clients.Company-documented. Zapier says 9,000+ apps and 40,000+ actions; Microsoft’s connector page lists 8,000+ apps and 30,000+ tools for its preview connector. What is Zapier MCP? – Zapier
GPT-6 Astra is top on Zapier’s AutomationBench private leaderboard at roughly 40%.Vendor/benchmark-host reported. Zapier lists GPT-6 Astra (Max) at 41.4%; OpenAI repeats the 41.4% result. Not independently reproduced in the same metric in sources reviewed. AutomationBench: AI Agent Benchmarks | Zapier
AutomationBench evaluates cross-app business workflows with deterministic final-state grading.Documented by Zapier and paper; independently profiled by Artificial Analysis. AutomationBench
Most agent usage should be deterministic code instead.Speaker-reported judgment. Technically plausible, but the “80%” figure in the transcript is not independently corroborated.
Zapier uses multi-agent internal review loops and weekly workflow-recommendation agents.Speaker-reported internal practice [00:24:30–00:26:26; 00:37:07–00:40:58]. No public technical write-up found.
MCP introduces security risks around prompt injection, tool poisoning, over-scoped credentials, and exfiltration.Independently supported. OWASP and academic analyses identify these as core MCP risks. MCP Security - OWASP Cheat Sheet Series

Context and prior work

Zapier’s repositioning follows the broader standardization of tool access around MCP. The MCP authorization specification defines OAuth-based transport-level authorization for HTTP transports and emphasizes resource-specific token validation. modelcontextprotocol/docs/specification/2025-06-18/basic/authorization.mdx at main · modelcontextprotocol/modelcontextprotocol · GitHub The July 2026 MCP specification update added hardening details such as issuer validation, credential binding, cacheable list results, and header-based routing for gateway/WAF use cases.

Read the full section

Zapier’s repositioning follows the broader standardization of tool access around MCP. The MCP authorization specification defines OAuth-based transport-level authorization for HTTP transports and emphasizes resource-specific token validation. modelcontextprotocol/docs/specification/2025-06-18/basic/authorization.mdx at main · modelcontextprotocol/modelcontextprotocol · GitHub The July 2026 MCP specification update added hardening details such as issuer validation, credential binding, cacheable list results, and header-based routing for gateway/WAF use cases. The 2026-07-28 Specification | Model Context Protocol Blog

AutomationBench also sits in a line of agent benchmarks that move beyond static QA and coding toward long-horizon tool use. Its distinctive emphasis is ordinary SaaS workflow execution: CRM, spreadsheets, inboxes, calendars, ticketing, messaging, policy documents, and final-state assertions. AutomationBench Artificial Analysis characterizes AutomationBench-AA as a benchmark for agentic task completion in simulated SaaS application environments and explicitly marks its evaluations as independently conducted. AutomationBench-AA: Agentic SaaS Workflow Benchmark | Artificial Analysis

Limitations, safety, and contested findings

The biggest limitation is evidentiary: the episode is commentary and interview testimony. Zapier’s leaderboard, Zapier’s public GitHub task set, and Artificial Analysis’s AutomationBench-AA differ in task sets and metrics. A 2026 academic threat-modeling paper found major client-side weaknesses in tested MCP clients, especially insufficient static validation and limited parameter visibility.

Read the full section

The biggest limitation is evidentiary: the episode is commentary and interview testimony. It gives valuable operator insight, but internal Zapier claims about adoption, architecture, token spend, and support automation are not independently verified.

The second limitation is benchmark comparability. Zapier’s leaderboard, Zapier’s public GitHub task set, and Artificial Analysis’s AutomationBench-AA differ in task sets and metrics. Researchers should not cite a single “AutomationBench score” without specifying hosted private strict pass/fail, public task set pass rate, or AutomationBench-AA objective-score metric. AutomationBench: AI Agent Benchmarks | Zapier

Security is the strongest contested area. MCP’s value proposition—connecting agents to tools and credentials—is also its attack surface. OWASP lists tool poisoning, rug-pull attacks, cross-server tool shadowing, confused-deputy problems, data exfiltration through legitimate channels, excessive permissions, and supply-chain attacks as MCP risks. MCP Security - OWASP Cheat Sheet Series A 2026 academic threat-modeling paper found major client-side weaknesses in tested MCP clients, especially insufficient static validation and limited parameter visibility. Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning Zapier’s own compliance page says its AI Guardrails are not 100% accurate and should be used as one defense-in-depth layer, not a replacement for human review and security controls. Zapier AI Automation Platform: Legal and Compliance Information

OpenAI’s Astra system card adds another caution: although OpenAI says Astra is more robust and aligned than GPT-5.6 Sol, it also reports decreased chain-of-thought monitorability and says absence of observed failures does not establish reliability across settings. GPT-6 Astra System Card - OpenAI Deployment Safety Hub That matters for agentic workflow systems because reliable external action depends on both capability and monitorability.

Business and practitioner implications

For business leaders, the practical message is not “replace workflows with agents.” Foster’s episode claim is that organizations will want to swap models by workflow, not standardize on one provider [00:10:21–00:12:53]. AutomationBench’s leaderboard already exposes a cost/task column, and Artificial Analysis separately tracks token usage and cost metrics.

Read the full section

For business leaders, the practical message is not “replace workflows with agents.” It is “use agents to discover, build, and maintain workflows, then harden repeatable work into governed automation.” That aligns with Foster’s “no-code is code” framing [00:18:02]: the user experience becomes conversational, but the production substrate should remain auditable, scoped, and testable.

For developers, Zapier MCP can reduce integration burden by giving agents access to prebuilt app actions, but it introduces governance work: action allowlists, workspace segmentation, approval flows, logging, and credential hygiene. Zapier’s enterprise materials emphasize audit logs, app/action controls, managed connections, domain restrictions, and log streaming; those controls are relevant prerequisites for serious deployment. Zapier AI Automation Platform: Legal and Compliance Information

For AI teams, model routing becomes a cost/performance discipline. Foster’s episode claim is that organizations will want to swap models by workflow, not standardize on one provider [00:10:21–00:12:53]. AutomationBench’s leaderboard already exposes a cost/task column, and Artificial Analysis separately tracks token usage and cost metrics. AutomationBench: AI Agent Benchmarks | Zapier

Sources

Read the full section
FOLLOW THE EVIDENCE

The source trail.

Sources (17)
A LITTLE LESS NOISE. A LOT MORE CONTEXT.

Stay curious.
Follow the evidence.

Independent perspectives, the original sources, and room for the questions that don't have easy answers.

How we build the brief