Why I wrote this
I did not write this as a market report. I wrote it because I have become genuinely curious about building products people can actually use. That curiosity turned into a lot of research, across the board.
I am a tech guy at heart. After building SAVRN, nearly three hundred agents now, each with a specific job role, I have a real understanding of how hard this actually is. We test every agent team on our own company first: our front end, our back end, our books, our sales desk. None of it goes near a customer before that.
That work taught me something simple. Small and mid-sized businesses, and a surprising share of enterprises, do not lack access to models. They lack the infrastructure, the technical skills, and the engineering discipline to run a structured agent team safely. So I went looking for the evidence. This essay is that research, in full, with every source linked. It ends where my own work starts.
What the evidence says
Two adoption curves are pulling apart. At the frontier, agentic AI has moved past single-task copilots. It is now multi-agent systems that plan, act, and coordinate across functions with little human input. McKinsey's 2025 Global Survey finds 23% of organizations already scaling an agentic AI system somewhere in the enterprise (McKinsey). Gartner's inaugural Hype Cycle for Agentic AI shows 17% of organizations have deployed AI agents, and more than 60% expect to within two years. In one analyst summary's words, that is the fastest adoption-intent curve Gartner has recorded (Enterprise DNA on Gartner).
At the same time, the evidence for genuinely autonomous, multi-step deployments is thin even at the enterprise level. For small businesses it is nearly nonexistent. I reviewed the flagship small-business surveys: U.S. Chamber, Intuit QuickBooks, GoDaddy, Salesforce, the Marketing AI Institute, Constant Contact. None of them separates "using a chatbot" from "deploying an autonomous agent." None reports a penetration rate for autonomous executive-function agent teams.
The number shaping the public narrative is "95% of GenAI pilots fail." It comes from MIT NANDA's State of AI in Business 2025 report. It rests on 52 executive interviews, 153 survey responses, and more than 300 implementation reviews. It found that only 5% of custom enterprise AI tools reach production with measurable financial impact, even as more than 90% of employees already use personal LLMs for work, off the books (MIT NANDA / AI Governance Library).
That gap, between informal individual use and durable, production-grade deployment, is what this essay is about. Model capability does not explain it. A missing engineering discipline does, and that discipline gets its own section below. That same discipline gap is what turns "adoption" into an affordability question for any company without dedicated AI engineering capacity.
The essay draws on primary data from Stanford HAI, MIT NANDA, McKinsey, Deloitte, Gartner, IBM IBV, Bain, BCG, PwC, Anthropic, a16z, Menlo Ventures, the U.S. Chamber of Commerce, Intuit, GoDaddy, the Census Bureau, and vendor pricing pages. It draws on the public architectural literature from Anthropic, AWS, LangChain, OpenAI, and NIST, and on production case studies from Klarna, Sierra, and Cognition. It builds a framework for the "can't afford to keep up / can't afford not to" paradox facing small and mid-sized businesses. It closes with what that means for anyone building infrastructure to serve them.
Adoption Reality: Pilots, Production, and the Size Divide

The topline numbers disagree by definition, not just by data
Every major 2025 survey confirms broad-based experimentation with generative AI. The surveys diverge sharply once the question shifts to production or agentic deployment. Most of that divergence comes from definitions: each organization defines "adoption," "agent," and "production" differently. Stanford HAI's 2025 AI Index reports that 78% of organizations used AI in 2024, up from 55% in 2023. It also reports that generative AI use in at least one business function more than doubled over the same period, from 33% to 71% (Stanford HAI). McKinsey's 2025 Global Survey (1,993 respondents across 105 countries, fielded June 25–July 29, 2025) puts regular use of AI in at least one business function at 88%, up from 78% the prior year. Nearly two-thirds of respondents, though, say their organizations have not yet begun scaling AI across the enterprise. Only about one-third report having begun enterprise-wide scaling (McKinsey State of AI). McKinsey is also one of the few major 2025 surveys to publish a clean agentic-AI-specific adoption figure. That figure is 23% of organizations scaling an agentic AI system somewhere in the enterprise (McKinsey State of AI).
Anthropic's Economic Index draws on Claude.ai and API usage data through August 2025. It offers a firm-level adoption measure, not just an employee-level one, and it is unusually concrete. The share of U.S. firms reporting AI use in production processes rose from 3.7% in fall 2023 to 9.7% by early August 2025. That is more than double. It still leaves roughly nine out of ten U.S. businesses reporting no AI use in production processes as of mid-2025 (Anthropic Economic Index). The same report notes that reported employee-level AI use at work (Gallup data) rose from 20% to 40% between 2023 and 2025. That is a doubling in two years, and Anthropic explicitly benchmarks it against historical technology diffusion. The internet took roughly five years to reach an adoption rate that AI reached in two. Electricity took over 30 years to reach farm households after urban electrification. The personal computer took roughly 20 years to reach a majority of U.S. homes after 1981 (Anthropic Economic Index).
The MIT NANDA "95% failure" finding, in full context
The MIT NANDA (Networked Agents and Decentralized Architectures) State of AI in Business 2025 report, variously reported as "The GenAI Divide," is the single most cited pilot-failure statistic in the 2025–2026 discourse. It deserves precise sourcing. The authors are Aditya Challapally, Chris Pease, Ramesh Raskar, and Pradyumna Chari. The report is based on 52 executive interviews, 153 survey responses, and over 300 public AI initiative reviews. The July 2025 report found that 80% of firms have tried generative AI, but only 5% of organizations are translating AI pilots into real operational or financial impact. Separately, only 5% of custom internal AI tools reach production (MIT NANDA / AI Governance Library summary). Forbes' coverage of the report frames the core diagnosis precisely: pilots fail "because companies avoid friction." Companies lean on generic tools that are fine for demos but too brittle for real workflows. The surviving 5% succeed by embedding GenAI deeply into high-value workflows with memory and learning loops, rather than treating it as a bolt-on chat feature (Forbes). Menlo Ventures' own enterprise report explicitly cites the MIT finding as having "rattled markets over the summer" of 2025. That shows how influential the statistic became, and how contested (Menlo Ventures).
The finding's critics generally do not dispute the underlying data so much as its framing. The 95% figure describes custom-built, internally engineered tools. It does not describe vendor-native AI features embedded in existing SaaS (Copilot, Agentforce, HubSpot Breeze), which show materially higher retention because they inherit the vendor's engineering, data plumbing, and support burden. This distinction is build versus buy, and specifically the gap between ad hoc internal builds and vendor-engineered production systems. It recurs throughout this essay and is central to the SMB affordability calculus.
I have my own reference point here. In 2023 I built a catalog of 80 AI case studies across eight industries: transportation and logistics, entertainment and sports venues, energy and gas pipelines, healthcare, agriculture, industrial manufacturing, insurance, and supply chain. Ten companies in each. I did it on my own, years before this essay, because I wanted to understand what AI was actually doing inside operating companies rather than what vendors said it would do. Two years later I re-verified every company: the initiative, the financials, and every quantified outcome, checked against current sources, with each result graded as operator-disclosed or vendor and secondary. The method is different from MIT's. MIT NANDA worked from 52 executive interviews, 153 survey responses, and roughly 300 public implementation reviews in a single window. Mine is longitudinal and company-level: the same 80 cases followed across two years, with every number traced to who actually disclosed it. Two findings from that work bear directly on this essay. First, across all eight verticals AI crossed from pilot to operational standard at the leading firms, but the gains are bounded, use-case-specific, and unevenly evidenced. The eye-catching numbers almost always trace to vendor projections or secondary blogs rather than operator disclosures. Second, by 2024–2026 the leading edge had already pivoted from AI that recommends to AI that acts across a workflow. That is what the leaders look like. The rest of this essay is about everyone below them. The full study, with the verified-versus-vendor split, is at savrn.com/p/ai-integration.
Enterprise leadership sentiment: aggressive intent, uneven results
IBM's Institute for Business Value fielded three complementary 2025 studies: the "AI at the Core" survey of 2,500 executives across 18 industries and 19 regions, an "Agentic AI Pulse" survey of 400 C-suite executives across 15 roles, 11 industries, and six countries, and a separate study of 2,000 CEOs. Together the results show a leadership class racing ahead of realized value. 70% of executives call agentic AI important to their organization's future. 83% expect agents to improve process efficiency by 2026. 61% of CEOs report they are actively adopting AI agents today while preparing to implement at scale. Yet only 25% of AI initiatives delivered expected ROI over the past few years, and only 16% were scaled enterprise-wide (IBM IBV, June 2025, IBM IBV, May 2025). Half of CEOs admit that rapid AI investment left their organization with disconnected, piecemeal technology. 64% acknowledge that fear of falling behind is driving investment before the organization understands the technology's value. That is a direct empirical anchor for the affordability paradox explored later in this essay (IBM IBV). Gartner's own 2026 figures (17% deployed, 60%+ intending to deploy within two years) sit inside a broader warning. Gartner's June 2025 press release, based on a January 2025 poll of 3,412 webinar attendees, predicts that over 40% of agentic AI projects will be canceled by the end of 2027 due to escalating costs, unclear business value, or inadequate risk controls. Gartner analyst Anushree Verma notes that "most agentic AI projects right now are early-stage experiments or proof of concepts that are mostly driven by hype and are often misapplied" (Gartner). Gartner's Hype Cycle for Agentic AI places the category squarely at the "Peak of Inflated Expectations." It maps more than 30 distinct agentic technologies (MCP, multi-agent orchestration, agent governance frameworks, agent management platforms). It states plainly that "fully autonomous agents are not ready for the majority of enterprise use cases today." The organizations seeing results are using well-scoped agents, constrained workflows, and human oversight rather than open-ended autonomy (Gartner Hype Cycle analysis).
Bain's 2025 Technology Report frames the maturity curve in four capability levels. Level 1 (LLM-powered retrieval: copilots, knowledge assistants) was scaled by tech-forward enterprises in 2023–2024. Together with Level 2, it delivered 10–25% EBITDA gains through "grab-a-coffee" microproductivity. Level 2 (single-task agentic workflows) and Level 3 (cross-system orchestration with supervision) are where 2025 capital and deployment velocity are converging. Level 4 (multi-agent constellations, any-to-any agent collaboration) remains, in Bain's words, "on the whiteboard" (Bain Technology Report 2025). This maps almost exactly onto where executive-function agent teams sit. Level 4 is the definition of an autonomous C-suite cell, and Bain's own framing places it beyond what even leading enterprises have operationalized in 2025.
Enterprise vs. mid-market vs. small business: the comparison table
Precise, comparable adoption figures broken out cleanly by company size are rare. Most surveys either sample only enterprises (McKinsey, Deloitte, Bain, IBM) or only small businesses (U.S. Chamber, Intuit, GoDaddy) without a shared methodology. The U.S. Census Bureau's Business Trends and Outlook Survey (BTOS) is the most rigorous exception. It offers a nationally representative, biweekly cross-section by firm size.
| Segment | AI usage rate | Source / period | Notes |
|---|---|---|---|
| Firms with 250+ employees | 37% currently using AI | U.S. Census Bureau BTOS, Dec 2025–May 2026 | Largest firms are the biggest AI users |
| Firms with 100–249 employees | 32% currently using AI | Census BTOS, same period | Mid-market lags large enterprise by ~5 points |
| Firms with fewer than 20 employees | No statistically significant change | Census BTOS, same period | Flat over the observation window |
| Firms with 4 or fewer employees | Less than 20% currently using AI | Census BTOS, same period | Smallest firms trail furthest |
| All organizations, general AI use | 78% (2024), up from 55% (2023) | Stanford HAI AI Index 2025 | Enterprise-weighted sample, not SMB-specific |
| Organizations scaling agentic AI somewhere in enterprise | 23% | McKinsey State of AI 2025 | 38% of respondents from firms with >$1B revenue |
| Small businesses using AI-enabled tools (any) | 98% | U.S. Chamber Empowering Small Business, 2024 ed. | Extremely broad definition of "AI-enabled" |
| Small businesses self-identifying as using generative AI | 58% | U.S. Chamber Empowering Small Business 2025 | Up from 40% (2024) and 23% (2023) |
| Small businesses (≤100 employees) using AI regularly | 68% | Intuit QuickBooks, April 2025 | Up from 48% in July 2024; n=2,200+ |
| Small businesses using AI daily | 28% | Intuit QuickBooks, April 2025 | Subset of the 68% figure |
Here is the pattern. Census data shows large firms far ahead of micro-firms, while the small-business-focused surveys (Chamber, Intuit) show high and rapidly rising "AI use." That is not a contradiction so much as a definitional gap. Chamber and Intuit numbers capture any AI tool touch point (a chatbot, an image generator, an autofill feature embedded in QuickBooks or Shopify). Census BTOS asks about AI use "in business operations," a bar that filters out casual or peripheral use. None of these surveys, at any company size, isolate autonomous multi-step agents from single-turn generative tools. I return to that gap in Section 3.
Ask “AI in business operations” and the smallest firms trail. Ask “any AI tool” and nearly everyone says yes.
Blue bars are the Census Bureau’s nationally representative question about AI in business operations. Copper bars are small-business surveys that count any touch point: a chatbot, an image generator, an autofill feature. Neither set isolates autonomous, multi-step agents.
Pilot-to-production conversion: what the data actually supports
No enterprise survey I reviewed provides a clean numerator/denominator pilot-to-production conversion rate with a stated methodology beyond MIT NANDA's 5% figure. IBM's data offers an adjacent proxy: only 25% of AI initiatives delivered expected ROI, and only 16% were scaled enterprise-wide, over "the last few years." That is consistent with MIT's 5% figure, though not identical to it, since IBM's "scaled" bar may be less strict than MIT's "reaches production with real financial impact" (IBM IBV). Deloitte's fourth-quarter State of Generative AI in the Enterprise survey (2,773 director-to-C-suite respondents across 14 countries, fielded July–September 2024) found that 26% of organizations are exploring autonomous agent development "to a large extent" and 42% "to some extent." It did not publish a combined or production conversion figure (Deloitte). Independent 2026 commentary triangulating multiple sources reports that roughly 75% of enterprise leaders told Forrester (June 2026) they are adopting agentic AI. Gartner's 2026 Hype Cycle data found only 17% have actually deployed agents. Deloitte's own December 2025 readiness data found just 11% have production-ready agentic systems. That is a consistent three-to-one gap between stated adoption intent and verified production deployment across independent research houses (Digital Applied synthesis of Gartner/Forrester/Deloitte).
The Vendor and Tooling Landscape SMBs Face
Platform tier: broad agentic infrastructure
The market has split in two. On one side sit enterprise-first platforms with opaque, quote-only pricing. On the other sit self-serve platforms with published usage-based pricing that a smaller buyer can actually read. Microsoft Copilot Studio comes bundled at no additional cost for Microsoft 365 Copilot licensees (from $30/user/month) for basic internal agents. Scaling usage is a different story. That requires Copilot Credit capacity packs at $200/month per 25,000-credit pack, or a metered pay-as-you-go tier billed per agent action. In practice, the real SMB cost is opaque until usage patterns are established (Microsoft). Salesforce Agentforce runs three pricing models at once: Flex Credits at roughly $500 per 100,000 credits (about $0.10 per agent action, with a standard action consuming ~20 credits and, per the third-party breakdown, covering up to 10,000 tokens of processing), a per-conversation rate (about $2, per a third-party pricing breakdown; Salesforce's page lists it as contact-for-pricing), and a flat per-user license at $125/user/month, plus a $5/user/month baseline "Agentforce User License" tier (Salesforce; pricing breakdown). Amazon Bedrock Agents inherit standard Bedrock token pricing (for example, legacy Claude 3.5 Sonnet at $6/$30 per million input/output tokens under extended-access pricing, current Sonnet-class models at $3/$15, with a 50% batch discount). So agent cost there is a function of model choice and orchestration efficiency, not a flat agent fee. That is a model that rewards engineering skill over a simple purchase (AWS).
At the self-serve end, two platforms are aimed most directly at SMB budgets and technical capacity. Zapier prices by task (free at 100 tasks/month, Professional from $19.99/month, Team from $69/month, Enterprise by quote). n8n sells tiered plans, explicitly marketed as "Pro: for solo builders and small teams" and "Business: for companies with <100 employees," with a 50% Start-up discount for firms under 20 employees (Zapier; n8n). Lindy AI publishes transparent per-seat pricing: $29.99/month (Plus, 3,000 credits), $99.99/month (Pro, 15,000 credits), and $199.99/month (Max, 35,000 credits), with Enterprise adding SSO, audit logs, and HIPAA compliance via signed BAA. That makes it a rare example of an agent-building platform with fully public, SMB-legible pricing (Lindy). Relevance AI, by contrast, publishes only an Enterprise tier with no listed price. It is positioned explicitly for "companies looking to decouple growth from headcount via an AI Workforce," which is to say enterprise-only by design (Relevance AI).
Executive-agent-adjacent and vertical specialist tools
Cognition's Devin is the most prominent standalone engineering agent, and it publishes accessible tiers: Free ($0, small quota), Pro ($20/month with daily/weekly quotas across OpenAI, Claude, Gemini, and Cognition's own SWE-1.6 models), Max ($200/month with the highest quota), Teams ($80/month base plus $40/month per full developer seat), and custom Enterprise pricing with SOC 2, SSO, and data residency (pricing summary). Sierra, the Bret Taylor-founded customer-experience agent platform, uses outcome-based pricing. Customers pay only for successful resolutions, not for unresolved queries or human escalations, and there is no published rate. Contracts combine a base platform fee, per-outcome charges, and professional-services fees that "can exceed licensing costs." It is positioned explicitly as an enterprise-only "Agent Operating System" sold through custom sales-led contracts to buyers like ADT, Chime, Cigna, Nordstrom, Nubank, Ramp, and Wayfair, including 40% of the Fortune 50 (Sierra pricing analysis; Sierra). Harvey (legal) and Hebbia (finance/research) sit at the far enterprise end. Harvey's reported per-seat pricing runs $100–$500/user/month, with median verified annual contracts around $175,000 and typical minimums of 25–50+ seats. Hebbia publishes no pricing at all and positions itself for "the world's leading financial, legal, and Fortune 100 firms" (Harvey pricing analysis; Hebbia). Sana (Sana Agents) sits closer to SMB reach, with a $30/user/month Team tier (up to 50 members) alongside a free tier (5 members, 10 meetings/month) and custom Enterprise pricing (Sana).
Open-source and framework layer
Anthropic's own engineering guidance is clear about where the framework layer fits. Frameworks like the Claude Agent SDK, AWS's Strands Agents SDK, and GUI builders like Rivet and Vellum simplify "standard low-level tasks like calling LLMs, defining and parsing tools, and chaining calls together." At the same time, "they often create extra layers of abstraction that can obscure the underlying prompts and responses, making them harder to debug," and "can also make it tempting to add complexity when a simpler setup would suffice." Anthropic explicitly advises reducing abstraction layers as systems move to production (Anthropic, Building Effective Agents). I read that as a direct signal about the open-source agent-framework layer (LangGraph, CrewAI, AutoGen, OpenAI's Agents SDK, which explicitly replaced the experimental "Swarm" framework). It lowers the barrier to a working prototype. It does not, by itself, close the gap to production reliability. Engineering discipline layered on top of the framework closes that gap. The framework's existence does not.
Vendor landscape comparison table
| Vendor / platform | Target segment | Pricing model | Approximate cost | Agent capability tier |
|---|---|---|---|---|
| Microsoft Copilot Studio | Enterprise, mid-market | License + credit packs | $30/user/mo base; $200/mo per 25k-credit pack (Microsoft) | Level 1–2 (retrieval, single-task) |
| Salesforce Agentforce | Enterprise, mid-market | Flex credits / per-conversation / per-user | ~$0.10/action; ~$2/conversation; $125/user/mo (Salesforce; pricing breakdown) | Level 2–3 |
| AWS Bedrock Agents | Developers, enterprise | Token usage (model-dependent) | $3/$15 per 1M tokens (current Sonnet-class); $6/$30 legacy 3.5 Sonnet (AWS) | Level 1–3, build-your-own |
| Zapier Agents / n8n | SMB, solo builders | Task-based / tiered subscription | Free tier; n8n Business <100 employees, 50% startup discount (n8n) | Level 1–2 |
| Lindy AI | SMB, solo builders, small teams | Per-seat + shared credit pool | $29.99–$199.99/user/mo (Lindy) | Level 1–2 |
| Cognition Devin | Individual developers, small eng teams | Seat + usage tiers | $20–$200/mo self-serve; custom Enterprise (pricing) | Level 2–3 (coding agent) |
| Sierra | Enterprise only | Outcome-based, custom contract | Undisclosed; base fee + per-outcome + services (eesel.ai) | Level 3 (production CX agents) |
| Harvey | Enterprise (legal), 25–50+ seat minimum | Per-seat, custom contract | $100–$500/user/mo; ~$175k median annual contract (costbench.com) | Level 2–3 (vertical specialist) |
| Hebbia | Enterprise (finance/legal) | Custom, undisclosed | Not published; Fortune 100-oriented (Hebbia) | Level 2–3 (vertical specialist) |
| Sana (Sana Agents) | SMB to enterprise | Per-seat + free tier | Free (5 seats); $30/user/mo (Team, 50 seats); custom Enterprise (Sana) | Level 1–2 |
| Generic no-code SMB agent tools | Small business (10–75 employees) | Subscription or managed setup | $100–$500/mo subscription; $3,000–$12,000 managed setup; $15,000+ custom build (Advantech IT) | Level 1 (single workflow) |
The pattern in this table is stark. The platforms with transparent, SMB-legible pricing (Lindy, n8n, Zapier, Sana's free/Team tiers, Devin's self-serve tiers) top out at Bain's Level 1–2 capability: retrieval and single-task workflows. Every platform operating at Level 3 (cross-system orchestration) or working on executive-function use cases (Sierra, Harvey, Hebbia) is enterprise-only, quote-gated, and priced in five- or six-figure annual contracts with 25+ seat minimums. There is no publicly priced, self-serve product in this market that delivers genuine multi-agent executive orchestration. The pricing gap is also a capability gap.
The pricing gap is also a capability gap. Every product that operates above Level 2 is enterprise-only, quote-gated, and priced in five- or six-figure annual contracts.
Every platform with SMB-legible pricing tops out at Bain’s Level 1–2. Everything above it is quote-gated.
Bain’s four capability levels, mapped against who publishes a price. There is no publicly priced, self-serve product that delivers genuine multi-agent executive orchestration.
LLM-powered retrieval
Copilots, knowledge assistants. Scaled by tech-forward enterprises in 2023–2024; Bain reports 10–25% EBITDA gains from “grab-a-coffee” microproductivity.
Public, self-serve pricingLindy · n8n · Zapier · Sana · Copilot Studio
Single-task agentic workflows
One agent, one workflow, a human on escalation. Where 2025 capital and deployment velocity are converging, together with Level 3.
Public, self-serve pricingDevin self-serve · Lindy · Agentforce
Cross-system orchestration with supervision
Agents that call tools across systems under human oversight. Sold through custom, quote-gated enterprise contracts.
Enterprise-only, quote-gatedSierra · Harvey · Hebbia
Multi-agent constellations
Any-to-any agent collaboration, the definition of an autonomous executive team. In Bain’s words, still “on the whiteboard.”
No packaged product existsNo one, yet
SMB-Specific Evidence: What Small Businesses Are Actually Doing
The headline SMB numbers, and their limits
The U.S. Chamber of Commerce's Empowering Small Business series, run with Teneo Research since 2022, is the most-cited SMB AI dataset. Its 2025 edition (published August 18, 2025) reports that 58% of small businesses self-identify as using generative AI. That is more than double the 23% reported in 2023 and up from 40% in 2024. It also reports that 82% of AI-using small businesses increased their workforce over the past year, and 84% plan to increase technology platform use going forward (U.S. Chamber). The prior year's edition (September 2024) reported an even broader 98% of small businesses using some AI-enabled tool, with 40% using generative AI specifically. That spread shows how sensitive these figures are to one question: does "AI-enabled" include the ambient AI features baked into existing software (autocomplete, spam filters, fraud detection), or only AI a business owner consciously chose to adopt (U.S. Chamber, 2024 ed.)? Intuit's QuickBooks Small Business Insights survey (April 2025, n=2,200+ businesses with up to 100 employees) reports that 68% now use AI regularly, up sharply from 48% in July 2024, with 28% using it daily (Intuit). Intuit's broader 2025 Small Business Index finds that businesses managing eight or more areas of their operations with digital tools were 1.6x more likely to forecast positive revenue growth and nearly twice as likely to report productivity gains (67% vs. 36%) compared to businesses using two or fewer digital tools. I take that as a strong signal that digital-tool breadth correlates with SMB performance, not any single agent deployment (the index measures digital tools, not AI specifically; the figures are U.S.-only) (Intuit QuickBooks Small Business Index 2025).
Constant Contact's Q1 2026 Small Business Now report (over 1,500 SMB owners across five countries) found 54% of small businesses already using AI marketing tools, with 27% more planning to start in 2026. That implies more than 80% adoption of AI marketing tools specifically by year-end. The most common uses were trend-data analysis (45%), content composition (44%), and image/visual generation (40%) (Forbes coverage of Constant Contact). The Marketing AI Institute and SmarterX's 2025 State of Marketing AI Report (n=1,882, with roughly half of respondents at companies under $10M in revenue and under 50 employees) finds respondents self-classify their AI adoption stage as: Curiosity 4%, Understanding 13%, Experimentation 40%, Integration 26%, Transformation 17%. The plurality of SMB marketers are still in the "actively testing tools" phase. They have not embedded AI into workflows (Marketing AI Institute / SmarterX, State of Marketing AI 2025).
The critical evidentiary gap
None of the SMB-focused primary sources I reviewed for this essay (U.S. Chamber, Intuit, GoDaddy, Salesforce SMB Trends, the Marketing AI Institute, Constant Contact) provide a quantitative breakdown of autonomous agent adoption distinct from generic AI tool use. None explicitly distinguish "using a chatbot" from "deploying an autonomous agent that takes multi-step action." That is not an oversight in the research. It reflects a genuine absence in the underlying survey instruments, and it is itself a finding. Five separate report years into the AI adoption boom, the SMB market research apparatus has not yet developed the vocabulary or measurement frameworks to separate generative-AI usage from agentic deployment. GoDaddy's Summer 2025 Small Business Survey (about 1,400 respondents, majority under nine employees) does not address AI adoption at all. A separate GoDaddy Venture Forward piece notes that "2024 saw a huge leap in the adoption of AI and the rollout of new tools" without giving a small-business adoption rate (GoDaddy Venture Forward). Salesforce's own Small & Medium Business Trends Report landing page, at the URL indexed for this research, contains no publicly accessible adoption statistics in the crawlable content (GoDaddy; Salesforce SMB Trends).
So what functions are SMBs actually automating, based on the available evidence? The pattern across every source is consistent: marketing content generation, customer-facing chat/support triage, bookkeeping and financial summarization (via QuickBooks' embedded AI), and scheduling. These are single-turn, narrow, supervised tasks that map to Bain's Level 1 (retrieval/copilot) capability tier. Not one source I reviewed documented an SMB deploying a genuinely autonomous, cross-functional, multi-agent executive decisioning system in production. The closest analogs are individual case narratives of solo-founder or very small AI-native companies (discussed in Section 6), not survey-representative evidence of incumbent SMB adoption. This is the single largest evidentiary gap in the current public research. Any claim about SMB "autonomous agent adoption rates" beyond generic AI tool use should be treated as unsupported until better-instrumented survey data exists.
The "using ChatGPT" vs. "deployed agent team" gap
The gap between casual generative AI use and deployed autonomous agent infrastructure is best explained by the same finding MIT NANDA reports at the enterprise level, scaled down: over 90% of employees already use personal LLMs for work, "bypassing stalled enterprise initiatives" (MIT NANDA). At the SMB level, this looks like an owner or employee using ChatGPT, Claude, or an embedded tool (QuickBooks Assist, Shopify Magic) for a specific task. Drafting a marketing email. Summarizing a spreadsheet. Answering a customer question. There is no orchestration layer, no persistent state, no tool integration, and no governance wrapper. That is a fundamentally different thing from an autonomous agent team that plans multi-step work, calls tools and APIs, maintains state across sessions, and takes actions with real-world consequences (moving money, publishing content, committing code) without a human in every loop. The SMB survey data overwhelmingly describes the former. There is essentially no survey-grade evidence of the latter at meaningful scale in the small-business segment.
Barriers to SMB Adoption
Ranked barriers from available survey data
IBM's Agentic AI Pulse survey (400 C-suite executives) is one of the few sources that ranks agentic-AI-specific barriers with numbers, even though its sample leans enterprise. Data concerns (49%), trust issues (46%), and skills shortages (42%) top the list (IBM IBV). BCG's global AI at Work survey (11 countries, 10,600+ respondents) finds that only about one-quarter of frontline employees get strong leadership support for AI. Only one-third report proper training. Just 13% see AI agents deeply integrated into their daily workflows. My read is that organizational and change-management friction, not technology access, is the binding constraint. That is before you even reach the small-business segment, where formal training budgets and dedicated leadership bandwidth are scarcer still (BCG AI at Work 2025).
| Barrier category | Evidence | Relative severity for SMBs |
|---|---|---|
| Data readiness / integration complexity | 49% of C-suite cite data concerns (IBM IBV); Gartner cites integration as "the defining technical challenge" (Gartner Hype Cycle) | Severe. SMBs typically lack structured data and dedicated data engineering |
| Trust / governance / liability | 46% cite trust issues (IBM IBV); Gartner's 40%+ project-cancellation forecast cites "inadequate risk controls" (Gartner) | Severe. no SMB-scale governance tooling exists off the shelf |
| Talent / skills shortage | 42% cite skills shortages (IBM IBV) | Severe. SMBs cannot compete for AI engineering talent; fractional AI leadership is an emerging but immature market |
| Cost beyond subscription (implementation, evals, monitoring) | Managed SMB agent setup runs $3,000–$12,000; custom builds $15,000+ (Advantech IT) | High. the subscription price is a fraction of total cost of ownership |
| Change management / training | Only one-third of employees report proper AI training; strong leadership support raises positive sentiment from 15% to 55% (BCG) | High. SMBs have no formal L&D function |
| Regulatory patchwork (state AI laws, EU AI Act) | 65% of small businesses worry about a patchwork of state AI laws; 77% of AI-using small businesses say technology limits would hurt growth (U.S. Chamber 2025) | Moderate-to-high. compliance burden is fixed-cost and regressive by firm size |
| Disconnected/piecemeal technology from rushed investment | 50% of CEOs admit this outcome from their own AI investment (IBM IBV) | High. SMBs have fewer resources to remediate architectural mistakes |
| Shadow AI (unsanctioned use outside governance) | Over 90% of employees use personal LLMs for work, bypassing formal initiatives (MIT NANDA) | High. SMBs have essentially no visibility into or control over this |
The EU AI Act, Colorado, and NYC: a regulatory patchwork tax on smaller firms
The U.S. Chamber's 2024 report found 86% of small businesses saying proposed technology regulations would harm their ability to grow, and less than one-third felt "very well" prepared to comply with pending AI regulations. Its 2025 edition finds 65% worried about a patchwork of state AI laws (U.S. Chamber). Structurally, this is a regressive tax. The compliance cost for frameworks like the EU AI Act's SME provisions, Colorado's AI Act, or New York City's Local Law 144 bias-audit requirement for automated hiring tools is largely fixed. It does not scale with revenue. So the burden per dollar of revenue is dramatically higher for a 20-person business than for a 20,000-person one. Enterprises absorb this cost inside their existing legal and compliance departments. An SMB has three choices: pay outside counsel per incident, avoid the regulated use case entirely, or accept unmanaged risk.
The Autonomous Executive Agent Team: A Deep Dive

Defining the category
When I say autonomous executive agent team, I mean something distinct from a copilot, a single-task agent, or even Bain's Level 3 "cross-system orchestration." It is a Level 4 construct: a coordinated fleet of agents, analogous to CEO, CFO, COO, CTO/CIO, CMO, and Chief-of-Staff functions, that plan, propose, review, and in bounded cases execute cross-functional business decisions with governance in place rather than step-by-step human scripting. McKinsey's own framing captures the aspiration: systems based on foundation models that can act in the real world, planning and executing multiple steps in a workflow (McKinsey State of AI). Salesforce CEO Marc Benioff describes the adjacent vision as a "digital workforce" in which human and automated agents collaborate toward shared customer outcomes (McKinsey, citing Benioff).
Who is building toward this, and how far they've gotten
No vendor I reviewed sells a complete, off-the-shelf, multi-function autonomous executive agent team as a packaged product. What exists are specialist agents (Sierra for CX, Harvey for legal, Hebbia for financial research, Devin for engineering) that do well inside one functional lane, plus orchestration frameworks (LangGraph, CrewAI, OpenAI's Agents SDK) that let a sophisticated engineering team stitch specialist agents into a cross-functional fleet. Sierra's own public framing of where it is headed is worth reading. Its "Ghostwriter" system is described as an "agent-building agent" that removes the need to manually edit journeys, write integrations, or triage issues. It accepts standard operating procedures as input and builds and improves sophisticated agents "reliably and autonomously" inside a sandboxed, headless version of Sierra's platform (Sierra, Agents as a Service). That is the clearest public signal I found of a vendor building toward agent teams that build agent teams. It is still scoped to customer experience, sold to enterprise accounts (including 40% of the Fortune 50), and priced through custom, outcome-based contracts that SMB budgets cannot reach.
Cognition's own internal use of Devin is the most concrete public data point on what a mature, narrow-domain autonomous agent looks like in production. In one recent week, Cognition merged 659 Devin-authored pull requests into its own codebase (versus a prior 2025 best week of 154). It used Devin Review on every PR, ran a daily automated design-system audit, and used MCP integration to reach hundreds of external tools and data sources. Hour-long bug investigations became background tasks (Cognition). This is a single-function (engineering) agent operating at real production scale with heavy human review gates. It shows what disciplined agent engineering looks like. It is a proof point for one function, not for a cross-functional executive team.
Case study: Klarna's customer-service agent as the clearest production benchmark
Klarna's OpenAI-powered customer service assistant is among the most thoroughly documented production agent deployments available to the public. It works as a useful proxy for what a single, well-scoped, executive-adjacent function (customer operations) looks like when it is built with real engineering discipline. Within the first month after its February 2024 launch, Klarna's assistant handled 2.3 million chats and automated 67% of customer service conversations. It did the equivalent work of roughly 700 full-time agents and supported 35+ languages. It cut average resolution time from 11 minutes to under 2 minutes, reduced repeat inquiries by 25%, and was projected to drive roughly $40 million in profit improvement in 2024 (OpenAI; Twig case study synthesis). By Q3 2025, the same system was handling the equivalent of roughly 853 agents' worth of work, an estimated $60 million in annual savings, and a customer NPS of 73. The other side of the ledger matters too. Roughly 30% of tickets still escalate to humans. Klarna publicly rebalanced toward human agents for part of its service in 2025. The system took approximately six months to move from kickoff to public launch, and that was inside one of the most AI-forward fintechs in the world. OpenAI noted Klarna was "the first European company and first fintech firm globally" to build on its plugin ecosystem, a first-mover advantage most SMBs cannot replicate (Twig; OpenAI). By contrast, Twig, which is itself a vendor in this category, claims that platforms like Twig, Intercom Fin, and Decagon deliver 4–8 week deployment times and 50%+ autonomous resolution rates for teams willing to buy rather than build. That is a substantial narrowing of the build-versus-buy gap for this specific function. It is still not a cross-functional executive agent team (Twig).
The vanishingly small SMB evidence base
Across every source I consulted for this essay (vendor case study pages, funding announcements, survey data, and press coverage) there is no verified, named example of a small business (under 50 employees) running a genuinely autonomous, cross-functional executive agent team in production. Emerging products marketed explicitly at this gap do exist: "AI Chief of Staff" tools (Hey Teo, Russell, Noltra, Team0), "AI CFO" agents (Archie/Arc, RunwayBot, StartupCFO.ai), and "AI COO" concepts (Notion's Startup COO agent template, Vybe). They are overwhelmingly early-stage and thinly documented, and they are aimed at very early-stage startups rather than incumbent small businesses with existing operations, staff, and legacy systems to integrate against ([search results across AI CFO/COO/Chief-of-Staff vendor pages]). One frequently cited anecdotal case is Polsia, a single-founder company reportedly reaching $10 million in ARR using AI tooling in place of a traditional team. It illustrates the AI-native greenfield case, not the incumbent-SMB retrofit case that describes most of the actual small business population (note.com case discussion). This absence is itself a central finding. The autonomous executive agent team, as a deployed reality rather than a vendor pitch or a thought-experiment, exists today almost exclusively inside well-capitalized AI-native startups and the internal tooling of frontier AI labs and infrastructure operators. It does not exist inside the incumbent SMB population this essay set out to understand.
Agent Engineering: The Missing Discipline

The evolutionary arc: from prompting to autonomous agents
The industry got to reliable agentic systems through four eras you can name. To read any adoption statistic in this essay correctly, you need to know where the frontier actually sits, as opposed to where the marketing says it sits.
The prompting era (2022–2023) treated the model as a chat interface. Single-shot completions. Prompt refinement by trial and error. Success meant "getting a good answer on the third try." The looping era (2023–2024) added tool use and basic orchestration: ReAct-style reasoning loops, plus AutoGPT and BabyAGI-style autonomous retry loops. The model calls tools and keeps retrying until it judges itself "done." Simon Willison's widely cited definition captures the mechanical core: "an LLM agent runs tools in a loop to achieve a goal," bounded by a stopping condition (Simon Willison). That definition describes the mechanism, not the discipline around it. The loop by itself guarantees nothing about reliability. This looping pattern, mostly without state discipline, guardrails, or evaluation infrastructure, is where the overwhelming majority of SMB agent experiments live today. That squares with Gartner's warning that most current agentic AI projects are "early stage experiments or proof of concepts... often misapplied" (Gartner).
The agent engineering era (2024–2026) is where the serious builders now operate. Anthropic, Cognition, Sierra, AWS, and the OpenAI Agents SDK team are among them. This era is the direct answer to the failure modes of the looping era. Anthropic's own engineering guidance draws a sharp line between workflows (LLMs and tools orchestrated through predefined code paths, so deterministic and auditable) and agents (LLMs dynamically directing their own process, so flexible but unpredictable). It recommends starting with the simplest workflow that works and adding agentic autonomy "only when simpler solutions fall short," because agentic systems trade higher latency and cost for better task performance, and give up materially more predictability in the bargain (Anthropic, Building Effective Agents). That is the core of the discipline. Treat the model as a powerful but stochastic component, an unreliable, high-latency microprocessor. Wrap it inside a rigid, deterministic software architecture that owns the rails (schema validation, database writes, retries, rollback). The model owns reasoning and generation, and only within those rails. AWS's own Well-Architected Agentic AI Lens formalizes the same distinction operationally. Because "LLM-powered decisions are inherently non-deterministic," reliability strategies for agentic systems must use behavioral monitoring, evaluation frameworks, and graceful degradation rather than deterministic testing alone. That is a fundamentally different SRE discipline than traditional software (AWS Well-Architected Agentic AI Lens).
The emerging autonomous-agent era (2026 and beyond) goes one step past agent engineering. Instead of a human scripting the workflow structure in advance, self-directed multi-agent teams operate under governance rather than scripts. A human sets intent and approves consequential side effects. The system itself plans, dispatches work, reviews its own output through independent checks, executes within bounded authority, audits its actions, and proposes improvements for human ratification. This is where Sierra's Ghostwriter (an "agent-building agent" that plans and improves other agents from natural-language SOPs) and Cognition's internal use of Devin to review and merge hundreds of its own pull requests weekly are pointing. It is the leading edge, and it is what makes the affordability gap in Section 7 existential for SMBs rather than merely uncomfortable. The frontier is moving toward governance-based autonomy at exactly the moment most SMBs are still in the undisciplined looping era.
The frontier is moving toward governed autonomy while most small businesses are still in the looping era.
The industry’s path to reliable agentic systems, in four recognizable eras. The distance between the two markers is the affordability paradox.
The reliability math: why multi-step agent workflows are unforgiving
Here is the math I keep coming back to. The central quantitative case for agent engineering as a discipline, rather than a nice-to-have, is a simple compounding-probability calculation that most agentic AI marketing skips. Say a model or agent step succeeds independently with probability p on a single action, and a workflow needs N sequential steps to succeed end-to-end. The overall success probability is not p. It is pN. A model or agent with an apparently strong 85% single-step success rate collapses to roughly 44% end-to-end reliability across just five sequential steps. That is worse than a coin flip on whether the whole workflow completes correctly.
| Per-step success rate | 1 step | 3 steps | 5 steps | 8 steps | 12 steps | 20 steps |
|---|---|---|---|---|---|---|
| 99% | 99.0% | 97.0% | 95.1% | 92.3% | 88.6% | 81.8% |
| 97% | 97.0% | 91.3% | 85.9% | 78.4% | 69.4% | 54.4% |
| 95% | 95.0% | 85.7% | 77.4% | 66.3% | 54.0% | 35.8% |
| 90% | 90.0% | 72.9% | 59.0% | 43.0% | 28.2% | 12.2% |
| 85% | 85.0% | 61.4% | 44.4% | 27.2% | 14.2% | 3.9% |
| 80% | 80.0% | 51.2% | 32.8% | 16.8% | 6.9% | 1.2% |
This table is the technical spine of why an executive agent team is qualitatively harder to build than a single-turn copilot. It is also why the gap between "impressive demo" and "reliable production system" is not a matter of degree but of kind. Take a CEO-CFO-COO agent fleet that must plan, gather data across systems, draft a recommendation, cross-check it, obtain approval, and execute. Call it a modest six or seven distinct logical steps. Even a genuinely excellent 95%-per-step model degrades to somewhere between 70% and 77% end-to-end reliability with no engineering intervention. At enterprise scale, run thousands of times per month, a 25–30% failure rate on consequential business actions is not a rounding error. It is an operational and reputational liability. This one piece of math explains two findings at once: why MIT NANDA finds only 5% of custom AI tools reach production with real impact, and why Gartner expects escalating costs and inadequate risk controls to cancel more than 40% of agentic AI projects by 2027. Most teams are deploying probabilistic components in a multi-step chain with no deterministic scaffolding to arrest compounding failure. The failure only becomes visible once the pilot scales past a handful of steps.
The gap between an impressive demo and a reliable production system is not a matter of degree. It is a matter of kind.
Agent engineering exists specifically to break this compounding-failure math. Instead of hoping each step clears a high bar on its own, the discipline separates deterministic logic from non-deterministic logic. Deterministic logic is schema validation, database writes, API calls, dependency checks: code that either succeeds or fails predictably and can be retried, checkpointed, or rolled back. Non-deterministic logic is natural-language generation, reasoning, code synthesis: the parts only a model can do. Then it wraps the non-deterministic components in verification. First, sandboxed execution, where a test suite or validator, not the model's own self-assessment, decides whether a step actually succeeded. Second, durable state and checkpointing. LangGraph's persistence layer, for instance, explicitly warns of in-memory checkpointers that "when the process restarts, all checkpoints are lost" and instructs "use a persistent checkpointer for production" backed by Postgres, precisely because production multi-step workflows must survive failures and resume rather than restart from zero (LangGraph docs). Third, circuit breakers and guardrails that can halt execution before an expensive or consequential action completes. OpenAI's Agents SDK formalizes this. It runs distinct input, output, and tool-level guardrails and can block execution outright when malicious usage is detected. The documentation puts it this way: "blocking execution guarantees that the expensive model does not start" (OpenAI Agents SDK). The practical effect of this discipline is to convert a chain of pN probabilistic steps into a chain where each step is independently checked and retried until it clears a verification bar. That decouples end-to-end reliability from the raw exponential decay shown in the table above. The cost is substantially more engineering work per workflow.
End-to-end success is p to the power of N. Move the dials.
A workflow that needs N sequential steps, each succeeding with probability p, completes with probability pN. The table above is this curve; the chart redraws as you move the sliders. Faint curves are 99%, 90% and 80% per step.
What production-grade governance actually requires

Public architectural literature converges on a common pattern for any agentic system trusted to take consequential action rather than merely draft suggestions. This pattern, not any single vendor's product, is the actual generic blueprint an SMB would need to replicate to field a genuine autonomous executive agent team. NIST's Center for AI Standards and Innovation launched an AI Agent Standards Initiative in February 2026. It is built on three pillars: industry-led standards, open-source protocols, and research on agent security and identity. It came with a request for information on AI-agent security and a concept paper on agent identity and authorization (NIST, AI Agent Standards Initiative). Practitioner profiles that map the NIST AI Risk Management Framework's Govern, Map, Measure, and Manage functions onto agentic systems add the same list of controls: containment and sandboxing for autonomous agents, human oversight of agentic task execution, monitoring for unintended goal pursuit and "reward hacking," and explicit handling of multi-agent coordination failures and cascading errors (Cloud Security Alliance, NIST AI RMF Agentic Profile). AWS's Agentic AI Lens formalizes "bounded autonomy" as a first-class architectural principle. Every agent operates within explicitly defined scope boundaries, with guardrails that constrain behavior "regardless of inputs received," alongside transparency and explainability requirements that log agent decisions for audit (AWS Well-Architected).
The human-in-the-loop approval pattern documented across production guidance is consistent and specific. Approval gates are required whenever an agent's next action is irreversible, costly, regulated, or carries a high "blast radius" (examples cited include disabling multi-factor authentication, rotating keys, writing to production databases, or changing customer attributes). The operating goal across every framework reviewed is "supervised autonomy": the agent moves quickly when it is safe and slows down for human review when it must (StackAI human-in-the-loop pattern guide). This pattern borrows directly from decades-old control disciplines outside AI. The separation-of-duties principle, that no single party should be able to both produce and approve their own work, is foundational to financial "four-eyes" controls, to secure software release processes (a different party must review and approve code before it deploys to production), and to SOC 2 / ISO 27001 audit frameworks. All of them require an independently reviewable, tamper-evident audit trail of who approved what, when, and under what authority. Layered on top is a durable "Work Order" or task-queue pattern. The request is written to persistent storage before execution begins, so a crash or timeout mid-workflow can be recovered rather than silently lost. That is standard distributed-systems practice (the outbox pattern) applied to agent orchestration. An evaluation service supporting canary releases and automatic rollback of agent behavior changes is the direct analog of modern CI/CD and SRE practice, applied to a system whose behavior, unlike traditional code, can drift between releases even without a code change.
None of this is exotic engineering on its own. Durable workflows, approval gates, audit logs, and canary rollouts are all well-understood patterns with mature tooling in traditional software. What makes it a genuine barrier for SMBs is the integration surface. A production-grade autonomous executive agent team requires all of these patterns working together at the same time, correctly, and tuned specifically to the non-deterministic, tool-calling, multi-agent context. That takes a combination of distributed-systems engineering, security engineering, financial-controls literacy, and AI evaluation expertise. Essentially no off-the-shelf SMB vendor sells that as a single integrated product today. No SMB's existing IT staff, if it has any at all, is resourced or trained to build it from primitives. This is precisely the discipline gap MIT NANDA's 95% pilot-failure statistic is measuring, expressed architecturally rather than statistically. The 5% that succeed are, empirically, the organizations that built (or bought, fully built) something resembling this governance stack. The 95% that fail did not.
Six patterns, none exotic on its own. The barrier is running all six together, correctly, at once.
The public architectural literature (NIST’s agent-standards work, the AWS Well-Architected Agentic AI Lens, OpenAI’s Agents SDK guardrails, LangGraph’s persistence guidance) converges on this stack for any agent trusted to take consequential action.
Durable work order
The request is written to persistent storage before anything runs, so a crash mid-workflow is recovered, not lost (the outbox pattern).
Bounded authority
Every agent operates inside explicit scope boundaries and guardrails that hold “regardless of inputs received” (AWS Agentic AI Lens).
Independent review
No party produces and approves its own work. This is the four-eyes principle from financial controls, release engineering, SOC 2 and ISO 27001.
Human approval gate
Required whenever the next action is irreversible, costly, regulated, or has a high blast radius. Supervised autonomy: fast when safe, slow when it must be.
Tamper-evident audit
Who approved what, when, and under what authority. Independently reviewable, exportable, hash-linked.
Evaluate, canary, roll back
Agent behavior can drift between releases without a code change; CI/CD and SRE practice applied to a stochastic component.
The Affordability Paradox: A Framework

Stating the paradox precisely
I keep naming the same tension. Small businesses "can't afford to keep up" and, at the same time, "can't afford not to." That is not a turn of phrase. It is a real decision-theoretic bind, and it comes from three facts this essay has already put on the table.
First, the engineering discipline you need to run autonomous agents reliably (Section 6) has a cost floor, and that floor does not shrink with company size. A durable workflow engine, an approval service, an audit chain, and an evaluation/rollback pipeline cost roughly the same in engineering hours whether the company has 10,000 employees or 10.
Second, the capability frontier is moving on a measurable, fast exponential. METR's analysis of six years of agent benchmark data finds that the length of tasks AI agents can complete autonomously (at 50% reliability) is doubling roughly every seven months. On SWE-bench Verified specifically, task length is doubling even faster, under three months. If that trend holds for two to four more years, generalist autonomous agents will be capable of week-long autonomous tasks. If it holds to the end of the decade, month-long autonomous projects become feasible (METR).
Third, the cost of the underlying compute is collapsing even faster than capability is rising. One analysis finds frontier-model output pricing fell from roughly $60 per million tokens in 2023 to roughly $15 per million tokens in 2026 (a 4x nominal decline), while quality rose 2–3x over the same period. That works out to an effective cost-per-unit-of-intelligence decline of an estimated 10–15x at the frontier and 30–50x for "good-enough" tier models (pricing trend analysis). a16z's own CIO survey separately reports enterprise LLM costs "coming down by an order of magnitude every 12 months," with one CIO reporting "what I spent in 2023 I now spend in a week" (a16z).
Put those three together and you get the mechanism of the paradox. The engineering cost of reliability is roughly fixed no matter how big the firm is. The capability you get for that fixed engineering cost is compounding on a seven-month doubling clock. The raw compute cost of reaching that capability is falling even faster than that. So the absolute dollar cost of building a competent agent system is falling. But the relative competitive gap between a firm that has the engineering discipline to harness that falling-cost, rising-capability frontier and a firm that does not is widening every single doubling cycle. The firm without the discipline cannot convert cheap tokens and capable models into reliable multi-step outcomes at all. The firm with the discipline turns each new model generation into compounding operating leverage.
The absolute cost of building a competent agent system is falling. The relative gap between the firm that can harness it and the firm that cannot is widening every doubling cycle.
Fixed engineering cost. Compounding capability. Collapsing compute price.
The absolute cost of building a competent agent system is falling. The relative gap between the firm that can convert cheap tokens into reliable outcomes and the firm that cannot widens every cycle.
- What it costs
- A durable workflow engine, an approval service, an audit chain and an evaluation/rollback pipeline cost roughly the same in engineering hours for 10 employees or 10,000.
- What it implies
- The length of tasks agents complete autonomously doubles roughly every seven months (SWE-bench Verified: under three). Two to four more years of trend means week-long autonomous tasks.
- What it implies
- Frontier output pricing ~$60 → ~$15 per million tokens, 2023–2026, while quality rose 2–3×. a16z’s CIOs report costs falling an order of magnitude every 12 months.
A framework: the four quadrants of SMB agent-adoption exposure
| Low engineering discipline | High engineering discipline | |
|---|---|---|
| Low urgency / low agent-exposure business model | Status quo. Safe to wait; the risk is opportunity cost only | Premature over-investment risk; low near-term payoff |
| High urgency / high agent-exposure business model | The paradox trap: falling behind competitors who deploy correctly, but any DIY attempt likely lands in the 95% pilot-failure bucket (MIT NANDA) | The compounding-advantage zone: captures falling compute cost and rising capability as operating leverage every ~7-month doubling cycle (METR) |
Most incumbent SMBs sit in the upper-right-to-lower-left transition. They operate in industries with real agent exposure (professional services, retail, logistics, financial services-adjacent functions), but they lack the engineering discipline described in Section 6. Their realistic options are narrow. They can attempt a DIY build, which very likely lands in MIT NANDA's 95% failure bucket, given that even well-resourced enterprises with dedicated AI teams cluster there. They can buy a narrow, vendor-engineered point solution (Klarna-style customer-service agents, QuickBooks-embedded bookkeeping AI) that captures Level 1–2 value but not executive-function orchestration. Or they can wait, and accept that the compounding-advantage zone is being seized by AI-native competitors and better-resourced incumbents in the meantime.
Most incumbent small businesses sit in the trap: real exposure, no engineering discipline.
Urgency of the business model against the engineering discipline available to it.
The paradox trap
Falling behind competitors who deploy correctly, while any DIY attempt most likely lands in the 95% pilot-failure bucket.
The compounding-advantage zone
Captures falling compute cost and rising capability as operating leverage every ~7-month doubling cycle.
Status quo
Safe to wait. The risk is opportunity cost only.
Premature over-investment
Low near-term payoff; discipline without a use for it.
Quantifying the cost of inaction
IBM's CEO data gives a direct behavioral signal that executives themselves see this dynamic. 64% of CEOs acknowledge that "the risk of falling behind drives investment in some technologies before the organization clearly understands their value," and 37% explicitly say it is better to be "fast and wrong" than "right and slow" on technology adoption. That is a rational response to a compounding-advantage dynamic. It also explains why 50% of the same CEOs admit their rushed investment left them with disconnected, piecemeal technology (IBM IBV).
PwC's Global AI Jobs Barometer supplies the productivity-gap side of the ledger with unusual precision. Labor productivity growth in industries most exposed to AI nearly quadrupled, rising from 7% (2018–2022) to 27% (2018–2024). Productivity growth in the least AI-exposed industries actually declined slightly, from 10% to 9% over the same comparison. Revenue per employee grew three times faster in AI-exposed industries (PwC). This is the empirical core of "can't afford not to." The gap between AI-exposed and non-exposed sectors is not stable. It is actively widening. An SMB in an AI-exposed sector that does not close its own capability gap is, by this data, on a path of relative margin and productivity erosion, whether or not it "feels" behind today.
Quantifying the cost of premature or undisciplined adoption
The cost-of-action side of the ledger is just as quantifiable and just as sobering. Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027 due to "escalating costs, unclear business value, or inadequate risk controls." That is not a prediction about AI capability. In Gartner's own framing, and echoed by independent analysis, it is a prediction about governance and infrastructure readiness. That means it applies with disproportionate force to organizations attempting agent deployment without the discipline described in Section 6 (Gartner).
Real-world SMB implementation costs bear this out directly. Managed-service agent setup for a small business runs $3,000–$12,000 plus $200–$600 a month ongoing, and fully custom development starts at $15,000 and climbs. Those costs dwarf the $100–$500/month subscription price quoted on any vendor's homepage. They also do not include the ongoing cost of monitoring, evaluation, and model-spend management once the system is live (Advantech IT). For a business exploring genuinely cross-functional, executive-level orchestration rather than a single narrow workflow, the applicable cost comparison is closer to Harvey's median $175,000 annual enterprise legal-AI contract or Sierra's undisclosed but reportedly services-heavy enterprise pricing than to any SMB-tier subscription. No vendor sells the executive-team pattern at SMB price points, for the architectural reasons detailed in Section 6.
The resolution: agents-as-a-service as the arbitrage
Here is the practical resolution of the framework, and the commercial opening for infrastructure operators. The fixed engineering cost of reliability can be amortized across many SMB customers by a managed-service or platform provider in a way no individual SMB can amortize alone. This is precisely the "agents as a service" pattern Sierra itself has begun articulating publicly. It is also structurally identical to how SMBs historically got access to capital-intensive infrastructure they could never build alone. Those historical analogs are next.
Historical Analogs: What SMB Technology Adoption Curves Actually Teach
Cloud computing
Cloud computing is perhaps the closest structural analog to the current agentic-AI moment. It was a capital-intensive, engineering-heavy infrastructure capability that became accessible to SMBs only once vendors abstracted away the underlying complexity into a rentable service. Industry timelines of SMB cloud adoption show a steady climb through the early 2010s: a reported 71% year-over-year rise in SMB cloud adoption in 2012 and nearly 90% of SMBs using the cloud by 2015. Mainstream small-business use arrived most of a decade after Amazon's 2006 launch of EC2 (iCorps cloud timeline). The lesson that applies directly to agentic AI: infrastructure that is prohibitively complex to self-build becomes SMB-accessible only through a managed abstraction layer. The businesses that waited for that abstraction layer to mature, rather than trying to build cloud infrastructure themselves in 2008, generally did not suffer competitively for the wait. Businesses in genuinely cloud-native, digitally exposed sectors that waited too long did lose ground to faster-moving, cloud-native competitors.
SaaS
SaaS adoption followed a similar but faster curve, sped up by the 2008 financial crisis pushing SMBs toward lower-capex software models. A 2014 North Bridge/Gigaom survey of 1,358 respondents reported SaaS adoption rising more than fivefold, from 13% in 2011 to 74% in 2014, as the category matured from Forrester's 2008 "road bumps" assessment into default infrastructure within roughly a decade (Telecompetitor). The compressed timeline relative to cloud infrastructure itself reflects a general pattern. Each successive wave of technology-enabled infrastructure diffuses to SMBs faster than the previous wave, because vendors learn from the prior cycle how to package complexity for smaller buyers more quickly.
Mobile and e-commerce
Mobile technology adoption among small businesses reached 66% by 2013 according to Constant Contact survey data (n=1,305). That is a materially shorter climb than either cloud or SaaS. It reflects consumer-technology diffusion dynamics (smartphone ownership) bleeding directly into business practice without requiring dedicated infrastructure investment (Constant Contact). E-commerce adoption in the mid-2000s, tracked by the U.S. Census Bureau across roughly 137,600 businesses through four separate annual surveys, shows the opposite pattern: a slow, multi-year climb, because SMBs had to build entirely new capabilities (payment processing, logistics, digital storefronts) rather than simply adopting a faster distribution channel for an existing capability (U.S. Census Bureau, E-Commerce 2005).
ERP
ERP adoption is the clearest cautionary analog for the executive-agent-team thesis specifically. ERP, like an autonomous executive agent fleet, is a cross-functional, integration-heavy system rather than a point solution. Academic literature on SME ERP adoption documents significant implementation-cost and ROI-realization challenges for smaller firms, which typically lack the dedicated IT departments larger firms bring to the same projects (SSRN, ERP adoption in SMEs). ERP took decades, not years, to reach majority SMB penetration, and a meaningful share of SME ERP projects historically failed or stalled during implementation. That is a direct precedent for Gartner's 40%+ agentic-AI project cancellation forecast and MIT NANDA's 95% pilot-failure statistic. The parallel is instructive. ERP eventually reached SMBs not because SMBs became capable of building integration platforms themselves, but because vendors (NetSuite, and cloud-native successors) packaged the entire integration burden into a rentable, pre-configured product. That is exactly the abstraction-layer dynamic in Section 7.5.
Comparative adoption curve table
| Technology | Approx. years from enterprise availability to majority SMB adoption | Primary failure/friction mode | Resolution mechanism |
|---|---|---|---|
| Cloud computing (AWS-era) | ~8–10 years | Complexity, security concerns, migration cost | Managed cloud platforms, MSPs abstracting infrastructure |
| SaaS | ~6–8 years (accelerated by 2008 crisis) | Integration concerns, "road bumps" per Forrester 2008 (Forrester) | Vendor-side onboarding, freemium tiers, self-serve UX |
| Mobile / smartphone | ~3–5 years | Minimal; consumer tech diffusion did the work | None needed; adoption was pulled by consumer behavior |
| E-commerce | ~10+ years for full SMB penetration | New capability build (payments, logistics, storefronts) | Turnkey platforms (Shopify-era) abstracting the entire stack |
| ERP | Decades; still incomplete for smallest firms | High cost, integration complexity, cross-functional scope, high failure rate | Cloud-native, pre-configured, vertical-specific ERP products |
| Agentic AI / autonomous agents (current) | Unresolved; capability doubling ~7 months (METR), cost falling 4x–50x (pricing analysis) | Engineering discipline gap (Section 6), governance, data readiness | Unresolved; agents-as-a-service is the emerging candidate |
The consistent lesson across every analog is that SMB adoption of cross-functional, integration-heavy technology has never happened primarily through SMBs building internal capability. It has happened through a vendor or managed-service layer absorbing the complexity and renting out the outcome. ERP's decades-long, still-incomplete SMB penetration is the most sobering precedent for executive-agent-team adoption specifically, precisely because it shares the cross-functional integration burden that makes Level 4 agent orchestration categorically harder than the Level 1–2 point solutions (chatbots, content generators) that dominate current SMB AI use. The critical variable that could compress agentic AI's SMB timeline relative to ERP's decades-long curve is the unprecedented pace of the underlying capability and cost curves (Section 8.6). No prior technology wave combined this degree of cross-functional integration complexity with this rate of underlying capability improvement.
Cross-functional, integration-heavy technology has never reached small business through DIY. It arrived through an abstraction layer.
Approximate timelines from Section 8. ERP is the sobering precedent for executive-agent teams; the capability and cost curves are the reason this cycle may not follow it.
Why this cycle may not follow the old timeline
Two forces set the agentic AI cycle apart from every historical analog above, and they argue against a simple extrapolation of ERP's multi-decade timeline. First, the capability curve itself is compounding faster than any prior general-purpose business technology. METR's seven-month task-horizon doubling time, if sustained, implies autonomous agent capability roughly doubles within a single fiscal year, compared to the multi-year hardware and software generation cycles that paced cloud, SaaS, and ERP adoption curves (METR). Second, and reinforcing the first, the cost of reaching that capability is falling on a similarly compressed timescale. a16z's CIO survey reports enterprise model costs falling "by an order of magnitude every 12 months," a rate of price decline with no precedent in cloud storage, compute, or SaaS seat pricing history (a16z). Bain's own frontier-versus-diffusion framing captures the resulting asymmetry precisely. Enterprises that scaled Level 1 tools in 2023–2024 already captured 10–25% EBITDA gains before most competitors had meaningfully started, and Bain explicitly frames 2025 as the year "capital, innovation, and deployment velocity are converging" around Levels 2 and 3. That is a compression of the adoption timeline that historical analogs like ERP, which took a full generation to reach majority SMB penetration, simply do not anticipate (Bain).
What Could Tip Adoption
Protocol standardization
Anthropic's Model Context Protocol arrived in late 2024. Anthropic has since donated it to the Linux Foundation's Agentic AI Foundation for stewardship. What it does is standardize how agents connect to external tools and data sources. That is precisely the "tool wiring" integration burden that IBM's data identifies as a top-three barrier (Anthropic MCP announcement; Anthropic MCP foundation donation). Google's Agent2Agent (A2A) protocol takes on the neighboring problem: agents talking to other agents across vendors and frameworks. If it holds, it addresses the kind of vendor lock-in that made past enterprise integration projects (including ERP) so costly (Google Developers Blog, A2A announcement). Gartner's 2026 Hype Cycle explicitly lists MCP and multi-agent orchestration standards among the more than 30 agentic technologies it now formally tracks. That is its own signal. The standardization layer is moving from experimental to infrastructural (Gartner Hype Cycle analysis). Standardization matters more for SMBs than for anyone else, because it is exactly the layer that used to require bespoke, expensive systems-integration work. That is the same work that kept ERP prohibitively costly for small firms for decades.
Lower-code agent builders and vertical SaaS embedding
Section 2 documented the self-serve pricing tier: Lindy, n8n, Zapier. That tier is the "no-code abstraction layer" mechanism that resolved SMB cloud and SaaS adoption historically, now applied to agent building specifically. Vertical SaaS embedding matters just as much. Instead of an SMB adopting a standalone agent platform, agent capability is arriving pre-integrated inside software SMBs already use and trust. Examples are QuickBooks' embedded AI insights, Shopify Magic, and HubSpot's "Agent Hub and Agent Builder," announced as "one place to build and manage AI agents with shared context" directly inside the CRM SMBs already operate (HubSpot). This embedding path sidesteps the integration-complexity barrier almost entirely. The vendor has already done the data plumbing and governance work as part of the core product. It is the same resolution mechanism that let Shopify's turnkey model finally unlock SMB e-commerce adoption, decades after e-commerce infrastructure first became available to enterprises.
Managed agent platforms and "agents-as-a-service"
Sierra frames its offering publicly as "Agents as a Service." Alongside it there is an emerging market of managed-service SMB agent implementers charging $3,000–$12,000 for turnkey setup rather than $15,000+ for custom builds. Together they are the direct SMB analog to the managed cloud service providers (MSPs) that resolved the cloud adoption gap a decade earlier (Sierra; Advantech IT). This is the most probable near-term tipping mechanism for executive-function agent teams specifically. It is the only model I identified that plausibly amortizes the Section 6 governance-stack engineering cost across many SMB customers, instead of each SMB building it once, alone, at full cost.
Regulatory clarity, or the lack of it
The U.S. Chamber found that 65% of small businesses worry about a patchwork of state AI laws, and that less than one-third feel well-prepared to comply. That suggests regulatory clarification could meaningfully reduce a barrier small businesses are currently reporting themselves. Clarification could come through federal preemption, harmonized state frameworks, or clearer safe-harbor provisions for supervised agentic use (U.S. Chamber). It could go the other way. Continued fragmentation across the EU AI Act's SME provisions, Colorado's evolving AI Act, and NYC's algorithmic bias-audit regime raises the fixed compliance cost, and that cost falls hardest on smaller firms. That could push SMBs toward managed and vendor-embedded solutions, where the vendor absorbs compliance risk, rather than DIY builds. That would reinforce the affordability paradox rather than resolve it. Or, put more carefully, it would resolve it through consolidation rather than broad-based capability building.
What This Means
The evidence assembled here supports several conclusions. Each one is load-bearing enough to sharpen a thesis built around the affordability paradox. First, the paradox is real and measurable, not just rhetorical. The reliability math in Section 6.2, the METR capability-doubling data, and the token-cost collapse data together show that the relative competitive gap between disciplined and undisciplined AI adopters compounds every few months, even as the absolute dollar cost of access falls. So "waiting for prices to come down" is not a way out of the paradox. The discipline gap, not the price, is the binding constraint.
Second, no vendor in the SMB market today sells the full executive-agent-team governance stack at SMB-accessible pricing. Every product in this space that operates above Bain's Level 2 capability tier is enterprise-only, quote-gated, and priced in five- or six-figure annual contracts. That leaves a structural vacuum between what SMBs can afford and what genuine cross-functional autonomy requires.
Third, the historical analogs converge on a single resolution mechanism: a managed or embedded abstraction layer that amortizes integration and governance engineering across many customers. "Agents-as-a-service," in the pattern Sierra has begun articulating publicly, is the most plausible version of that mechanism for this cycle specifically.
Fourth, this essay has pointed repeatedly to an evidentiary gap on SMB autonomous-agent adoption. No survey instrument yet distinguishes "used a chatbot" from "deployed an autonomous agent team." That gap is itself strategically significant. It means the market for rigorously measuring, and therefore rigorously selling into, SMB agentic-AI adoption barely exists yet. For an infrastructure operator positioned to define the category on its own terms, that is as much an opportunity as a research limitation.
Finally, the four-era arc runs from prompting to looping to agent engineering to autonomous agents. Most SMBs are currently stuck in the undisciplined looping stage. The arc suggests that the operators who move fastest from that stage directly into governed autonomy, skipping a slow, incremental crawl through ad hoc agent engineering, will capture the compounding-advantage zone in Section 7.2. They will get there before the abstraction-layer vendors most SMBs will eventually rely on have fully matured their own governance stacks.
The homework, and the work
I will end where the research ends and where my own work begins. SAVRN is building governed agent packages for exactly the customers this essay describes: small and mid-sized businesses, and enterprises without an AI engineering bench.
We test every package on our own company before it goes near a customer. Our own front end, back end, books, and sales desk. The design follows the pattern the public literature converges on in Section 6. One named human owner per fleet. A control plane that brokers every agent action. Durable work orders. A separate approval step that no agent can grant itself. An append-only audit chain. Independent review that reports to the owner rather than to the orchestrator that produced the work.
Agents draft and recommend. Humans approve. The platform executes.
What that costs, and what it makes possible for a twenty-person company, is a separate essay. This one was the homework.
Questions this raises
What is the SMB affordability paradox?
Small and mid-sized businesses can’t afford to keep up with agentic AI and can’t afford not to. The engineering cost of deploying agents reliably is roughly fixed regardless of company size, while the capability available at that cost is doubling roughly every seven months and the price of compute is falling even faster. The absolute cost of access is dropping, but the relative gap between disciplined and undisciplined adopters widens every cycle.
How many small businesses actually use AI?
It depends entirely on the question asked. The U.S. Chamber’s 2025 survey finds 58% of small businesses self-identify as using generative AI; Intuit finds 68% of businesses with 100 or fewer employees use AI regularly. The Census Bureau, which asks about AI “in business operations,” finds fewer than 20% of firms with four or fewer employees. No survey at any size isolates autonomous, multi-step agents from single-turn generative tools.
Is the “95% of AI pilots fail” statistic real?
It comes from MIT NANDA’s State of AI in Business 2025 report, based on 52 executive interviews, 153 survey responses, and more than 300 implementation reviews. It found only 5% of custom internal AI tools reach production with measurable financial impact. Critics note it describes custom-built internal tools, not vendor-native features embedded in existing software, which retain far better because they inherit the vendor’s engineering.
What is an autonomous executive agent team?
A coordinated fleet of agents, analogous to CEO, CFO, COO, CTO, CMO, and chief-of-staff functions, that plans, proposes, reviews, and in bounded cases executes cross-functional business decisions under governance rather than step-by-step human scripting. In Bain’s four-level maturity model it is Level 4, multi-agent constellations, which Bain describes as still “on the whiteboard” even for leading enterprises.
Why do multi-step agent workflows fail so often?
Because success compounds. If each step succeeds with probability p and the workflow needs N steps, end-to-end success is p to the power of N. A model that is right 85% of the time on one step completes a five-step workflow about 44% of the time. Even an excellent 95% per-step model degrades to roughly 70–77% across a six- or seven-step executive workflow without engineering intervention.
What does “agent engineering” actually mean?
Treating the model as a powerful but stochastic component inside a deterministic software architecture. Code owns the rails (schema validation, database writes, retries, checkpoints, rollback) and the model reasons and generates strictly within them. Each step is verified by a test or validator rather than by the model’s own judgment, so the chain stops compounding raw model error. Anthropic’s guidance is to start with the simplest workflow that works and add autonomy only when simpler solutions fall short.
What does production-grade governance for agents require?
The public literature converges on a stack: a durable work order written before execution begins, bounded authority for every agent, independent review so no party approves its own work, a human approval gate for irreversible or high-blast-radius actions, a tamper-evident audit trail, and evaluation with canary releases and rollback. None is exotic alone. The barrier is running all of them together, correctly, on a non-deterministic multi-agent system.
Can a small business buy an executive agent team today?
Not as a packaged product. Every platform with SMB-legible, self-serve pricing tops out at retrieval and single-task workflows. Every product operating at cross-system orchestration or executive-function use cases is enterprise-only, quote-gated, and priced in five- or six-figure annual contracts, often with 25-plus seat minimums. Managed setup for a narrow small-business agent runs $3,000–$12,000; custom builds start at $15,000.
What do cloud, SaaS, and ERP adoption teach about this cycle?
Cross-functional, integration-heavy technology has never reached small business through do-it-yourself capability. It arrived through a vendor or managed layer that absorbed the complexity and rented out the outcome: managed cloud, self-serve SaaS, turnkey e-commerce, pre-configured cloud ERP. ERP’s decades-long, still-incomplete SMB penetration is the sobering precedent; the unprecedented capability and cost curves are the reason this cycle may run faster.
What could tip small-business adoption of agent teams?
Four mechanisms: protocol standardization (Anthropic’s Model Context Protocol, Google’s Agent2Agent) that removes bespoke integration work; agent capability embedded in vertical software SMBs already use; managed “agents-as-a-service” providers that amortize the governance stack across many customers; and regulatory clarity that lowers a fixed compliance cost that falls hardest on the smallest firms.
Sources
Standalone sources pageEvery source on one page: grouped, linked, citableEvery figure in this essay is linked inline to its source at the point of use; this list collects them once. Where a source is a secondary write-up (an analyst summary, a pricing aggregator, a practitioner blog) the essay says so in the sentence that cites it. Before publication each link was checked live and each cited number was compared against its source; the corrections that check produced are already reflected in the text.
- McKinsey · mckinsey.com
- Enterprise DNA on Gartner · enterprisedna.co
- MIT NANDA / AI Governance Library · aigl.blog
- Stanford HAI · hai.stanford.edu
- Anthropic Economic Index · anthropic.com
- Forbes · forbes.com
- Menlo Ventures · menlovc.com
- savrn.com/p/ai-integration · savrn.com
- IBM IBV, June 2025 · newsroom.ibm.com
- IBM IBV, May 2025 · newsroom.ibm.com
- Gartner · gartner.com
- Bain Technology Report 2025 · bain.com
- U.S. Census Bureau BTOS · census.gov
- U.S. Chamber Empowering Small Business, 2024 ed. · uschamber.com
- U.S. Chamber Empowering Small Business 2025 · uschamber.com
- Intuit QuickBooks, April 2025 · quickbooks.intuit.com
- Deloitte · deloitte.com
- Digital Applied synthesis of Gartner/Forrester/Deloitte · digitalapplied.com
- Microsoft · microsoft.com
- Salesforce · salesforce.com
- pricing breakdown · techparrot.io
- AWS · aws.amazon.com
- Zapier · zapier.com
- n8n · n8n.io
- Lindy · lindy.ai
- Relevance AI · relevanceai.com
- pricing summary · automationatlas.io
- Sierra pricing analysis · eesel.ai
- Sierra · sierra.ai
- Harvey pricing analysis · costbench.com
- Hebbia · hebbia.com
- Sana · sanalabs.com
- Anthropic, Building Effective Agents · anthropic.com
- Advantech IT · advantechits.com
- Intuit QuickBooks Small Business Index 2025 · quickbooks.intuit.com
- Forbes coverage of Constant Contact · forbes.com
- Marketing AI Institute / SmarterX, State of Marketing AI 2025 · 20757840.fs1.hubspotusercontent-na1.net
- GoDaddy Venture Forward · godaddy.com
- GoDaddy · godaddy.com
- Salesforce SMB Trends · salesforce.com
- BCG AI at Work 2025 · bcg.com
- McKinsey, citing Benioff · mckinsey.com
- Cognition · cognition.com
- OpenAI · openai.com
- Twig case study synthesis · twig.so
- note.com case discussion · note.com
- Simon Willison · simonwillison.net
- AWS Well-Architected Agentic AI Lens · docs.aws.amazon.com
- LangGraph docs · docs.langchain.com
- OpenAI Agents SDK · openai.github.io
- NIST, AI Agent Standards Initiative · nist.gov
- Cloud Security Alliance, NIST AI RMF Agentic Profile · labs.cloudsecurityalliance.org
- StackAI human-in-the-loop pattern guide · stackai.com
- METR · metr.org
- pricing trend analysis · tokenmix.ai
- a16z · a16z.com
- PwC · pwc.com
- iCorps cloud timeline · blog.icorps.com
- Telecompetitor · telecompetitor.com
- Constant Contact · news.constantcontact.com
- U.S. Census Bureau, E-Commerce 2005 · census.gov
- SSRN, ERP adoption in SMEs · papers.ssrn.com
- Forrester · forrester.com
- Anthropic MCP announcement · anthropic.com
- Anthropic MCP foundation donation · anthropic.com
- Google Developers Blog, A2A announcement · developers.googleblog.com
- HubSpot · hubspot.com
