You Aren't Choosing an AI Tool. You're Choosing Who Gets Paged at 2 AM.
Part of AI Leadership in the Real World — how leaders turn scattered pilots into governed, adopted, measurable capability. TLDR: A support agent doing 50,000 chats a month needs ~3.5 FTE and $500k+/year just to stay accurate — while a typical 100-seat Copilot rollout sees only 20-30 seats used weekly. For SMBs, build-vs-buy isn't about features. It's about what you can afford to own for 24 months. We thought we were choosing a tool. We were really choosing a future dependency, a support queue, a governance burden, and a second bill that arrives a year later. Every vendor demo promised acceleration, control, and simplicity at once. Every internal proposal promised flexibility, ownership, and leverage. Nobody said both bills arrive late — one in engineering on-call, the other in consumption meters. Good platform decisions feel a little boring at first and very smart a year later. Why AI is special (and why old build-vs-buy math breaks) Traditional software mostly stays still when you leave it alone. AI doesn't: It drifts. Knowledge changes, customer language shifts, users ask harder questions once they trust it. Accuracy quietly drops from 90% to 70% with no error log. It speaks for you — legally. A wrong Confluence page is embarrassing. A wrong chatbot answer is a commitment a tribunal can enforce. It lives on someone else's deprecation clock. OpenAI gives at least 6 months before retiring a GA model. That's a hard deadline, not a backlog item. Prompts, evals, and output parsers all need rework. It multiplies cost per request. One human click = one action. One agent resolution = 6 lookups, drafts, updates, and logs — each potentially metered. It turns connectors into permanent work. Salesforce, SharePoint, Jira, Zendesk all change auth, rate limits, and APIs. Your agent keeps running while its knowledge goes stale. Gartner predicts 40%+ of agentic AI projects will be canceled by end of 2027 on cost, unclear value, and weak risk controls. McKinsey's State of AI 2025 (Nov 5, 2025) found 88% report regular AI use in at least one function, yet only 39% report any enterprise EBIT impact (most under 5%; ~6% high performers). BCG's Widening AI Value Gap (2025, n=1,250) found only 5% generating value at scale while 60% see minimal gains. The gap isn't building. It's owning. Here are 6 public cases every SMB CTO should know — scenario, what happened, what didn't work, and why. 1. The chatbot that created legal liability: Air Canada (2024) Scenario: Jake Moffatt, booking emergency travel after his grandmother died, asks Air Canada's website chatbot about bereavement fares. The bot says: book full fare now, claim the discount within 90 days. What happened: That was wrong. The real policy barred retroactive claims. Moffatt flew, applied with screenshots and a death certificate, was denied. He took it to the British Columbia Civil Resolution Tribunal — Moffatt v. Air Canada, 2024 BCCRT 149 (Feb 14, 2024). The tribunal ordered Air Canada to pay C$812.02 (C$650.88 fare difference + interest + fees) for negligent misrepresentation. It explicitly rejected Air Canada's "remarkable" defense that the chatbot was a separate entity responsible for its own actions: "it is still just a part of Air Canada's website." What didn't work and why: No grounding to the actual policy page, no policy-layer guardrail, no correction path. The bot hallucinated a generous policy linked to the real bereavement page that contradicted it. SMB lesson: If you ship a customer-facing agent — bought or built — you own what it says. For a 20-person company, one such incident is a support crisis and a trust crisis. Budget the eval + human escalation before launch, not after. 2. The "AI platform" that was mostly services: Builder.ai (2025) Scenario: Builder.ai, founded 2016 as Engineer.ai, marketed "Natasha" as AI that builds software "as easy as ordering pizza." Raised $450M+, valued at ~$1.5B in 2023 with Microsoft and Qatar Investment Authority backing. What happened: On May 20, 2025 it entered insolvency. Audits revised claimed 2024 revenue of $220M down to ~$55M. Creditor Viola Credit seized ~$37-40M after a $50M debt facility. The US Attorney's Office for SDNY subpoenaed records. UK entity Engineer.ai Global Ltd went into compulsory liquidation. Viral coverage said "700 engineers faked the AI." That shorthand is contested — Gergely Orosz (Pragmatic Engineer), after talking to ex-employees, corrected it: there was a real AI team doing spec, estimation, and code workflows, alongside hundreds of outsourced developers doing delivery. WSJ had flagged heavy human reliance as early as 2019. What didn't work and why: The hybrid model can be legitimate. The failure was marketing opacity + financial misrepresentation: buyers couldn't tell which part was automated, which was human, and what would happen if the vendor collapsed. Customers who outsourced their roadmap to it lost the platform overnight. SMB lesson: When buying an "AI platform," diligence the labor model: what % is model vs. reusable components vs. humans? What happens to your code, data, and uptime if they go under? Get escrow and export in writing. 3. The seat license nobody sat in: Microsoft 365 Copilot Scenario: A 100-person SMB buys 100 Copilot for Microsoft 365 seats at $30/user/month ($36,000/year) — plus the required underlying M365 license upgrade many teams forget to price. List price: Microsoft lists Microsoft 365 Copilot at $30/user/month (annual commitment) as an add-on requiring a qualifying M365 base plan — true all-in $42-$90/user/month depending on base tier. Copilot Studio is separate consumption: ~$200/mo per 25,000-credit pack. What breaks: Independent utilization surveys consistently report only 20-30% of purchased seats see weekly active use at scale (directional, not audited). Change management (training, comms, support) is repeatedly estimated at 30-50% of license cost and rarely in the proposal. Separate cautionary tale (consumer, not enterprise): Australia's ACCC filed Federal Court proceedings Oct 27, 2025 alleging Microsoft misled ~2.7M Personal/Family subscribers by bundling Copilot rises ($109→$159 Personal, $139→$179 Family) while hiding the cheaper Classic no-AI option in the cancellation flow. The widely cited 100%+ ROI figures trace to a Microsoft-commissioned Forrester TEI study — real data, not independent data. SMB scale-down: At 25 seats instead of 100, the waste is smaller in dollars but identical in ratio — 5-8 active seats carry the other 17. That's why seats must be earned, not allocated. What didn't work and why: Per-seat pricing assumes uniform adoption. In SMBs adoption is spiky: 15 power users love it, 60 try twice, 25 never enable it. You pay for 100, get value from 25. The prerequisite upgrade + idle seats kills ROI by month four. SMB lesson that worked: Pilot with a 25-seat power cohort, instrument weekly active use from day one, and only expand when active-use >60% for 4 weeks. Disciplined teams treat seats as earned, not allocated. 4. The second meter: Salesforce Agentforce ($2/conversation) Scenario: A mid-size team handling 50,000 service interactions a month turns on Agentforce. List price: $2 per conversation — on top of existing Service Cloud seats. List price: $2 per conversation at list, before discounts — and before the Service Cloud seats underneath. Salesforce's current official model is Flex Credits ($500/100k; 1 standard action = 20 credits = $0.10) as the flexible alternative, with unused credits not rolling over and Flex vs. Conversations not mixable in one org. What breaks: Conversation definitions (failed/escalated/abandoned handling) and per-workflow action counts are contract-specific — which is exactly why pre-sign modeling matters. Internal IT/HR agents with lots of short queries get the worst unit economics on per-conversation pricing — same $2 whether the answer saved $200 or $2. SMB scale-down: At 5,000 interactions/mo instead of 50,000, the meter drops 10x (~$10k/mo at list) — but the modeling work doesn't. You still must define the billable unit in writing before signing. What didn't work and why — usage amplification: A human resolving a ticket = 1 action. An agent resolving it = query record + search KB + draft response + update case + send email + log interaction = 6 metered steps, depending on config. Teams estimated on user requests but were billed on agent actions. Copilot Studio has the same pattern: $200/mo for 25,000 credits, but generative actions burn credits faster than classic flows, plus separate Azure token + compute charges. Three bills, three consoles. SMB lesson that worked: Demand a cost-modeling pilot in writing: define "conversation/credit" for your workflow, exclude tests, cap overages, negotiate rollover + 24-month price lock. Design agents to cache lookups and confirm intent before chaining actions — teams report 20-40% consumption cuts from flow design alone. 5. The build that became a hidden product: the 3.5-FTE support agent Scenario (composite SMB, modeled on SearchUnify's 50k/mo reference — not a single named company): A 30-person SaaS company builds a RAG support agent on LangGraph + Confluence + Zendesk. Pilot hits 88% answer accuracy. CEO calls it "done." What happened (per SearchUnify's 2026 vendor field analysis — directional, not audited): To keep a 50k-interaction/mo agent production-grade you need roughly: 1.0 AI/ML engineer + 1.0 data engineer + 0.5 platform + 0.5 security/governance + 0.5 product analyst = 3.5 FTE, ~$500k-700k/year before infra, tokens, observability, and audits. Major model migrations take 6-10 weeks each (re-benchmark, re-prompt, re-test, re-secure). Connectors drift: auth changes, endpoints retire, docs grow by thousands of pages. Three drifts compound: knowledge drift (docs change), data drift (tickets change), behavioral drift (users ask harder questions). Without evals, you find out from angry customers. SMB scale-down: At 5k interactions/mo the token/infra meter drops ~10x — but the governance floor doesn't. You still need someone to own drift review, connector fixes, and the model-deprecation deadline. That's the firewall that pages you at 2 AM when the model retires or the connector breaks. What didn't work and why: The team budgeted 8 weeks to build, zero heads to operate. Two backend devs now spend Fridays fixing retrieval, tuning prompts, and patching connectors — the "temporary" bot became an unstaffed platform product. Software maintenance is 80-90% of lifecycle cost historically; AI maintenance runs 2-3x the initial build in the first few years because drift + vendor deprecation never pause. SMB lesson: Don't ask "can we build it?" Ask "can we staff 0.5-1.0 FTE for 24 months to own it — including who gets paged?" If not, buy the boring platform and own only prompts + evals. 6. What did work: narrow scope + outcome pricing (Intercom Fin + BCG 5%) Scenario: Same SMB instead picks one high-volume, well-logged workflow — e.g., "refund-status + plan-change" tickets — and buys Intercom Fin at ~$0.99 per resolved conversation. What happened: Intercom's official Fin meter is $0.99 per resolution (standalone available with a 50-resolution minimum, no seats required) — you pay only when Fin resolves, not per query or token. Why this pattern works for SMBs: price maps to value (resolution, not login), the team can calculate cost-per-resolution vs. human cost, and scope stays narrow enough to govern. Industry studies widely report the same shape even where exact numbers vary by survey: BCG reported only ~5% generate substantial AI value at scale with "deep and narrow" pilots outperforming sprawl, McKinsey reported high AI usage but far lower EBIT impact, and mature portfolios skew heavily buy (commodity) with selective build (differentiating). Treat those survey percentages as directional, not audited — the mechanics (narrow scope + outcome meter) are the reliable part. Why it worked: Narrow blast radius, verifiable outcome, and a vendor whose meter matches your P&L. The team kept ownership of KB quality + weekly refusal/deflection review — 2 hours/week, not 2 FTEs. What TCO actually means for AI (the 4 lines everyone underestimates) The "second bill that arrives a year later" — that's TCO. Most teams price the build and forget the other three lines: Build / purchase — the demo number everyone quotes (Copilot $30/seat, Agentforce $2/conv, or your 8-week sprint). Operate & maintain — 2-3x the build in the first few years (drift detection, connector fixes, model migrations at 6-10 weeks each, prompt rework). Governance & change management — 30-50% of license cost (training, comms, support) or 0.25-0.5 FTE for evals, security review, incident response. Meter / exit reserve — the second meter (overages, usage amplification, repricing on renewal) + the lock-in reserve (what it costs to leave: data export, re-integration, M&A repricing). Score the rubric's "24-mo TCO" row on the sum of these four lines, not the first one. Here's a worked example for a 25-seat SMB handling 5,000 tickets/month: Line (24 months) Buy (Copilot + Agentforce/Fin) Build (LangGraph + RAG) Purchase / build 25 seats × $30 × 24 = $18k + Fin 5k × $0.99 × 24 = $119k → $137k 8-week sprint ~$40k one-time → $40k Operate & maintain vendor-absorbed → $0 2.5x build = $100k → $100k Governance / change mgmt 40% of license = $55k + 0.25 FTE evals $30k → $85k 0.5 FTE × $150k × 2 = $150k Meter / exit reserve overages + repricing reserve → $20k lock-in: data export + re-integration → $15k 24-mo total $242k $305k Read honestly: in this scenario the build runs ~26% more over 24 months and hands you the on-call burden. Swap your seat counts, ticket volumes, and rates — the shape stays. Build wins only when the workflow differentiates you and you can staff the operate/govern lines. An SMB CTO rubric (steal this for your next review) Score each use case 1 (buy) to 5 (build). Default to buy unless 2+ factors score 4-5. Factor 1 = Buy 5 = Build Differentiation Generic across companies (follow-ups, expense coding, support deflection) Encodes how you win (your underwriting, your triage, your data moat) Data sensitivity Public/internal-low-risk, vendor has certs you lack (SOC 2, EU residency) Proprietary + regulated, must stay in your VPC with audit trail you control Change frequency Stable workflow, vendor ships monthly improvements you want free Changes weekly with your product; vendor roadmap would block you Ownership capacity No one to carry 0.5 FTE on-call, evals, connector fixes for 24 mo Named owner + backfill funded for upgrades, drift review, incident response 24-mo TCO Build TCO >2x buy on honest math (build + 2-3x maintenance + infra + governance) Build TCO within 30% of buy and exit/lock-in reserve priced (re-pricing, M&A, seat+usage dual meter) Non-negotiables before any demo: SSO, audit logs, data residency/boundaries, admin kill-switch + human escalation, observability (resolution, hallucination, cost/interaction). If a vendor can't meet these, they're off the shortlist — no matter how good the demo. Run pilots in real workflows, not sandboxes. Document re-evaluation triggers (e.g., "if cost/resolution >$X or active-use
This is a summary aggregated from Dev.to. Read the complete article on the original site:
Read full article at Dev.to