Dev.to · 17 min read

AI Made Prototyping Free. That Is Exactly Why Your Portfolio Strategy Matters Now.

AI Made Prototyping Free. That Is Exactly Why Your Portfolio Strategy Matters Now.

Part of the "AI Leadership in the Real World" series: how leaders turn AI from scattered pilots and executive excitement into governed, adopted, measurable business capability. TLDR: 60 ideas on a board. 4 product teams to build them. $1.2M in annualized run cost for pilots that produced $340K in measurable value. The week we made tradeoffs visible was the week AI stopped being a budget line item and started being a product strategy. My AI backlog looked like a menu with 60 "top priorities" and no kitchen to cook them. Every function had a smart idea. Every idea came with urgency and a sponsor. Customer support wanted a deflection chatbot. Engineering wanted a code review assistant. Sales wanted lead scoring. Operations wanted anomaly detection. HR wanted resume screening. Finance wanted invoice reconciliation. Each one was a good idea. That was the problem. When every idea is good, prioritization becomes political. The loudest pitch keeps winning. People start optimizing for being seen, not for being useful. And the organization quietly trains everyone to be louder. I run a product P&L. I do not have the luxury of treating every good idea as a funded initiative. My job is not to maximize the number of AI pilots. My job is to maximize the return on the engineering capacity, infrastructure budget, and organizational trust I have been entrusted with. Those are finite. Every pilot I approve is a pilot I cannot fund somewhere else. Every dollar of inference cost is a dollar that did not go to a product feature, a reliability improvement, or a person. We were treating AI like a lottery ticket instead of a managed investment. The P&L was telling us that before I was willing to listen. What Changed: The Cost of the First Prototype Collapsed Two years ago, prototyping an AI use case took weeks — data pipeline, model, inference endpoint, UI, deployment path. The cost itself was a prioritization mechanism. Only ideas that survived a viability check got built. That barrier is gone. Today, a prompt, an API key, and an afternoon get you a working demo. A LangGraph agent can be built in an evening. A RAG pipeline can be running by lunch. This is great for product. I can test a hypothesis before writing a quarterly business case. A product manager can answer "would this help our users?" in days, not months. But it also means the filter is gone. Now every idea can get a prototype, every prototype a demo, every demo enough excitement to justify keeping it alive. And every live pilot has a run cost: API calls, cloud infrastructure, developer attention, security review cycles, roadmap slots. The prototype is cheap. The pilot is not. The production system is expensive. The distance between those three stages is where most AI budgets quietly bleed out. HBR Noticed the Same Pattern I started seeing this pattern in my own portfolio before I saw it in print. Then HBR published three pieces in six months that described exactly what I was living through. In November 2025, Goutam Challagalla, Mahwesh Khan, and Fabrice Beaulieu (IMD/BCG) published "Stop Running So Many AI Pilots". Subtitle: "Instead of testing lots of use cases across the company, pick one area and go deep." They used Reckitt as their case study. Reckitt found use cases spanning the business — presentations, customer support, procurement. Each guaranteed time savings. But the executives realized "the effort wouldn't transform the company's strategy or create a meaningful advantage. They were hoping for something more dramatic, not just marginal efficiency improvements." That sentence hit me. We had pilots that worked, saved time, produced decent demos. But they were not changing the product. They were making the same product slightly faster. The same month, Ania Masinter published "Prioritizing AI Investments That Create Real Value" as an HBR Executive Playbook. Her argument: "It's time for companies to move from experimentation to disciplined, focused investment in AI." She referenced the 2025 Wharton-GBK AI Adoption Report on executive pressure for AI ROI. That pressure is not theoretical — I feel it every quarter. Then in January 2026, Faisal Hoque, Erik Nelson, Tom Davenport, and Paul Scade published "Manage Your AI Investments Like a Portfolio". Their framing: "Business leaders now face intense pressure to transform their organizations with AI, even though the technology, public attitudes, and the competitive landscape are all still in flux. The result is often too many pilots with too little coordinated oversight." They cited an IBM study finding "isolated, piecemeal deployments," a Deloitte report on "limited buy-in by senior executives," and McKinsey's State of AI 2025 finding "weak linkage to strategic goals." Three articles. Three author groups. One diagnosis: too many pilots, too little oversight, too little connection to strategy. The Question That Changed How I Think Larry Page, co-founder of Google, famously said: "Put more wood behind fewer arrows." The phrase is sometimes attributed even earlier to Scott McNealy, co-founder of Sun Microsystems. The idea is old and simple: when you have limited strike capability, concentrate resources on fewer bets so each one hits harder. The opposite is the scattergun: build lots of small arrows, fire them all, hope something sticks. In AI, that means launching pilots across every department, demoing them all, declaring whatever survives as "the strategy." It feels productive — your board sees activity, your teams feel empowered. But every live pilot consumes maintenance, infrastructure, security review, and trust. A pilot that does not get retired becomes a permanent tax. And it fragments your roadmap — you cannot ship a coherent product when engineering is maintaining six pilots that were never supposed to become products. The question is not whether you should experiment. You should. The question is whether you experiment to explore — cheap and healthy — or to avoid committing, which is expensive and corrosive. I do not think there is a universal right answer. There is the answer that fits your capacity, your P&L, and your market's window. The Framework That Helped: GIST I did not invent the portfolio model from scratch. I borrowed from GIST — Goals, Ideas, Step-Projects, Tasks — a framework created by Itamar Gilad during his time at Google. The piece that matters for this conversation is the transition between Ideas and Step-Projects. Ideas are a hypothesis bank. Everything goes in. Nothing is committed. The low cost of AI prototyping makes this stage richer than ever — you can test more hypotheses in a week than you used to in a quarter. That is the advantage. Step-Projects are the gate. Small, timeboxed experiments with a measurable question and a kill criterion. An idea only becomes a step-project when it has a hypothesis, an owner, and a metric. The prototype proves technical feasibility. The step-project proves it is worth scaling. That transition — from idea to step-project — is where most AI portfolios fail. The prototype is cheap, the demo is impressive, and the idea skips the gate and goes straight to production. GIST forces a different discipline: scaling is earned at the step-project gate, not granted at the demo. I am not claiming we adopted GIST wholesale. We borrowed the principles and adapted them to an organization with monthly planning, security review cycles, and a finance team that wanted ROI narratives. But the core insight — ideas are hypotheses, experiments are gates, scaling is earned — is what made the portfolio work. Two things nobody tells you about the gate There are two problems I see teams hit at the step-project gate that GIST does not fully solve on its own. Both are culture problems dressed up as process problems. How much do you build to test the hypothesis? This is the question I get asked most. Some teams build too little — a thin wrapper around an API call that cannot answer the actual question. Some teams build too much — a full production system with auth, monitoring, and a deployment pipeline, as if the step-project is already the product. And some teams build the whole thing, skip the gate entirely, and ship it. The right amount is the smallest thing that answers the hypothesis with enough confidence to make a scaling decision. Not more, not less. That judgment comes from practice, not from a framework. I have seen teams spend three months building what should have been a two-week step-project, and I have seen teams ship a one-evening prototype to production because it looked good in a demo. Both are failures of the same instinct: the inability to distinguish "enough to learn" from "enough to ship." Killing a bad idea is a success. This is the culture shift. GIST treats a killed step-project as a healthy outcome — you learned the idea was not worth scaling, and you saved the organization months of wasted effort. Most organizations do not work that way. In most cultures, a killed project feels like a failure. The person who proposed it feels embarrassed. The team that built it feels like they wasted their time. The sponsor who championed it feels political loss. So instead of killing, the project gets extended. It gets reframed. It gets "one more quarter." And it quietly becomes a permanent tax on the organization. This is not a process problem. You can have the best kill criteria in the world, and it will not matter if the culture treats a kill as a career risk. The shift has to come from leadership. When a step-project gets killed, the response should be: "Good. We learned something. What is the next hypothesis?" Not: "What went wrong?" The first response builds a culture of evidence. The second builds a culture of self-protection. I made this change deliberately. The first three times we killed a step-project, I said the same thing in the review: "This is the system working. We saved the organization from scaling something that would not have earned its budget." By the sixth time, teams stopped defending dying projects. They started bringing their own kill recommendations. That is when I knew the culture had shifted. What We Tried We shifted from an ideas list to a portfolio model. The goal was to make experimentation deliberate and scaling earned. Every stage had a gate, and every gate had a P&L implication. Four stages, not a backlog I created a simple portfolio board: explore, validate, scale, retire. Explore — sandbox. Small budget, two weeks, no production commitments. You get a slot if you can state the question and the metric that answers it. The prototype is free. The slot is not. Validate — pilot stage. Real users, real data, real success metrics. Required: risk tier, baseline measurement, rollout plan, named product owner. No owner, no pilot. Scale — production. Not everything that passed validation scaled. We scaled what fit the month's strategy and had operational backing. Scaling meant a roadmap commitment, a budget line, an adoption plan. Retire — pilots that stalled, demos without owners, use cases too small to justify maintenance. We made retirement a normal part of the lifecycle, not a failure. A retired pilot freed capacity. That capacity was a product asset. Kill criteria upfront Every use case entered the portfolio with kill criteria defined before it started: No data after 30 days? Killed. No baseline, no decision to scale. No product owner after pilot? Killed. An orphan pilot is a future maintenance burden. No adoption path? Killed. A tool with no users is a demo, not a product. No measurable impact against baseline? Killed. If you cannot prove it worked, it did not work. The criteria were kindness, not punishment. They prevented zombie work and the slow accumulation of run costs nobody audits. Monthly portfolio review as a decision forum Not a status meeting. A decision forum. Finance, Security, and product owners in the room. Moves between buckets required a decision. Kills required a decision. Scale commitments required a decision. By the third review, people came prepared with data instead of defending work that had not produced evidence. The Uncomfortable Moments The first exec-sponsored kill. A VP pushed for a customer-facing support chatbot. Six-week pilot, 12% deflection — decent but not transformative. The kill criteria said no adoption path: the contact center team had not committed to workflow changes. I had to say no to a senior leader in a room full of peers. The P&L reality: extending meant another quarter of API costs, another developer's partial allocation, another security review — all for a use case with no path to production. The cost of saying yes was not zero. It was the cost of every other bet that slot could have funded. Two teams building the same thing. Month two: the portfolio review surfaced that engineering and operations were both building anomaly detection for the same data pipeline. Neither knew. We merged the efforts. The transparency helped — it was not about who was right, it was about paying for the same outcome twice. The quick pilot that consumed three months. A "two-week" invoice reconciliation pilot stretched to three months. The tool worked. The demo was impressive. But finance had not been consulted on workflow integration, the data pipeline needed unscoped changes, and the success metric was never baselined. Three months of developer time and API costs. Zero documented impact. That one was on me — I used to reward activity because it was easier to see than outcomes. The P&L does not reward activity. It rewards outcomes. What I Think the Evidence Shows The surprising result: fewer bets created more momentum, not less. With 15 pilots simultaneously, each got a fraction of the attention it needed — the platform team context-switched across five integrations, security review was a bottleneck, adoption planning was an afterthought. None had enough depth to become a real feature. They were all shallow. With four bets, each got real ownership. The platform team went deep on one integration. Security review happened early. Adoption planning started before launch. The four survivors had enough depth to become products — not demos, but features with users, metrics, and a roadmap. The Reckitt case followed the same pattern. They picked one area and went deep. Their executives said they "were hoping for something more dramatic, not just marginal efficiency improvements." That instinct — to look for the bet that changes the product, not just shaves hours off a process — is the real work of a product leader. AI did not create that instinct. It made the cost of ignoring it visible faster. What I Would Do Differently Start the rubric simpler. Our first scoring had twelve dimensions. Teams spent more time gaming the rubric than building. We cut it to four: value, feasibility, risk, adoption readiness. That was enough. Invest earlier in baselines. "Impact" is a debate without a baseline. Without one, every product decision is a story. With one, it is a number. Finance trusts numbers. Invite Finance and Security in from day one. We brought them in at month four. Every late-stage challenge could have been an early-stage conversation. Finance in the room makes the cost of yes visible. Track fully loaded cost from day one. API costs, cloud, developer time, security review, PM attention. The prototype is cheap. The pilot is not. Know the difference before you scale. Create a small fast-experiments fund outside the portfolio. Not every idea needs the formal process. The rule: fast experiments cannot enter validate without the portfolio review. Exploration should be cheap. Scaling should not be accidental. What I Am Still Uncertain About I do not have a clean ROI number for the portfolio approach. I can tell you we retired 14 pilots in six months, merged 6 duplicates, and avoided scaling two use cases that would have failed in production (one with a projected $180K annual run cost). But I cannot give you a single number that proves the portfolio was worth it. The strongest evidence is cultural and financial. Teams stopped bringing half-formed ideas. They started bringing ideas with baselines, owners, and cost estimates. The run cost of our AI portfolio stopped growing without corresponding growth in measurable impact. This approach can feel overly bureaucratic to small teams. If you have three engineers and one AI use case, you do not need a portfolio board. You need to ship. The portfolio earns its keep when demand exceeds capacity and politics distorts prioritization. I am not sure where the exact threshold is. I am also uncertain about the balance between exploration and focus. Too much focus and you miss emergent opportunities. Too much exploration and you never build depth. The Reckitt case suggests going deep wins. But I have seen the scattergun surface an unexpected winner nobody would have prioritized. I do not think there is a formula. There is a discipline — the discipline of making the tradeoff visible and revisiting it regularly. Closing If everything is transformative, nothing is. AI lowered the cost of the first prototype to near zero. That is a gift. But it also removed the natural filter that forced teams to think before they built. The result, as HBR documented across three pieces in six months, is too many pilots with too little oversight and too little ROI. The fix is not to stop prototyping. The fix is to make scaling a deliberate decision — a product decision, a financial decision, a strategic decision. Larry Page was right. Put more wood behind fewer arrows. Not because small experiments do not matter — they do, and the low cost of running them is an advantage you should use aggressively. But because the organization that can choose which bets to concentrate on is the one that turns AI from a backlog of demos into a product strategy that ships, earns its budget, and shows up on the P&L as something other than a cost center. I do not think this is the only answer. I think it is the answer that fit my organization, my P&L, and my moment. Yours may be different. But the question is the same. Your AI backlog is not a strategy. It is a P&L liability until you prove otherwise. Three Questions for You Are you building many small arrows or putting more wood behind fewer? What would change if your CFO asked you to justify the run cost of every live pilot — could you? What kill criteria do you use, and do you actually enforce it — or does it bend for the loudest sponsor? Does it bend for your VP? Where has the low cost of AI prototyping helped you, and where has it created pilots that are quietly costing you engineering capacity you cannot afford? I genuinely want to hear both sides, because I am still calibrating the balance myself. References Challagalla, G., Khan, M., & Beaulieu, F. (2025). "Stop Running So Many AI Pilots." Harvard Business Review, November–December 2025. hbr.org/2025/11/stop-running-so-many-ai-pilots Masinter, A. W. (2025). "Prioritizing AI Investments That Create Real Value." HBR Executive Playbook, November 2025. hbr.org/2025/11/prioritizing-ai-investments-that-create-real-value Hoque, F., Nelson, E., Davenport, T., & Scade, P. (2026). "Manage Your AI Investments Like a Portfolio." Harvard Business Review, January 2026. hbr.org/2026/01/manage-your-ai-investments-like-a-portfolio Gilad, I. GIST Planning: Goals, Ideas, Step-Projects, Tasks. Framework developed at Google. See: ProductPlan Glossary — GIST Planning. Page, L. (attributed). "Put more wood behind fewer arrows." Co-founder, Google. The phrase is also attributed to Scott McNealy, co-founder of Sun Microsystems. See: English Stack Exchange discussion. IBM Institute for Business Value. "CEOs Double Down on AI While Navigating Enterprise Hurdles" (May 2025). newsroom.ibm.com. Key finding: 50% of surveyed CEOs report that rapid investment has resulted in disconnected, piecemeal technology within their organization. Wharton-GBK AI Adoption Report (2025). Referenced in Masinter (2025) regarding executive pressure for AI ROI. ai.wharton.upenn.edu.

This is a summary aggregated from Dev.to. Read the complete article on the original site:

Read full article at Dev.to

More Startup & VC News