Build vs. Buy AI Agents: A Decision Framework for Technology Leaders
Late last year I sat in on a steering meeting at a logistics company that had, by any reasonable measure, done everything right. They had identified a genuine bottleneck , a customer-service queue where roughly forty per cent of inbound tickets were repetitive shipment-status queries that agents answered by copy-pasting from three internal systems. They had scoped an AI agent to handle those queries end-to-end. They had a budget, an executive sponsor, and a two-quarter runway. And they were completely stuck, not on the technology, but on a single question that had already consumed six weeks of calendar time: should they build the agent themselves or buy one of the platforms three different vendors were circling them about.
The CTO wanted to build. His argument was reasonable on its face , the workflow touched proprietary data, the routing logic was specific to their operation, and he did not want to hand a core customer-facing capability to a vendor who could raise prices or disappear. The Head of Customer Operations wanted to buy. Her argument was equally reasonable , she had a queue that was growing faster than her headcount, and every month spent building was a month the problem got worse. Both of them were right. That was precisely the problem. The meeting kept circling because the two of them were optimising for different things and neither had a shared way to make the trade-off explicit. What should have been a decision had become a standoff.
I have now watched this exact standoff play out enough times, across enough organisations, that I am confident it is not a quirk of personalities. It is structural. The build-versus-buy question for AI agents feels like a technology decision, so it gets handed to technologists, who answer it with technology criteria. But it is not primarily a technology decision. It is a decision about where an organisation wants to hold risk, where it wants to hold control, and how much it is willing to pay , in money and in time , for each. When you frame it that way, the standoff dissolves, because you can finally see that the CTO and the Head of Customer Operations were not disagreeing about facts. They were disagreeing about which risks mattered more, and no one had ever asked them to say so out loud.
Why the Question Feels Harder Than It Is
The reason build-versus-buy has become genuinely difficult with AI agents , harder than it ever was with, say, a CRM or an accounting package , is that the market has collapsed the middle ground that used to make the choice clear. With traditional enterprise software, “build” meant writing an application from scratch and “buy” meant licensing a finished product, and the gap between those two was so vast that most decisions were obvious. AI agents live in a different landscape. You can now assemble a capable agent from foundation models you rent by the token, orchestration frameworks you install for free, and vector stores you spin up in an afternoon. “Build” no longer means building from nothing. And “buy” no longer means buying something finished, because most agent platforms still require you to configure prompts, connect data sources, define tools, and tune behaviour , which is to say, they require you to build inside their walls.
This is the first thing I try to get leadership teams to see clearly, because it reframes everything that follows. The real spectrum is not build versus buy. It is a question of how much of the stack you own and how much you rent, and at which layers. An organisation that “buys” an agent platform but writes all its own prompts, tools, and evaluation logic has, in truth, built most of what matters and bought a runtime. An organisation that “builds” on an open-source framework but leans entirely on a hosted foundation model has bought the single most expensive and most strategically consequential component. The vocabulary of build-versus-buy hides these distinctions, and hidden distinctions are where bad decisions breed.
Here I want to draw a line between two things that get conflated constantly in these discussions: control and ownership. They are not the same. Ownership is about who holds the asset , whose code it is, whose infrastructure it runs on, whose intellectual property the configuration represents. Control is about who can change the agent’s behaviour, how quickly, and without whose permission. You can own something and have little practical control over it, as anyone who has inherited an undocumented internal system knows. And you can control something you do not own, as anyone running a well-configured SaaS platform with a good API knows. The logistics company’s CTO was arguing about ownership. His counterpart was arguing about control , specifically, her ability to change the agent’s behaviour the day a shipping partner changed its status codes. Until you separate these two ideas, the conversation is doomed to talk past itself, because the words people are using do not mean the same thing to each of them.
The second confusion worth naming is between cost and time-to-value. Building almost always looks cheaper on a spreadsheet, because you are only counting the marginal cost of engineering time you have already paid for. What the spreadsheet does not count is the opportunity cost of the queue that keeps growing while you build. In the logistics case, the buy option was roughly three times the annual cost of the build option on paper , and the build option was, in reality, far more expensive, because it deferred value by two quarters during which the ticket backlog would have grown by an amount that dwarfed the licence fee. Cost and time-to-value are different axes, and treating them as one number is how organisations end up building things that were cheap to build and ruinously expensive to have waited for.
The Four-Lens Framework
What I use to break these decisions open , both when I am advising a leadership team and when I am designing an AI operating model from scratch , is a framework I call the Four-Lens Framework for Agent Sourcing. The idea is simple: a build-versus-buy decision about an AI agent should be examined through four independent lenses, each of which answers a different question, and the decision should only be made once all four have been looked through separately. The failure mode I described in that steering meeting comes almost entirely from people looking through one lens each and assuming they are seeing the whole picture.
The first lens is Strategic Proximity. The question here is: how close is this agent to the thing your organisation actually competes on? Not how important it is , importance is a distraction, because everything feels important. The question is differentiation. A shipment-status agent, however valuable, does not differentiate a logistics company from its competitors; every logistics company answers those queries. A pricing-optimisation agent that encodes a decade of hard-won margin intelligence is a different matter entirely. The closer an agent sits to your source of competitive advantage, the stronger the case for building, because you do not want your differentiation living inside a product your competitors can also license. The further it sits from that core, the more comfortable you should be buying, because you gain nothing by owning a commodity capability except the burden of maintaining it.
The second lens is Change Velocity. How often will this agent’s behaviour need to change, and who needs to be able to change it? Some agents operate in stable domains where the logic barely shifts from one year to the next. Others live in environments that churn constantly , regulatory regimes that update, partner integrations that break, product catalogues that turn over monthly. The higher the change velocity, the more it matters that control sits with the people closest to the change, and this cuts in a direction that surprises people: high change velocity often argues for building, because a good internal build puts the tuning controls in your own hands, whereas a bought platform can make you wait for a vendor’s release cycle or, worse, a professional-services engagement to change behaviour you need to change tomorrow. But it can also argue for buying, if the vendor’s platform is specifically designed for fast self-service configuration. The lens does not give you the answer; it forces you to ask whether your control model matches your change model.
The third lens is Capability Distance. This is an honest assessment of the gap between what building well requires and what your organisation can actually do , not aspirationally, but today. Building a production AI agent is not the same as building the demo that impressed everyone in the sprint review. It requires evaluation harnesses, monitoring, fallback design, prompt versioning, data pipelines, and the operational discipline to keep all of it running when the underlying model changes underneath you. I have seen more agent builds fail on capability distance than on any other lens. An organisation looks at a working prototype, concludes that building is easy, and discovers eighteen months later that it has built a fragile artefact no one knows how to maintain. The larger the capability distance, the stronger the case for buying , or, more precisely, for buying the parts you cannot yet build and building only where you have genuine, demonstrated strength.
The fourth lens is Exit Cost, and it is the one most often skipped, because it asks an unglamorous question at the moment of greatest enthusiasm: if this goes wrong, or if something better comes along, how hard is it to get out? Vendor lock-in with AI agents is subtler than with traditional software, because the lock-in is not only in the data and the integrations but in the accumulated configuration , the prompts, the tuning, the institutional knowledge encoded in a specific platform’s idioms. An agent you have spent a year tuning inside a proprietary platform is expensive to leave even if the licence itself is cheap to cancel. Building does not eliminate exit cost , a bespoke system nobody understands is its own kind of trap , but it changes its character. The discipline the Exit Cost lens imposes is to make the exit path explicit before you commit, so that you are choosing lock-in knowingly rather than discovering it later.
The framework’s real value is not in the four lenses individually , most experienced leaders could name similar considerations if pressed. Its value is that it separates them, and insists they be examined one at a time. The standoff in that steering meeting happened because Strategic Proximity, Change Velocity, and Exit Cost had all been mashed into a single argument about whether to “trust a vendor.” Pulled apart, the arguments stopped competing and started composing.
What Actually Happened
We took the logistics company through the four lenses in a single ninety-minute session, and the picture that emerged was not the one either protagonist had been arguing for. On Strategic Proximity, the shipment-status agent scored low , it was valuable but not differentiating, which weakened the CTO’s build case considerably. On Change Velocity, it scored high, because shipping partners changed status codes and SLAs constantly, which was exactly the Head of Operations’ fear. On Capability Distance, the honest answer was uncomfortable: they had two capable engineers who had built the prototype, but no evaluation harness, no monitoring, and no one who had ever run an agent in production. The distance was large. On Exit Cost, the buy options varied enormously , one vendor’s platform was a walled garden, another exposed everything through open standards.
The decision that fell out of this was a hybrid that neither party had proposed. They bought a platform , but specifically the one that scored well on Exit Cost, exposing its configuration through open, portable formats , and they used it to close the growing queue immediately, satisfying the time-to-value pressure that was genuinely urgent. In parallel, they used the six months of running that bought platform to close their capability distance deliberately: the two engineers built the evaluation harness and monitoring against a real production workload rather than a hypothetical one. The plan, revisited quarterly, was that if the pricing-optimisation agent , which scored high on Strategic Proximity , ever became a priority, they would now have the operational muscle to build it, having practised on something lower-stakes.
What worked was the sequencing. By buying first, they stopped the bleeding and bought themselves the time to build capability rather than pretending they already had it. What did not work, at least not smoothly, was the internal messaging: the CTO initially read “buy” as a verdict on his team’s ability, and it took a deliberate conversation to reframe it as a staged path toward building the things that actually mattered. The framework had resolved the technical decision cleanly; it had not, on its own, resolved the human one. That took a separate effort.
Where This Framework Breaks Down
I want to be honest about the limits, because a framework sold as universal is a framework that will eventually embarrass you. The Four-Lens Framework assumes you have enough clarity about your own strategy to score Strategic Proximity meaningfully, and many organisations do not. If a leadership team cannot say what it competes on, the first lens produces noise, and the whole exercise tilts toward the more measurable lenses , Capability Distance and Exit Cost , which quietly biases every decision toward buying. That bias is not always wrong, but it should be recognised as an artefact of strategic vagueness rather than a genuine conclusion.
The framework is also weakest in fast-moving foundational technology, where the ground shifts under all four lenses simultaneously. An Exit Cost assessment made against today’s agent platforms may be obsolete in a year, because the entire category is immature and consolidating. I have watched organisations make a careful, defensible build decision that was correct on the day and wrong within twelve months, simply because a bought option that did not exist at decision time arrived and made their bespoke work redundant. The lenses discipline the decision; they cannot see the future. This argues strongly for revisiting agent-sourcing decisions on a deliberate cadence rather than treating them as permanent, and for weighting the Exit Cost lens more heavily in immature categories precisely because you expect to change your mind.
The most common failure mode in adoption is subtler than getting the answer wrong. It is using the framework to ratify a decision that has already been made emotionally , the CTO who wants to build and scores the lenses to justify it, the operations leader who wants to buy and does the same. The framework only works if the scoring is done honestly and, ideally, by people who do not have a stake in a particular outcome. When it becomes a theatre of justification, it produces worse decisions than a coin toss, because it dresses a predetermined conclusion in the costume of rigour. If you cannot get an honest scoring, you are better served by naming the political reality directly than by pretending the framework has spoken.
The Real Decision Underneath the Decision
The thing I have come to believe, after watching a great many of these choices play out over the years that follow them, is that build-versus-buy for AI agents is rarely the decision it appears to be. It appears to be a choice between two options. It is almost always a choice about sequencing and capability , about what an organisation should be good at building itself, eventually, and what it should be content to rent forever. The organisations that handle this well are not the ones that always build or always buy. They are the ones that treat each agent as a deliberate bet on where to invest their own competence, buying the commodity capabilities without apology and building, patiently, only where owning the capability compounds into advantage.
The question worth carrying out of any steering meeting like the one I sat in on is therefore not “build or buy?” It is: if we build this, is it because we will be meaningfully better for having built it , or merely because we could? The word “could” has bankrupted more engineering roadmaps than any vendor ever has. A senior technology leader’s job is not to build everything that is buildable. It is to know, precisely and honestly, the difference between the capabilities worth owning and the ones worth renting , and to have the discipline to rent the rest.

