What Every Enterprise Should Ask an AI Vendor

Every answer a vendor gives you has a price attached, and most of them are free. The few that cost the person saying them something real are the only ones carrying information.
Somewhere in the second or third meeting, the evaluation team gets to its list. Does the platform integrate with the identity provider, does it support deployment inside the customer's own environment, can a human approve an action before it executes, how is data isolated between tenants, what happens when the model is wrong. The vendor's answers come back smooth and complete, because they have been given a hundred times and refined against a hundred rooms exactly like this one. Then someone near the end of the table, half out of curiosity, asks whether anyone has stopped using the product, and what they said on the way out. The register of the room changes. The account executive glances at the solutions engineer. Whatever comes next — a name and a story, or a graceful pivot to logo retention rates — is worth more than the previous forty minutes combined, and it is worth more for a reason that has nothing to do with the content of the reply.
The reason is that the first set of questions could be answered at no cost. Saying yes to an integration, to an enterprise deployment model, to human oversight, costs the vendor nothing in the moment; the claim is unfalsifiable until long after the contract is signed, and the salesperson bears none of the consequence when it turns out to have been aspirational. The last question could not be answered without spending something. A churned customer named out loud is a lead handed to the buyer that may end the deal. This asymmetry is the most reliable instrument a buyer has, and almost nobody uses it deliberately, because evaluation processes are built to grade the content of answers rather than to measure what each one cost the person giving it.
The cheapest sentence in the room is "yes, we do that"
Consider what a capability question actually asks a vendor to do. It invites them to describe a state of the world that is expensive for the buyer to verify, unattributable to any individual after the fact, and immediately rewarded with progress toward a close. Under those incentives, a yes is not dishonesty so much as the path of least resistance, and it is available to every supplier in the category regardless of what they have actually built. This is precisely the dynamic behind what Gartner calls "agent washing" — the relabeling of existing chatbots and rule engines as autonomous systems — in the same analysis where the firm predicts that more than forty percent of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear value, and inadequate controls. Relabeling works because the label is free to apply and expensive to test, which means a buyer who scores vendors on the presence of claims is scoring them on something the market has already optimized to zero.
Now consider the opposite kind of statement. A vendor who says that a particular workflow is not built yet, that the last customer in your industry took nine months rather than the six on the slide, that a named account left because the operating model never took hold internally, or that a specific class of decision still requires a person and probably will for another year, is doing something visibly against interest. Every one of those sentences raises the probability that the buyer walks. There is no version of the incentive structure in which they are the easy thing to say. That is exactly what makes them informative: a claim that would have been made regardless of whether it was true tells you nothing about the world, while a claim that would only be made by someone willing to absorb its cost tells you a great deal — about what they have seen, and about what they are optimizing for.
The practical shift this implies is small in mechanics and large in effect. The questions on the evaluation list mostly stay; what changes is that the buyer stops grading the answer and starts estimating its price. When the reply to a question about failure modes is a fluent taxonomy of failure modes with no incident attached to any of them, the price was zero and the answer should be scored accordingly. When a solutions engineer says that in one deployment the system quietly produced confident output from a stale knowledge source for three weeks before anyone caught it, and describes what they changed afterward, they have paid real money to say so, and the buyer has learned something no reference call would have surfaced.
Not every currency is the same, and rehearsed candor is counterfeit
The complication is that cost is relative to position, and a naive version of this test punishes the wrong vendors. A twelve-customer startup asked to name someone who left is being asked to spend a much larger fraction of its total credibility than a large incumbent would spend on the same disclosure, and a large incumbent frequently cannot name anyone at all because its own contracts and counsel forbid it. Treating silence in that second case as evasion misreads the situation. The right move is to look for expenditure in whatever currency the vendor actually holds. A supplier who cannot discuss customers by name can still spend control: letting your engineers run the system unsupervised against your own messy data rather than a curated demo set, handing over the raw logs of what the agents did and where they got it wrong, agreeing to a paid pilot whose success criteria you write and they do not get to renegotiate. Each of those costs them something they would prefer to keep, and each is a substitute for the admission the lawyers blocked.
The mirror-image error is more common and more dangerous, which is mistaking the performance of candor for the real thing. Sales training has caught up to this test, and the market is now full of rehearsed vulnerability — the volunteered admission that the interface is dated, that onboarding documentation lags the product, that the company is not the cheapest option. These land as honesty and cost nothing, because none of them would lose the deal; several of them advance it by making everything said afterward sound more credible. The test for whether an admission is genuine is not how uncomfortable it sounds but whether it plausibly puts the contract at risk. An answer that could reasonably cause a serious buyer to walk away has been paid for. An answer that makes the vendor seem endearingly self-aware has not, and a buyer who cannot tell the difference has simply moved the thing being gamed one level up.
What the expenditure is actually evidence of
It is worth being precise about why this signal predicts anything, because the argument is not that honest vendors are nicer to work with. It is that the willingness to pay for an honest answer is evidence of two specific things, both of which matter enormously after the contract starts. The first is direct operational experience. Nobody can describe a failure mode they have not watched happen; the texture of a real deployment that went sideways — which assumption broke, who noticed, how long the fix took, what the customer's team had to do differently — cannot be generated from a product roadmap. A vendor who has that texture available has run the thing in production under conditions like yours, and a vendor who deflects into abstraction about robustness and guardrails may simply have nothing concrete to deflect from. The second is a statement about time horizon. Spending deal probability today to avoid a mismatch tomorrow is only rational for someone who expects the relationship to last long enough for the mismatch to matter, which is the closest thing a buyer gets to observing, before signing, how a supplier will behave in month nine when something breaks and nobody is watching. The body of deployment writing collected at the autonomous-enterprise category publication keeps circling the same pattern: programs rarely fail at the technical claim that was verified in evaluation, and frequently fail at the organizational reality that was never discussed because raising it would have been awkward for whoever was selling.
None of this exempts the vendors who make the argument, and it should not be read as one. StudioX sells an Enterprise AI Platform in exactly this category, and a buyer who takes the position seriously should apply it there with no softening: press on which deployments of Autonomous AI Workers stalled and what stalled them, on where Human-in-the-Loop turned out to be needed far more often than the operating model implied, on which Enterprise Knowledge sources produced confidently wrong Observations before anyone caught the drift, on which integration that was supposed to be an afternoon of Instant MCP work became a quarter of plumbing. If the answers come back frictionless and flattering, the correct inference is the same one you would draw about any other supplier in the room, and the fact that a vendor articulated the test is not evidence that it passes the test. A framework that resolves into a reason to buy from the party who supplied the framework is a sales asset wearing a lab coat.
The reframe worth carrying into the next vendor meeting is that you are not there to collect answers. You are there to run a price discovery exercise on a single commodity — the truth about how this system behaves in an organization like yours — and the useful measurement is not what each answer said but what it cost the person to say it. Keep score that way for an hour and the room reorganizes itself: the polished vendor with the complete yes column drops toward the bottom, because a complete yes column is what a free answer set looks like, and the one who gave you a name, a date, and a story they would rather you had not heard rises to the top, because that is what an expensive answer set looks like. The purchase you are making is not a set of capabilities. It is a multi-year relationship with people who will either tell you when something is wrong or manage your perception of it, and the only preview of that behavior you will ever get is what they were willing to spend on you before they had your money.
Discussion
No comments yet — start the conversation.