Enterprise DeploymentEnterprise AI PlatformEnterprise KnowledgeupgradedEnterprise Autonomy

Ask Your AI Vendor When, Not Whether

TS
Trevor Solis · Lead AI Engineer, Missions
August 7, 2026

Every AI vendor can say yes to "can it do that." Almost none of them can say when it worked, for whom, and in what form — and the difference between those two answers is the whole of your diligence.

Somewhere in the middle of the second vendor call, a senior operations leader asks the question she has been building toward for twenty minutes. Can the system read an inbound request, pull the relevant history out of the system of record, decide what should happen next, and act on it without a person babysitting each step? The answer comes back immediately and without qualification: yes, absolutely, that is exactly the kind of thing it does. A screen share follows, and for four minutes the thing does precisely that, cleanly, on a case that behaves. Everyone nods. The evaluation team leaves with a positive impression and no new information whatsoever, because the question they asked was one that no vendor in the market has any reason to answer differently, and the demonstration they watched was built by people whose job it was to make sure it worked.

That meeting repeats itself thousands of times a quarter across enterprise software, and it is the single largest source of wasted diligence in AI procurement right now. Not because buyers are careless — the people in that room were experienced, skeptical, and asking about the right underlying capability — but because they were using a question format that cannot discriminate. Capability questions have a structural flaw: they are answerable in the affirmative by anyone with a roadmap, a demo environment, and a reasonable engineering team, which describes every serious vendor and a great many unserious ones. A question that everyone passes is not a test. It is a ritual that feels like a test, and the feeling is the dangerous part, because it lets an evaluation conclude with confidence it has not earned.

Yes is free, and everyone knows the price

It is worth being precise about why "can it do X" collapses so reliably, because the reason is not that vendors are lying. In most cases they genuinely are not. A roadmap item is a real intention, and a person answering yes about something scheduled for the next two quarters is making a statement they believe. A demo is a real artifact too — the code runs, the integration is live, the output on screen was produced by the system and not by a slide. What makes both of them uninformative is that neither one carries any load. The roadmap costs nothing to extend, and it will be reprioritized without anyone calling you. The demo was built against data the vendor chose, in an environment the vendor controls, on a path the vendor has walked a hundred times, and its entire purpose is to be the one case that works. When you ask a question whose honest answer is yes for a prototype, yes for a roadmap, and yes for a production deployment running at a thousand cases a day, you have designed a question that deliberately erases the only distinction you care about.

The market has also learned, quite efficiently, to optimize for exactly this kind of interrogation. Gartner's analysts have gone as far as coining a term for the practice, warning about "agent washing" — existing chatbots, rule engines, and workflow tools relabeled as agents without a change in what they can actually do unattended — in the same research note where the firm predicts that over forty percent of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. Read that prediction alongside the demo you just watched and an uncomfortable implication surfaces. Those cancelled projects were not bought by people who forgot to ask whether the system could do the work. Every one of them asked, and every one of them was told yes, and the yes was true in the narrow sense and worthless in the sense that mattered. The failure was not detected at purchase because the purchase process was not built to detect it.

The information lives in when, in what form, and for whom

Shift the grammar of the question and the same conversation starts producing signal almost immediately. Instead of asking whether the system can handle exceptions, ask when it first handled them in a customer's production environment rather than a test one, and what share of cases it currently completes without a human touching them at that customer. Instead of asking whether it integrates with your systems, ask which specific integration has been running longest, who has been running it, and what broke in the first month. Instead of asking whether a human can stay in the loop, ask what the system does today when it is uncertain, and whether that behavior is configured per workflow or hardcoded, and how a customer changed it last quarter. None of these are cleverer questions. They are the same questions with a time index and an evidence requirement attached, and that small change moves them out of the space where every answer is yes and into the space where answers differ.

What makes this work is that timing and evidence are expensive to fake in conversation. A vendor can extend a roadmap for free, but they cannot invent a customer who has been running a workflow for eighteen months, cannot invent the specific way that customer's data broke their assumptions in week three, and cannot invent the operational detail that only accumulates when software has been surviving contact with a real organization. The texture of a true answer is unmistakable once you are listening for it. Someone describing a deployment that actually exists volunteers the awkward parts without being pushed — the integration that needed a custom adapter, the category of case they still route to a person because the accuracy was not good enough, the three weeks of tuning before the completion rate stopped embarrassing everyone. Someone describing a capability that exists only as intent produces fluent, general, forward-tense language, and produces it comfortably, because there is nothing specific to be uncomfortable about.

This is why the specificity of the answer is itself the measurement, and arguably a better one than the content of the answer. If you ask when a capability first ran in production and you get a quarter, a customer profile, a workload volume, and a story about what went wrong, you have learned that the thing is real regardless of whether the numbers impress you. If you ask the same question and get a reassurance about architecture, an assertion that it is designed for exactly this, or a pivot to the roadmap, you have also learned something definitive, and you learned it without an argument. The vendor did not have to admit anything, and you did not have to catch them out. You simply asked a question that required a fact, and observed whether a fact arrived. The evaluation stops being adversarial and becomes something closer to instrumentation.

Diligence is a claim about the present tense

There is a deeper reason this reframing matters, and it goes beyond protecting yourself from an overenthusiastic sales cycle. The gap between what AI systems can be made to do and what they reliably do unattended, at volume, inside an organization that did not design its data for them, is the central unresolved fact of this technology's enterprise moment. That gap is where the cancelled projects live. The body of work now accumulating around how autonomous operations actually take hold inside real enterprises keeps returning to the same finding: the organizations that succeed are not the ones who bought the most capable-sounding system, but the ones who were honest with themselves about which parts were operational today and which were aspirational, and who sequenced their rollout accordingly. Timing questions are how you get that honesty into the room before the contract rather than after the pilot.

It is fair to point out that this standard is inconvenient for vendors, including the one I work on. StudioX describes a platform of autonomous AI workers running missions across enterprise systems, with specialist agents, a reasoning core, and human-in-the-loop gates, and every one of those phrases is exactly the kind of language that a capability question would wave through and a timing question would put under pressure. That is the correct outcome. A buyer who asks us when a given mission type first ran in a customer's production environment, how many organizations run it now, and what the human intervention rate looks like in month six versus month one is asking questions we should be able to answer with dates and specifics, and if we cannot, the honest reading is that the capability is younger than the vocabulary describing it. The argument in this essay is not a reason to prefer any particular vendor. It holds identically if you buy from someone else, and it is most useful precisely when applied to whoever you are currently most excited about.

The mental model worth carrying out of this is that a vendor conversation is not a test of the product at all. It is a test of the distance between the product and its description, and capability questions are blind to distance by construction because they only ask whether the description is conceivable. Time is the instrument that makes distance visible. Ask when, ask for whom, ask in what form it runs today, and the same sentence that sounded like a specification a moment ago resolves into one of three very different things — something operating, something being built, or something being imagined — and you will find that you can tell them apart with a fluency that no amount of demo-watching would ever have given you.

Discussion

No comments yet — start the conversation.

Join the discussion

See StudioX run.

Put autonomous AI workers to work on your own systems and knowledge.