24 Questions to Ask an AI Agent Vendor Before You Sign
Before you sign with an AI agent vendor, get written answers to 24 questions across six areas: scope and fit, proof, data and security, architecture, commercials and ownership, and life after launch. Five of them decide most deals: the fixed price on the first workflow, the written success metric, the agreed kill threshold, who owns the prompts and code if you part ways, and what the agent costs to run each month at your real volume. A vague answer on any of those five is the signal.
Gartner estimates that only about 130 of the thousands of vendors selling agentic AI are real, with the rest engaged in what it calls agent washing: rebranding chatbots, assistants, and RPA scripts as agents. Same firm, same release: over 40 percent of agentic AI projects will be canceled by the end of 2027, driven by escalating costs, unclear business value, and inadequate risk controls.
Those three failure modes are all detectable in a sales call. Below are the questions that detect them. Print the list, ask them in order, and require the answers in writing.
Key takeaways
- Ask for a fixed price on the first workflow. A vendor who will not price one bounded job is telling you they cannot scope one.
- Require a written success metric and an agreed kill threshold before any build starts. Gartner’s cancellation drivers are cost, unclear value, and weak risk controls; all three are contract problems, not technology problems.
- Ownership of prompts, code, and configuration is the new lock-in fight. Get it answered in writing or you will pay to leave.
- Ask what the agent costs to run per month at your real volume, not per seat. Model usage plus monitoring is the line most quotes omit.
- Only 21 percent of organizations report a mature governance model for agentic AI, so ask who approves what, and where the audit trail lives.
Section 1: Scope and fit
1. Which of our workflows should we NOT automate yet, and why? The single most revealing question on this list. A vendor selling capacity will automate anything. A vendor selling judgment will name two or three workflows that are too ambiguous, too low-volume, or too dependent on data you have not cleaned up.
2. What is the one bounded workflow you would start with, and what does done look like? You want a single job with a clear beginning and end, not a platform. If the answer is a roadmap, the scope is not defined.
3. What is the success metric, how is it measured, and who measures it? Hours returned, error rate, cycle time, cost per transaction. Pick one primary. If the vendor measures it themselves with no independent read, the metric is decoration.
4. What is the kill threshold, and who can call it? The number below which the pilot stops and nobody argues. Named person, written number, agreed before the build. A vendor who resists a kill threshold is telling you about their confidence.
Section 2: Proof and track record
5. Show me an agent you run in your own business, in production, today. Building agents for other people while running none of your own is a warning. Ask what it does, what it broke, and how they found out.
6. What is the longest an agent you built has been running in production? Demos are easy. Month nine is hard, because by then three connected tools have shipped breaking changes.
7. Walk me through an agent build that failed and what you changed afterward. Anyone with real deployments has one. A vendor with no failure story has no production history or no candor, and both cost you.
8. Who exactly will build this, and are they the people in this room? Ask for names, and ask whether the work is subcontracted. If the build team is different from the sales team, get the build team on the next call.
Section 3: Data, security, and compliance
9. Do the models you use train on our inputs and outputs? The answer must be no, and it must be traceable to the vendor’s commercial terms with the model provider, not to a verbal reassurance.
10. Where does our data go, which subprocessors touch it, and can we see the list? A vendor who cannot produce a subprocessor list has not thought about your procurement team.
11. Do you have SOC 2, ISO 27001, or ISO 42001, and if not, what do you actually operate? Certification is a real gate for many enterprises. But an honest description of operating controls beats a certification claim that turns out to be Type I from two years ago, or “in progress” with no auditor engaged. Ask which, ask for the date, ask for the report.
12. Will you sign a DPA, and how fast can you return a completed security questionnaire? Speed here predicts everything else. The vendors who answer in days have done it before.
Section 4: Architecture and integrations
13. Which of our systems will the agent touch, and what permission does it need in each? Least privilege, per system, in writing. Integrations, not the AI, are where most of the budget and most of the breakage live.
14. What happens when one of those tools changes its API? You want a named owner, a monitoring approach, and an answer about who pays. “We would fix it” is not a support model.
15. Where is the human in the loop, and which actions can the agent take without approval? Draw the line explicitly. Anything irreversible, anything that spends money, anything that reaches a customer should require a person by default.
16. What is the audit trail, and can we read it without asking you? Deloitte’s 2026 survey of 3,235 IT and business leaders found only 21 percent have a mature governance model for agentic AI, and the named gaps are exactly this: decision boundaries, real-time monitoring, and audit trails. Do not accept a black box.
Section 5: Commercials and ownership
17. What is the fixed price for the first workflow? Not a range, not a discovery fee that becomes a range. One number for one bounded job. Our published AI agent cost bands put a competently built single-workflow agent at $3,000 to $8,000, so you have a reference point for whether an answer is serious.
18. What are the monthly model and monitoring costs at our real volume? Ask them to compute it from your transaction count. Budget 15 to 25 percent of build cost per year for the running of it. A quote with no ongoing line item is incomplete, not cheap.
19. Who owns the prompts, the code, the configuration, and the evaluation set if we part ways? This is the new website-hostage problem. The answer you want is that you own all of it and can export it. Get it in the contract.
20. What does it cost to add the second and third workflow, and is that price fixed now? Scope creep is priced at the moment of maximum leverage, which is after the first one works. Fix it early.
Section 6: After launch
21. What does “maintained” include, in a list? Monitoring, incident response time, model updates, integration repairs, prompt tuning, reporting. If it is not enumerated, it is not included.
22. How will you hand this over if we want to run it ourselves? Documentation, runbook, access, a training session. Ask what that costs. A vendor confident in their value does not need to hold the keys.
23. What is your response time when the agent does something wrong in production? Not uptime. Wrongness. Agents fail by being confidently incorrect, which no uptime SLA catches.
24. What would make you tell us to stop? The last question, and the one that tells you whether you are buying judgment or capacity.
Three answers that should end the conversation
- “It depends” to question 17, twice. Serious vendors price a bounded workflow.
- “We own the prompts” to question 19. You are buying a dependency, not a system.
- A certification claim that cannot produce a report and a date at question 11. If they will bend the truth in the sales call, the audit trail will not save you.
How we answer these ourselves
Fair is fair, so here are our answers to the four sharpest ones.
We run our own agent fleet in production every day, and it reclaims 50 or more hours a week across the team. We also operate a public MCP server at mcp.strataigize.com, which means an AI agent can query our business directly, and we have not found another agency that runs one.
On security and compliance, we do not have SOC 2, and we say so on our trust and security page rather than in a footnote. That page carries what we do operate: least-privilege client-granted access, credentials in managed stores, a governance summary aligned to the four functions of the NIST AI Risk Management Framework, our subprocessor list, and a DPA available on request. We complete vendor security questionnaires when asked.
On price, we quote a fixed number for the first workflow and we agree the kill threshold in writing before the build starts. On ownership, you own the prompts, the code, and the configuration.
FAQs
What is agent washing?
Rebranding an existing product, usually a chatbot, an assistant, or an RPA script, as an autonomous agent without the underlying capability. Gartner estimates only around 130 of the thousands of agentic AI vendors are genuine. Question 5 and question 13 are the fastest tests: ask what they run themselves, and ask what permissions the agent actually holds.
How long should AI agent vendor evaluation take?
Two to four weeks for a bounded first workflow, most of it spent on the security review rather than the demo. Send questions 9 through 12 before the first call so the compliance answers arrive in parallel with the commercial ones.
Should we run a paid pilot or a free proof of concept?
Paid, and bounded. Free pilots get staffed with whoever is available and measured against nothing. See what a pilot that survives procurement looks like for the structure that holds up in a security review.
What if we are still deciding whether to build this internally?
Run the build versus buy comparison first. The vendor questions only matter once you have decided a vendor is in scope.
Bring this list to us
We will answer all 24 on a call, in order, and put the answers in writing. Book a growth audit or read what our AI agent development engagements actually include.