The AI Agent Pilot That Survives Procurement
A pilot that survives procurement has five things agreed in writing before any build starts: one bounded workflow, a single success metric with a named owner who measures it, a kill threshold that stops the project without an argument, a data handling answer your security team has already read, and a fixed price. Run it paid, for 30 to 90 days, on real data. Pilots do not die in the demo. They die in security review, and in the meeting where nobody can say what success was supposed to look like.
Gartner expects over 40 percent of agentic AI projects to be canceled by the end of 2027, and names three causes: escalating costs, unclear business value, and inadequate risk controls. Every one of those is decided before the first line of code, in how the pilot was structured. Here is the structure that avoids all three.
Key takeaways
- Paid, 30 to 90 days, one bounded workflow, real data. Free proofs of concept get staffed with whoever is free and measured against nothing.
- Write the success metric and the kill threshold before the build. The kill threshold is the part that makes finance comfortable and the part vendors resist.
- Answer the security questions in week zero, not week six. Security review is where pilots actually die.
- Name one human owner on each side. Deloitte found only 21 percent of organizations have mature agentic AI governance, and decision boundaries are the top gap.
- End the pilot with a decision, not a demo: expand, extend once, or stop. Write which one before you start.
The five things that must be agreed before you build
1. One bounded workflow
Not a department. Not a platform. One job with a clear trigger, a clear end state, and a countable unit of work. Invoice exceptions routed. Inbound leads qualified and booked. Support tickets triaged and tagged. If you cannot describe the workflow in one sentence with a noun and a verb, it is not bounded yet, and the pilot will expand until it fails.
The test: can you state how many times this workflow ran last month? If the number does not exist, get it before the pilot, because it is also your baseline.
2. A written success metric with a named measurer
One primary metric, one number, one person who reads it. Hours returned per week, percentage of items handled without human touch, cycle time, error rate, or cost per transaction. Pick the one that would justify the spend on its own.
Two rules that matter more than the metric you choose. First, the measurer should be on your side, not the vendor’s; a vendor grading their own homework is not evidence. Second, capture the baseline before the agent goes live. Retrofitted baselines are always flattering and never believed.
3. A kill threshold, agreed in writing
The number below which the pilot ends and nobody relitigates it. If the metric is “60 percent of tickets triaged with no human correction,” the kill threshold might be 35 percent. Miss it at day 60, the engagement stops, and you have spent a bounded amount instead of an unbounded one.
This is the clause that turns an AI experiment into a normal procurement decision, and it is the one that separates vendors. A vendor who resists a kill threshold is telling you what they believe about their own odds.
4. Data handling answered up front
This is where pilots die. Not in the demo, not in the pricing, in the security questionnaire that arrives in week five and takes six weeks to answer badly.
Get these answered before the kickoff: what data the agent reads and writes, where it is processed, which subprocessors touch it, whether the models train on your inputs and outputs (the answer must be no, and it must be traceable to the vendor’s terms with the model provider), how access is granted and revoked, and whether a DPA exists. Ask for the vendor’s trust documentation as a link, not a promise. If they have to write it for you, you have found the timeline risk.
Ours is public at trust and security: security posture, data handling, subprocessors, a governance summary aligned to the four functions of the NIST AI Risk Management Framework, and a DPA available on request. We do not hold SOC 2, and that page says so rather than burying it, because a procurement team finds out either way.
5. A fixed price, and what the second workflow costs
One number for the bounded job. Our published AI agent cost bands put a competently built single-workflow agent at $3,000 to $8,000, which is the right order of magnitude for a pilot. Also fix the price of workflows two and three now, before the first one works and your leverage disappears.
The 30, 60, 90 day shape
Days 0 to 7, before the build. Security questionnaire and DPA in flight. Baseline captured. Success metric, kill threshold, and both named owners signed off. Access granted at least privilege, in your own systems, revocable by you.
Days 7 to 30, build and shadow. The agent runs against real data but takes no irreversible action. Every output is reviewed by a human, and the disagreements are the training signal. You are measuring accuracy, not yet saving time.
Days 30 to 60, supervised live. The agent acts, a person approves anything irreversible, and the approval rate becomes the second metric worth watching. If approvals are near universal, widen the autonomy. If corrections are frequent, you have found the boundary of the workflow, which is a useful result even if it kills the pilot.
Days 60 to 90, decide. Three outcomes, written down in advance: expand to the next workflow, extend once with a specific fix and a new date, or stop. “Keep going and see” is not one of them, and it is how the 40 percent become the 40 percent.
What your security and legal teams will actually ask
Assemble this before the first call and you remove weeks from the cycle:
- The model question. Which models, from which providers, under what terms, and do they train on our data.
- The data flow. What leaves our environment, where it lands, how long it is retained, how it is deleted.
- The subprocessor list. Names, not categories.
- Access and revocation. Platform-native roles you grant and can pull yourself, versus credentials the vendor holds. The first is much easier to approve.
- The audit trail. What the agent did, when, on whose authority, readable by you without asking the vendor.
- Human approval boundaries. Which actions require a person, in writing. Deloitte’s 2026 survey of 3,235 IT and business leaders found only 21 percent have a mature agentic AI governance model, with decision boundaries, real-time monitoring, and audit trails named as the missing pieces. Your reviewer knows this.
- Regulatory posture. If you operate in the EU, note that the Digital Omnibus agreement deferred most high-risk obligations under the AI Act to December 2027 and August 2028, but Article 50 transparency obligations remain live from 2 August 2026, including disclosure when a person is interacting with an AI system. Ask the vendor how the agent discloses itself.
The pilot design mistakes we see most
Automating the interesting workflow instead of the expensive one. The interesting one demos well. The expensive one funds the next three.
No baseline. Without a before number, every after number is an argument.
Success defined as “it works.” It always works in a demo. Define the number.
A pilot that cannot fail. If there is no threshold that stops it, you have not bought a pilot, you have bought a first instalment.
Running two vendors on the same workflow at once. It feels rigorous and it is not: neither one gets clean data, both blame the other, and you spend twice to learn less.
What we do inside our own operation
We run this structure on ourselves. Our internal agent fleet reclaims 50 or more hours a week across the team, built one bounded workflow at a time, reviewed weekly, with a person approving anything irreversible. We also operate a public MCP server at mcp.strataigize.com, so an AI agent can query our business directly. We have not found another agency running one, and it exists because we needed it, not because it was a launch.
That is the reason we can price a first workflow with a straight face: we have run the month-nine version of this on our own operation, not just on a slide.
FAQs
How long should an AI agent pilot run?
Thirty to ninety days. Shorter than 30 and you are measuring novelty. Longer than 90 without a decision and the pilot has become the project, which is the failure mode Gartner’s cancellation data describes.
Should an AI agent pilot be paid?
Yes. Free pilots are staffed with spare capacity on both sides and measured against nothing, which is why they end in a demo rather than a decision. A paid, bounded pilot at the single-workflow price gets real people and a real success metric.
What is a reasonable kill threshold?
Roughly half to two thirds of the target metric, set so that missing it clearly means the workflow is not ready rather than that the agent had a slow week. The exact number matters less than the fact that it is written down and that someone is named to call it.
Who should own the pilot internally?
One person who owns the workflow today, not a committee and not the AI enthusiast. They know the exceptions, they can judge whether an output is right, and they will be honest about whether the hours actually came back.
What questions should we ask the vendor before this starts?
The full list is in 24 questions to ask an AI agent vendor. If you are still deciding whether to use a vendor at all, start with build versus buy.
Run one with us
We scope pilots exactly this way: one workflow, fixed price, written success metric, agreed kill threshold, security answered up front. Book a growth audit and we will name your highest-value workflow, or read how AI agent development works here.