You evaluate a GTM AI pilot against a document you wrote before it started. A pilot that begins without a pre-registered primary metric, a kill threshold, a holdout group, and a decision date backed off the contract’s notice window does not produce an evaluation — it produces a dashboard, and the dashboard belongs to the vendor. Four design choices decide whether the pilot answers anything: what you pre-register, how you isolate a control, how you count human rework, and where the decision date sits relative to the renewal clock. This page gives each one a default you can start from.
The failure is upstream of the model
Two numbers get quoted at every AI steering committee. MIT’s Project NANDA reported in The GenAI Divide: State of AI in Business 2025 that 95% of enterprise generative-AI pilots produced no measurable P&L impact, drawn from 52 structured interviews, 153 surveyed leaders, and a review of 300 publicly disclosed deployments between January and June 2025. Separately, IDC research with Lenovo found that 88% of AI proofs of concept never reach production — four graduating for every 33 launched.
Both numbers are softer than they look. The NANDA document is labelled preliminary findings rather than peer-reviewed work, and its methodology drew public criticism on sample construction and on the leap from “no measurable P&L impact” to “failure.” Do not cite either figure as evidence that your pilot will fail. Cite them for the thing both research teams actually name as the lead obstacle: not model quality, but unclear success criteria and no agreed definition of good enough. That failure happens in the first week, before a single message goes out, and it is the only part of the pilot you fully control.
1. Pre-register four lines
Write these before the first run, in a document with a date on it, circulated to whoever will sit in the renewal meeting.
The primary metric — exactly one. For an outbound agent it is fully-loaded cost per held qualified meeting. Not meetings booked, because booked and held diverge most in precisely the population an AI agent overproduces. Not reply rate, because a reply is not revenue and an offended reply counts the same as a positive one. Secondary metrics are fine to collect; only one decides.
The threshold, expressed against your own baseline. Measure the trailing 90 days before the pilot on the same metric, and write the pass mark as a ratio to it. Published benchmark reply rates are useless here: they encode someone else’s list quality, offer, domain reputation, and segment. Your baseline is the only comparison that holds those constant.
Kill criteria that end the pilot early. These are breach conditions, not performance misses — a spam-complaint rate crossing your ceiling, a compliance or brand incident, an agent writing bad data into the CRM at a rate that costs more to clean than the pilot can return. Name the threshold and the person who can pull the plug without convening a committee.
The decision date. Section 4 sets it. It is not “end of Q3.”
One more line, and it is the one that gets skipped: the amendment rule. Criteria changed after data starts arriving invalidate the comparison. If you must amend, the memo records both versions and reports the result as retrospective.
2. Build the holdout
This is the step nearly every GTM pilot skips, and skipping it means no result the pilot produces can be attributed to the tool. A quarter with a new agent in it is also a quarter with different seasonality, a repriced competitor, two SDRs ramping, and a changed offer. Without a control arm you are measuring all of that at once.
- Randomize at the account level, not the lead level. Two contacts at the same account landing in different arms contaminate both — they talk to each other, and multi-threading means the treatment reaches the control.
- Hold out 20-30% of eligible accounts. Below 20% the control arm produces too few outcome events to resolve a realistic effect. Above 30% you pay more pilot cost per unit of learning without changing the answer.
- Stratify on whatever predicts the outcome — usually ICP tier or segment — so the two arms are not accidentally different on the variable that matters most.
- Run both arms in the same window. Before-and-after is not a control; it is a control plus every other thing that changed.
- Set an outcome floor. If the control arm yields fewer than about 30 outcome events across the window, the honest finding is “not enough signal to decide,” and that is a legitimate pilot result. It is also a reason to negotiate a month-to-month term rather than sign an annual one on a coin flip.
3. Count the human rework
The strongest published evidence that self-reported AI productivity runs in the wrong direction comes from METR’s randomized trial of 16 experienced open-source developers across 246 real tasks. They forecast that AI tools would speed them up by 24%. They were measured 19% slower. Afterwards, having lived the slowdown, they still estimated AI had made them 20% faster.
METR itself labels that result historical and disclaims generalization beyond software development, so do not import the 19% into a GTM model. Import the sign error. The participants were not confused about the magnitude of the effect; they were wrong about its direction, on their own work, immediately afterward. Any renewal decision that reconstructs rework from memory is running the same experiment.
So log it during the pilot, in four buckets: configuration and prompt tuning, output review and editing, escalation and reply handling, and cleanup — bad CRM writes, wrong-contact apologies, sequences pulled mid-flight. Sample rather than log everything; two tracked hours per week per participating rep is enough to build a rate. Rework hours belong in the numerator of cost per outcome at fully-loaded rates, alongside the license fee and the credit spend. See seat-based vs usage-based AI pricing for the consumption half of that number.
4. Work the calendar backwards from the notice window
The renewal-date trap runs the same way every time. The pilot ends, results are ambiguous, somebody reasonably asks for one more month — and the non-renewal notice deadline passes during the extension. The contract renews at list price, and the extension you granted to gather more evidence has become the decision.
Schedule backwards, in this order:
- Read the notice window in your own MSA. SaaS non-renewal notice periods run 30 to 90 days before term end. The only number that matters is the one in your agreement, and it is worth confirming that the clock starts at term end rather than at an anniversary date.
- Decision date = notice deadline minus 14 days. That fortnight is memo writing, finance review, and one exec calendar slip.
- Pilot end = decision date minus 14 days. Meetings booked in the final week have not held yet, and deliverability damage surfaces on a lag. A pilot scored on its last send date is scored before its own consequences arrive.
- Pilot start = pilot end minus the pilot length. Six to eight weeks for outbound email; longer only if your sales cycle makes the outcome event rarer.
- Baseline = the 90 days before pilot start. You already have this data; pulling it is an hour, and skipping it is what makes the result unreadable.
There is no regulatory backstop underneath this. The FTC’s Negative Option Rule — the “click-to-cancel” rule — was vacated by the Eighth Circuit on 8 July 2025 on procedural grounds, days before it took effect, and the Commission sent a draft advance notice of proposed rulemaking to OIRA on 30 January 2026 to restart the process. Even had it survived, it governed consumer negative-option plans and would not have reached a B2B SaaS agreement. The notice window in your contract is the entire protection.
Calibrated defaults
Start here and move a number only when you can say what about your motion makes it wrong.
| Decision | Default | Move it when |
|---|---|---|
| Primary metric | Fully-loaded cost per held qualified meeting | The agent’s output is not a countable outcome |
| Comparison | Your trailing-90-day baseline | Never — published benchmarks do not control for your domain or list |
| Randomization unit | Account | Never for outbound; multi-threading leaks treatment |
| Holdout share | 25% of eligible accounts | Below 20% only with an outcome rate high enough to hit the floor |
| Outcome floor | 30 events in the control arm | Raise it when the expected effect is under 20% |
| Pilot length | 6-8 weeks | Longer for sales cycles where the outcome event is rare |
| Rework sampling | 2 tracked hours per rep per week | Raise it in the first two weeks, when configuration load peaks |
| Tail before scoring | 14 days after the last send | Longer where meetings routinely book 3+ weeks out |
| Decision date | Notice deadline minus 14 days | Earlier when the memo needs a board or procurement cycle |
Pitfalls and guards
- The vendor’s dashboard becomes the scoreboard. It counts sends, replies, and bookings — the three things it can see — and cannot see held meetings, rework hours, or cleanup. Guard: score from your CRM, your finance export, and your timesheets, and agree the data sources in the pre-registration before anyone has a stake in the answer.
- Sales hand-picks the pilot accounts. The agent gets the warm list, the control gets the remainder, and the comparison is dead before the first send. Guard: randomize from the eligible pool programmatically, and have the list assignment reviewed by someone who does not carry the number.
- “Extend the pilot” replaces “decide.” Guard: an extension is permitted only when the notice deadline still falls after the new decision date. Otherwise extending is renewing, and the memo should say so in those words.
- Booked substituted for held. Guard: the metric definition names the held event and states the no-show treatment before the pilot starts, because an agent tuned for volume moves booked and held in opposite directions.
- Rework reconstructed at the renewal meeting. Guard: the sampled time log, kept weekly. METR’s participants misjudged the direction of the effect on their own work within hours of doing it.
- A pilot that produces no signal gets read as a pass. Guard: the outcome floor is written down in advance, and “below floor” is a named verdict with its own action — a month-to-month term, not an annual one.
Where this framework breaks down
It assumes enough volume to split. If you touch fewer than roughly 200 accounts a month, a holdout large enough to resolve anything leaves a treatment arm too small to be worth running, and the measurement apparatus costs more than the tool. Buy on judgment instead, say so explicitly, and put the discipline into the contract term — month-to-month or a short initial term with a negotiated notice window — rather than into a control group that cannot carry a conclusion.
It also assumes a countable outcome. For a research assistant or an in-CRM copilot, cost per meeting is the wrong frame entirely; drop it and run the rework ledger alone as a task-level time study against a matched set of tasks done without the tool. And for the scoring step once the pilot closes, the AI SDR pilot scorecard skill implements the four axes and refuses a verdict when criteria were amended mid-flight or the sample is short. For the architecture question that decides what you should be piloting at all, start with AI SDR and signal-driven vs autonomous AI SDR.
One last signal that costs nothing to read. If a vendor will not sell a term that ends after your decision date, or resists a pilot with a holdout in it, that is information about the vendor’s confidence in the result — and it arrives before you have spent anything.