A 90-day agency pilot is a limited paid engagement to test how an agency works before making a longer commitment. Give it one defined channel or project, agree on a starting baseline, and write down how you’ll assess progress.
The review should judge both delivery and results. If the test produces too little data, an inconclusive verdict is more useful than declaring success or failure without evidence.
This playbook covers scope, checkpoints, contract terms, and the final decision. Our day-by-day accountability checklist explains what the agency should deliver during the pilot.
The short version: Scope the pilot to one defined job, record a starting baseline, and agree on delivery and outcome measures before day one. Review progress during the 90 days, not only at the end. Keep account access and exit terms clear. If the results lack enough data, treat the verdict as inconclusive and decide what evidence is still needed.
Why Does a 90-Day Pilot Beat a 12-Month Contract?
A 90-day pilot gives the agency time to launch tests and produce an initial cost read while limiting your commitment to one quarter. Both sides get evidence before agreeing to more. A bad fit costs one quarter, not one year.
An annual retainer makes agency revenue predictable, but it does not guarantee your results. Marketing performance does not need 12 months to become observable. It needs a clean test, enough budget, and an honest readout.
Good agencies like pilots too. Defined criteria protect them from clients who move goalposts mid-engagement. A documented win is also a stronger renewal argument than a signature from last January.
One caveat: a pilot is what you run after an agency passes basic screening. The 12 questions to ask before hiring a growth agency come first. The pilot tests the answers.
How Should You Scope an Agency Pilot?
Scope the pilot to one channel or system, never “everything.” Pick one paid platform, a lifecycle rebuild, or a creative testing program. Narrow scope makes the verdict easier to interpret. Broad scope blurs attribution.
When an agency touches five channels at once, nobody can say which effort moved the number by day 90. You end up judging vibes. A single channel against a single baseline keeps the pilot honest.
Scope also means resourcing. Agree on the media budget floor before launch. An underfunded test may lack enough data for a useful verdict, regardless of agency quality.
Alongside agency deliverables, list what your team must provide and give each item an owner and due date. Access delays, slow approvals, and missing creative can consume the test window. A written record helps separate agency execution failures from client-side delays at the review.
Choosing the first lever may require strategy work. Our growth strategy service starts with that diagnostic. It finds the highest-value pilot before you commit test budget.
What Success Metrics Should You Fix in Writing Before Day 1?
Agree on three checkpoints before the pilot starts: leading indicators by day 30, cost metrics by day 60, and pipeline by day 90. Record each metric, baseline, and owner in writing. Otherwise, the final review can turn into a debate about what success meant.
Here is the arc, phase by phase.
| Pilot phase | What to expect | What to measure |
|---|---|---|
| Days 1 to 30 | Access, tracking audit, signed baseline, first tests launching | Leading indicators: CTR, conversion rate, CPC vs baseline |
| Days 31 to 60 | Learning phases cleared, losers cut, weekly reporting running | Cost: CAC or cost per qualified lead trending |
| Days 61 to 90 | Winners scaled, playbook written, verdict review held | Pipeline: qualified opportunities and attributed revenue |
Day 30 is often too early for a final cost verdict. Meta’s current budget guidance recommends funding campaigns for at least seven days so delivery can learn. Many buyers use 50 weekly optimization events as a planning heuristic, not a guaranteed exit rule. Early reads should stay directional and use the agreed baseline.
Cost metrics may become useful around days 45 to 60 when conversion volume supports them. That is when CAC or cost per qualified lead becomes a fair discussion. By day 90, ask whether the work produced credible pipeline.
Need senior strategy and hands-on execution?
Our fractional growth leadership establishes the baseline, chooses the strongest opportunities, and gives the agency team clear experiments to run.
Book a Free Strategy CallWhat Contract Terms Make a Pilot Safe?
Four terms reduce pilot risk. Use a fixed 90-day term with no auto-renewal. Continue month to month with 30-day notice. Keep ownership of accounts, data, and creative. Use a flat fee that does not rise with spend.
Founders often overlook ownership. If the pilot runs in agency-owned accounts, important history may leave with the agency. You should be able to exit on day 91 with every account and asset intact.
Fee structure matters during a pilot. Percentage-of-spend pricing often runs 10% to 20% of ad budget, according to Feedbird. The agency earns more as spend rises during your efficiency test. A flat pilot fee removes that incentive.
Write each deliverable into the agreement. Require the baseline within two weeks, weekly reports, and a final playbook. You keep that playbook regardless of the verdict.
Define the data source for each metric in the same document. A CRM pipeline report and an ad-platform dashboard can tell different stories. Pick the source of truth before either side sees the result.
Include a minimum sample or spend threshold where possible. If the pilot misses it, label the verdict inconclusive. Do not turn a low-volume result into a pass or fail.
How Do You Judge the Pilot Verdict?
Judge the pilot against criteria set before day 1. Book the final review when the pilot begins. A pass beats the baseline with credible measurement. A fail misses both performance and process standards. An inconclusive result has promising evidence but too little data.
A pass does not mean the work is finished. It means the system works, so you continue month-to-month and expand scope one lever at a time.
If the pilot fails, explain why. The agency may have underdelivered, or broken tracking and insufficient budget may have left any agency unable to prove results. Record the cause in the review, keep the playbook, and use the agreed exit terms to leave cleanly.
An inconclusive verdict can invite endless extensions. Allow one 30-day extension with a single named metric. If that metric still lacks a usable result, close the pilot.
What Should You Own at the End of the Pilot?
At the end of a 90-day pilot, you should own the diagnosis, experiment record, channel data, creative files, and a sequenced 6 to 12 month roadmap. The agency should also make a written recommendation about whether to scale, extend one test, change direction, or stop.
Put those ownership terms in the contract before day 1. Your accounts, data, creative files, and roadmap should remain yours. The day-90 review should close with a written recommendation, including when the right answer is to take the roadmap in-house.
A useful pilot ends with a decision you can defend and work you can keep. The next step is choosing a problem narrow enough to test with the budget and data you have.
Let’s work through your proposed pilot before the clock starts. We can help define a manageable scope, a baseline, and the evidence you’d need to continue.
Like this? Get the next one.
Short emails. New posts as they ship.