How to structure a 90-day agency pilot that proves something
Phase-by-phase structure, the right KPIs for each month, pre-agreed decision gates — and the contract terms that make a pilot mean anything.
In short: A 90-day pilot beats a 12-month leap of faith — if you structure it. Write the baseline and success memo before day one, stabilise in month one, test in month two, scale and decide in month three against pre-agreed thresholds. Judge process and learning velocity, not week-two ROAS.
Why pilots beat annual contracts
The 12-month lock-in mostly protects the agency. A structured pilot caps downside for both sides: you risk one quarter, they get a fair shot at proving process. We'll say the quiet part out loud — The Shizz proposes pilots in most new engagements, because an agency confident in its process shouldn't fear a 90-day scoreboard. What a pilot cannot be is 30 days: algorithm learning phases, creative cycles and Indian delivery realities need a quarter to produce honest signal, which is exactly why we've documented how long performance marketing takes to show results. Ninety days is also long enough that a bad-faith agency can't simply stall through it.
Pilots discipline the client too. Writing down what success means before spending forces conversations about margins, capacity and honest baselines that annual contracts let everyone postpone for a quarter too long.
Before day one: the setup that decides the pilot
Most failed pilots fail in week minus-one — the agency inherits broken tracking, an unagreed baseline and three private definitions of success, and month three becomes archaeology instead of evaluation. Get four things in writing before any spend moves:
- Access done properly. Ad accounts, pixel and analytics live in your Business Manager, agency added as partner — never the other way around.
- A frozen baseline. Last 90 days of spend, revenue, blended CAC, contribution margin and repeat rate, agreed as the honest “before”.
- A one-page success memo. The three numbers that will decide the pilot, their thresholds, and who arbitrates disputes about measurement.
- Reporting cadence. Weekly numbers, fortnightly calls, and an agreed template — so month three isn't a fight about formats.
Days 1–30: stabilise — judge on process, not ROAS
Month one is tracking fixes, account restructure, stopping obvious bleed and getting the creative pipeline moving. Judging ROAS here punishes the agency for surgery you asked them to perform. Judge instead: how fast they diagnosed, how sharp their questions were, whether the first creative batch shows real thinking about your customer. Weekly written notes matter more than dashboards this month. One warning sign worth acting on: big wins claimed in week two are usually an attribution change, not growth — ask exactly what changed in measurement before celebrating anything.
Watch how they handle the awkward finds, too: undisclosed test budgets, dead subscriptions, double-counted conversions. An agency that surfaces problems it could have quietly hidden is showing you what the next two years would look like.
Days 31–60: test — judge on learning velocity
Month two is structured testing: hooks, offers, audiences, landing pages — each with a written hypothesis and a recorded result. The KPI is clean learnings per week, not ROAS wiggle; a team that banks eight real learnings in a month is building something compounding, whatever the topline says. Volume matters less than cleanliness — three contaminated tests teach nothing. This is also where scope honesty shows. A good agency flags problems outside its retainer — a leaking product page, a broken offer — instead of quietly optimising around them. Ask what they found that isn't their job to fix. And insist the learnings live in a shared document from day 31 onward, not in the agency's heads — if the pilot ends, those learnings are most of what you paid for.
Days 61–90: scale, and the decision gate
Month three scales what tested well, and now ROAS and CAC are fair judges — against the memo's thresholds, not against hope. The gate has three pre-agreed outcomes: continue (thresholds hit), extend 30–45 days (trajectory clearly positive but the starting account was a mess), or exit (numbers missed and process weak). Deciding on vibes after 90 days means the setup failed, not the agency. Hold the decision meeting in week 13 with the memo on screen — not in week 16, after two reschedules, when memory has replaced measurement.
| Phase | Focus | Judge on | Ignore |
|---|---|---|---|
| Days 1–30 | Tracking, structure, stop the bleed | Diagnosis speed, question quality | ROAS swings |
| Days 31–60 | Structured tests: hooks, offers, pages | Clean learnings per week | Single-week winners |
| Days 61–90 | Scale winners, kill losers | Thresholds from the success memo | Anything not written down |
Contract terms that keep a pilot honest
- Everything runs in your ad accounts; creative files and audience data are yours on exit.
- No auto-conversion into a 12-month contract — renewal is an explicit yes after the gate.
- Standard market pricing, not a heavily discounted “trial rate” — pilots priced below cost attract agencies that mass-produce them and staff them with juniors.
- 30-day notice either way, and a handover document owed regardless of outcome.
- A named senior owner on the agency side for the full 90 days.
If an agency resists these terms, that is the pilot result arriving early — and every one of them is easier to agree before the first invoice than after the last. Put the gate date in everyone's calendar on day one with the success memo attached. For the fuller diligence list, see the questions to ask a marketing agency before signing anything.
What a good pilot proved — and what it didn't
Ninety days proves process, communication, learning speed and directional economics. It does not prove your ceiling: compounding effects — the creative library, audience learnings, retention flows — take six to twelve months to show fully, which is why our average client relationship runs 1.5 years. Judge the pilot on trajectory and process quality, then judge the year on outcomes. Brands that get this order wrong either fire good agencies at day 90 or stay loyal to bad ones for two — both expensive, as the patterns in why D2C brands switch agencies show. The pilot's real output is confidence with evidence behind it — in the agency, or in your decision to walk.
Frequently asked questions
How long should an agency pilot be?
Ninety days. Thirty is too short for algorithm learning phases and creative cycles to produce honest signal, and six months is just a contract wearing a pilot's name tag.
What KPIs should a 90-day agency pilot measure?
Different ones per phase: diagnosis speed and process quality in month one, clean documented learnings per week in month two, and ROAS or CAC against pre-agreed thresholds in month three. Judging month one on ROAS is the most common way good pilots get misread.
Should you pay full price during an agency pilot?
Yes, standard market rates. Heavily discounted pilots select for agencies that run them in volume with junior staff. Negotiate exit terms and data ownership instead of the fee.
What if results are flat after 90 days?
Separate process from outcome. If learning velocity was high and the account inherited was a mess, a 30-45 day extension is often rational. If both the numbers and the process were weak, exit — flat results plus weak process rarely improves in month four.
Do good agencies accept pilots?
The confident ones often propose them — The Shizz does in most new engagements. Resistance to any structured trial, as opposed to negotiation over its terms, tells you something about how the agency expects the first quarter to go.
Start with an audit, not a contract
Before any pilot, we run a free Growth Audit — so the success memo is built on your account's real numbers, not guesses. It's the same first step that started 160+ brand relationships averaging 1.5 years.
Book a Growth Audit →