A paid media partnership can become one of the largest commitments a marketing team makes, and one of the hardest to reverse. It gets harder when the partner will run several markets at once. Swapping agencies later means a new search, a new onboarding and months of lost learning.
The search alone is expensive. The ANA and the 4A's put the average client cost of an agency pitch at $408,500.
This guide shows how to scope an agency pilot project that answers the real question before the long contract starts. It covers scope, spend floor, pricing, success gates, pilot markets and contract terms.
P.S. If you'd rather start with a scoped test than a year-long contract, our paid media team runs 90-day pilots with the success gates written in before launch.
TL;DR
- A strong pilot usually runs for about 90 days, focuses on one or two channels, and limits the initial market scope.
- Set the budget high enough to generate enough conversions for the platform and test design to produce useful evidence.
- Pay a fixed pilot fee, kept separate from media, with no auto-renewal.
- Agree delivery, performance, learning and relationship success gates before launch.
- Keep every ad account, pixel and creative file in the brand's name.
Before the mechanics, one principle matters. We treat media buying as a laboratory, so ad spend should produce either performance or useful learning. An agency pilot applies the same principle to the agency itself.
The main purpose is to see how a partner thinks, tests and reports when creative, media and data sit in the same room, over a window long enough to learn something real. A pilot judged only on CPA is grading the market. A pilot judged on the learning system is grading the agency.
What Is an Agency Pilot Project?
An agency pilot project is a paid, fixed-scope engagement with an agency, usually about 90 days on one or two channels. It has written success gates and a set decision date, and a brand runs it before signing an agency-of-record contract.
Some teams call it a marketing agency pilot, a trial period or a test project. The label matters less than the structure: a defined scope, a defined budget and a written answer to "what would make us sign?"
How Does an Agency Pilot Differ From an Agency-of-Record Contract?
An agency-of-record (AOR) contract hands an agency ongoing responsibility for a channel or a brand, usually for a year or more. A pilot, on the other hand, is a bounded test of whether that agency deserves it.
The table below places the pilot next to the engagement types it usually gets confused with.
Engagement | Typical length | What it proves | What it can't prove | Exit terms |
|---|---|---|---|---|
Account audit | 2-4 weeks | Quality of diagnosis | Execution | None |
Agency pilot | About 90 days | Execution, learning system, fit | Long-run scale economics | No notice period, no auto-renewal |
Project engagement | Varies | One deliverable | Ongoing management | Ends at delivery |
AOR or retainer | 12-24 months initially | Compounding performance | Nothing it hasn't already earned | 60-90 days' notice |
The AOR terms reflect what Everything-PR reports as typical: 12 to 24 months for an initial engagement, with 60 or 90 days' notice to exit. If you only need a diagnosis, a paid media account audit is faster and cheaper than a pilot.
Why Run a Marketing Agency Pilot Before Signing an Agency of Record?
A pilot caps the downside of one of the most expensive procurement decisions in marketing. You find out how an agency actually works before a multi-year relationship makes switching slow and costly.
The Cost of a Wrong Agency Choice
A wrong pick is expensive to make and slow to undo. According to the ANA and 4A's 2025 tenure report that we shared above, the average client-agency AOR relationship now lasts about 7 years. Media-only agencies average 3.7 years.
The same report shows a clear difference in tenure between review approaches. Brands that review their agencies frequently average as little as 3.8 years, while those without mandatory reviews average 8.1 years. An early exit can also mean another selection process, another onboarding period and additional disruption.
Why Agency Pilots Are Gaining Ground
Brands are switching agencies more frequently while putting agency spend under greater scrutiny. In 2025, incumbents kept just 21% of the media spend put up for review, according to COMvergence data, the lowest retention rate in eight years. Gartner's CMO survey found 39% of CMOs plan to cut agency budgets.
When more accounts move and budgets are under review, proving fit before committing becomes standard practice.
A pilot makes the most sense when:
- You're entering a new region and need a partner with real local execution.
- An incumbent's results are declining and you want evidence before switching.
- You're consolidating several regional agencies into one global partner.
If you're still building the shortlist, our guide on how to choose the right growth agency covers that earlier stage.

How to Pilot an Agency: Scope a 90-Day Pilot in 8 Steps
To pilot an agency, write one testable hypothesis, lock the baseline, narrow the scope, then put market, budget, team, gates and fee on a single page both sides sign.
The eight steps below run in order, and each one depends on the step before it.
- Write the pilot hypothesis in one sentence. Use this format: "If {agency} runs {channel} in {market}, we expect {metric} to move from {baseline} to {target} by the end of the pilot."
For example: "If the agency runs Meta prospecting in Germany, we expect cost per first purchase to beat our trailing-quarter baseline by week 12." One sentence forces both sides to agree on the channel, the market and the number that matters. Our ad testing service starts every engagement with a 90-day testing roadmap tied to ROAS, CPA or CAC, which is the same discipline applied to creative.
- Lock the baseline and a metric dictionary in the first two weeks. The baseline is your trailing 90 days of CAC, CPA and MER. The metric dictionary defines what counts as a conversion, which attribution window applies, and which system is the source of truth: the platform, GA4 or your MMM. Once it's signed, nobody gets to redefine success halfway through. Our marketing agency metrics guide covers which metrics deserve a place in it.
- Pick one or two channels and exclude everything else in writing. Site rebuilds, brand campaigns and SEO stay out of scope. Every extra workstream adds a variable, and a pilot with too many variables produces a result nobody can read.
- Choose the pilot market. Select a market large enough to generate reliable data and representative enough to reflect the planned rollout. Market selection is covered in detail below.
- Set the media budget. Size the budget so each test cell can exit the platform's learning phase. The math is in the ad spend section below.
- Name the team. Ask for a named senior strategist, media buyer and creative lead, with hours per week for each, plus a clause that keeps them on the account for the whole pilot. The people who pitch are not always the people who run the work, and a pilot is your one chance to test the real team.
- Agree the success gates and kill criteria before launch. Write down what passes, what fails and what ends the pilot early before the first dollar is spent. The scorecard is below.
- Put the pilot on one page. Seven lines are enough: scope, dates, deliverables, success gates, fee, account ownership and the decision date. This one-pager becomes the document both sides sign, and later the starting point for the AOR contract.
Pro tip: We recommend asking each finalist to walk through its week-two test plan during the pitch. How specific that plan is tells you more about the pilot than the credentials deck does.

What a 90-Day Agency Pilot Timeline Looks Like
A 90-day agency pilot runs in five phases: two weeks of setup, two weeks of launch, a month of iteration, a month of scaling and scoring, then a final review in week 13.
Weeks | Checkpoint | What to verify |
|---|---|---|
1-2 | Plan approved | Baseline, tracking and test plan agreed |
3-4 | Campaigns live | Launch happened within the agreed scope |
5-8 | First learning read | Tests are producing documented insights |
9-12 | Gates scored | Performance and learning evidence is complete |
13 | Final review (day-91 decision) | Decision-makers have the full pilot readout |
Let’s see what the agency should be working on in each phase:
- Weeks 1-2: Complete the account and tracking audit, finalize the written test plan and agree on the metric dictionary.
- Weeks 3-4: Launch the first campaigns and creative tests within the agreed scope. Confirm that tracking works and each test has enough variation to produce a useful comparison.
- Weeks 5-8: Run iteration rounds and produce the first written test readouts, including what worked, what failed and what should change next.
- Weeks 9-12: Scale stronger performers, review holdout results where applicable and refresh creative as creative fatigue starts to affect performance.
- Week 13: Deliver the final readout, the scale plan and the recommendation for the next stage.
Add one mid-point review in week 6. Ask for a written "on track" or "at risk" verdict against each success gate, so the final review holds no surprises.
How Much Ad Spend Does an Agency Pilot Need?
For Meta and TikTok, budget should be large enough to generate roughly 50 optimization events within the platform's learning window. Google requires a different approach based on conversion volume and bidding strategy.
Below those thresholds, the algorithm never stabilizes, and the pilot ends up measuring noise.
Here's what each platform says:
- Meta: An ad set exits the learning phase after about 50 results in the week after its last significant edit. An ad set that's unlikely to get there is marked "learning limited", and Meta lists low budget as one of the causes.
- TikTok: Its help center calls 50 conversions the most significant indicator that an ad group has passed the learning phase.
- Google: Smart Bidding targets take one to two conversion cycles to settle, and Target ROAS needs at least 15 conversions in the last 30 days.
For Meta and TikTok, that gives a working formula: minimum weekly spend per ad set is roughly 50 × your target CPA. At a $40 CPA, one ad set needs about $2,000 a week.
A pilot testing three ad sets at once needs about $6,000 a week, or roughly $26,000 a month in media. This figure is derived from the platform rule, so treat it as a planning floor.
Platform | Learning threshold | Pilot budget implication |
|---|---|---|
Meta | About 50 results in the week after the last significant edit | Budget about 50 × CPA per ad set per week, and avoid edits that reset learning |
TikTok | 50 conversions | Same formula; consolidate ad groups to reach it faster |
Google (Smart Bidding) | 1-2 conversion cycles; 15+ conversions in 30 days for Target ROAS | Allow two to four weeks before judging bid performance |

Our Meta advertising and TikTok advertising teams size every pilot around this floor before anything launches.
What to Do When the Pilot Budget Can't Reach the Learning Floor
Adjust the test design before reducing the budget further. Three options can preserve the quality of the evidence:
- Optimize for a higher-funnel event. Add-to-cart or lead events happen more frequently than purchases, so they reach 50 a week at a lower spend.
- Cut the number of ad sets. Two well-funded test cells beat five starved ones.
- Extend the pilot. A 120-day pilot with a readable result is worth more than a 90-day pilot with an unreadable one.
If the pilot includes a geographic holdout, both markets need enough budget to generate a usable comparison. The market-selection section below covers that structure.
How Should an Agency Pilot Be Priced?
Pay a fixed pilot fee, kept separate from media, at or slightly above the agency's normal retainer rate. Free pilots can create weak incentives and make senior staffing harder to secure.
A pilot costs the agency more per month than a steady retainer does, because the audit, tracking fixes and account restructuring all land in the first few weeks. Heavy discounting can also make it harder to allocate senior staff to this setup work.
These are the five models you'll usually see:
Model | How it works | Best for | Risk to brand | Risk to agency |
|---|---|---|---|---|
Flat pilot fee | Fixed fee for the pilot period | Most pilots | Low | Absorbs setup cost |
Fee credited to AOR | Pilot fee is deducted from the first retainer months if you convert | Brands likely to convert | Low | Low |
Fee-at-risk | A share of the fee, typically 10-20% in our experience, tied to hitting the gates | Brands with a mature baseline | Low | Medium |
Percentage of media spend | Fee scales with spend | Rarely suits a pilot | Rewards spending over learning | Low |
Free or spec work | No fee | Avoid | Weak incentives; senior staff harder to secure | High |
Before signing, we suggest checking these four things in the pilot quote:
- Is media billed separately from the agency fee?
- Are setup fees refundable?
- Is the fee fixed for the whole pilot, with no overage charges?
- Will the pilot fee be credited against the retainer if you sign an AOR?
Which Success Gates and Kill Criteria Should an Agency Pilot Use?
Score four success gates, performance, learning, delivery and relationship, from 1 to 5 at the go/no-go review. Then set five kill criteria that end the pilot early.
This scorecard evaluates a marketing agency using evidence from your own account. The success gates answer "Did this agency earn the contract?"
The kill criteria answer "Should we stop now?" Write both down before launch so the evaluation criteria do not change in week 12.
Below is the weighting our team typically uses as a starting framework for agency pilots:
Success gate | Weight | What a 5 looks like | What a 1 looks like |
|---|---|---|---|
Performance vs. baseline | 35% | Beats baseline CPA, and holdout lift is positive | Worse than baseline, with no diagnosis |
Learning system | 25% | Documented tests, clear winners and losers, next briefs written | No test log |
Delivery | 20% | Every launch on time, reporting on schedule | Two or more missed launches |
Relationship and transparency | 20% | Full account access, senior team present, replies within the SLA | Screenshots only, team changed |
How to Set Performance Gates for an Agency Pilot
Set performance gates as staged thresholds against the baseline you locked in step 2. A single end-of-pilot number doesn't work, because paid media rarely improves in a straight line.
A workable pattern is "within 10% of baseline CPA by week 8, and beating it by week 12". Where you can, add incrementality: compare the pilot market with its holdout so platform-reported gains don't carry the whole verdict. If measurement isn't set up in-house, a marketing measurement agency can help you set up the comparison before launch.
How to Set Learning Gates for an Agency Pilot
Learning gates measure whether the agency leaves you smarter than it found you. Ask for three artifacts by the final review:
- A test log: Every hypothesis tested, the result and the decision it drove.
- A count of hypotheses tested: A pilot that tested three ideas in 90 days hasn't built a learning loop.
- A creative insights brief: Which angles, hooks and personas won, written so your team keeps the learning even if you don't sign.
In our experience, this is the gate that best predicts the next twelve months. Market conditions can change CPA, while a strong creative testing system keeps producing evidence for the next round of decisions. .
Which Kill Criteria Should End an Agency Pilot Early?
Kill criteria are the few signals serious enough to stop the pilot before the final review:
What you see | What it likely means | Action |
|---|---|---|
Two consecutive missed launches | The account is understaffed | Stop at the next checkpoint |
Metric definitions changed without sign-off | Reporting can't be relied on | Stop |
No access to the live ad account | Lack of transparency | Stop |
Named lead replaced without agreement | Staffing doesn't match the pitch | Escalate, then stop |
Ad sets still learning-limited after week 6, with no plan | Poor account structure | Fix within 7 days, or stop |
Our guide to marketing agency reporting covers what weekly reporting should include.
How to Choose Pilot Markets for a Global Agency Partner
Pilot in one representative mid-tier market, hold out a comparable market, and test the global operating model alongside local results. This combination tells you whether the agency can perform and whether it can repeat that performance across markets.
1. Pick a Representative Pilot Market
Choose the market that looks most like the rest of the rollout. The largest market may carry too much revenue risk for an initial test, while the smallest may not generate enough volume to produce a useful read.
A good pilot market has three things:
- A known baseline, with at least 90 days of clean data.
- Enough spend potential to clear the platform learning thresholds.
- The language, creative and regulatory load typical of the markets that follow.
Market option | Learning speed | Risk | How well it predicts the rollout | Verdict |
|---|---|---|---|---|
Largest market | Fast | High | High | Avoid for pilots |
Mid-tier representative market | Medium | Medium | High | Recommended |
Smallest market | Slow, often below the learning floor | Low | Low | Avoid |
Two markets (test plus holdout) | Medium | Medium | High | Best if budget allows |
2. Use a Matched Holdout Market
Pair the pilot market with a comparable market the agency doesn't touch. The holdout keeps running as it did before, so seasonality, pricing changes and category shifts show up in both markets. What's left over is the agency's effect.
Match the two markets on size, seasonality and channel mix, and agree the comparison method before launch. Our guide to incrementality testing walks through geo holdouts and the calculation behind them.
3. Test a Global Agency's Operating Model
A global partner has to show that its performance travels. Local results only tell you what happened in one market. While the pilot runs, check how the agency handles:
- Consolidated reporting: One dashboard, one currency, one metric dictionary across markets.
- Localization workflow: How briefs, creative and copy move from one market to the next, and how long it takes.
- Time-zone coverage: Who answers when a campaign breaks at 9 a.m. in your pilot market.
- Local vs. hub staffing: Whether a central hub or an in-market team runs your account.
- Network vs. bespoke structure: In H1 2026, 26% of reviewed media spend went to bespoke units run by the big holding companies. Ask which structure you're actually buying.
- Continuity after rollout: Whether the people running the pilot will stay on the account once it expands.
Localization demand rises quickly as more markets are added . Our partnership with NielsenIQ involved 250+ content creators across 19 countries and content in 15+ languages. A pilot should show whether the agency's workflow can absorb that kind of volume before you sign it up for all of it.
How to Structure Responsibilities, Ownership, and Contract Terms in an Agency Pilot
Give the agency ownership of the test plan and execution, keep every account and asset in the brand's name, and sign a pilot contract with no auto-renewal and no notice period.
1. Define Brand and Agency Responsibilities
Brand-side delays distort pilot results as often as agency mistakes do, so write both sides' duties down:
Area | Brand | Agency |
|---|---|---|
Access | Grants partner access on day 1 | Confirms tracking works |
Approvals | Approves creative and budget changes within the agreed SLA | Submits changes with a clear rationale |
Data | Provides baseline and finance data | Maintains the metric dictionary |
Test planning | Reviews and signs off | Writes and runs the test plan and log |
Budget decisions | Approves | Recommends |
Reporting | Attends weekly reviews and makes the final decision | Reports weekly and presents the end-of-pilot readout |
Add one clause to the one-pager: if a brand-side approval runs past the SLA, the pilot clock pauses. Otherwise the agency gets graded on weeks it couldn't use.
2. Keep Accounts and Assets Brand-Owned
Everything the pilot creates belongs to the brand from day one:
- Ad accounts, pixels, Conversions API setup, GA4 properties and tag containers, all registered in the brand's own Business Manager or Google Ads manager account. The agency works through partner access.
- Creative source files, including editable project files.
- Creator content and its licences, with usage rights that survive the end of the pilot.
3. Set the Pilot Contract Terms
Five terms protect the brand in an agency pilot:
- No auto-renewal. The pilot ends on its date unless both sides sign something new.
- No notice period during the pilot. AOR contracts typically carry 60 to 90 days' notice. A pilot that needs notice to end has stopped being a pilot.
- Media billed at cost. Any rebates, volume discounts or platform credits are disclosed.
- Learnings transfer. The test log, insights brief and reports belong to the brand, signed or not.
- A written decision date. Both sides know when the verdict happens and who makes it.
Pricing and fee credits sit with the pricing models above, so the contract only needs to reference them.
Should You Pilot Two Agencies or Test Against an Incumbent?
Split agencies by market or channel, and never put two agencies on the same audience on the same platform. Two agencies bidding into one auction raise each other's costs and argue over attribution.
The table below compares the main pilot structures and the trade-offs each one creates.
Structure | How it works | Fair test? | Main risk | Use when |
|---|---|---|---|---|
Parallel, same channel and market | Two agencies share one audience | No | Auction overlap and attribution disputes | Never |
Parallel, split markets | Each agency gets a matched market | Yes | Market differences | Choosing between two finalists |
Challenger vs. incumbent | The challenger takes one market or channel | Mostly | Different incentives during the test | Replacing a declining incumbent |
Sequential | One agency runs, then the other | Weak | Seasonality between the two windows | Budget too small for a parallel test |
Testing a challenger against the incumbent is the most common setup and the easiest to get wrong. The incumbent knows it's being tested, and its effort can shift in either direction. Give it a written brief covering the same 90 days and score both agencies on the same scorecard, so the comparison rests on identical rules.
How to Convert a Successful Pilot Into an Agency of Record (AOR)
At the day-91 decision, write the pilot's artifacts straight into the AOR: the metric dictionary, test log, named team, rate card and expansion triggers. The pilot already proved how the partnership works, so the contract should lock that in.
Five conversion terms carry the most weight:
- A 12-month rate lock at the rates quoted during the pilot.
- The pilot fee credited against the first retainer months, if you agreed that in the pricing model.
- Named-team continuity: the people who ran the pilot stay on the account.
- Market expansion triggers: for example, "add market B once market A holds its target CAC for 60 days".
- A performance review at month 6, scored against the same gates as the pilot.
A clean conversion also removes the need for repeated re-pitches. As the tenure data earlier in this article shows, brands that skip mandatory review cycles keep their agencies the longest.
When an Agency Pilot Is the Wrong Choice
A pilot is the wrong tool in four situations, and each has a better alternative:
- The budget can't reach the learning floor. Run an account audit, or plan a longer pilot at a lower weekly spend.
- The 90 days overlap your peak season. Move the window, or peak-season noise will swamp the result.
- The work is brand or creative-platform work with a long payback. Scope it as a defined project with creative deliverables.
- A full rebrand or replatform is underway. Wait until tracking and the site are stable, or start with an audit.
Scope Your Agency Pilot With inBeat's Paid Media Team
A good agency pilot sizes its budget to the learning floor, scores the partner on four gates, and reads results against a holdout market. This gives you evidence about how the agency thinks, tests and reports, which is the thing a credentials deck can't show you.
That's how our paid media team runs pilots. Performance creative, media buying and measurement sit together from week one, so the test log, the creative insights and the performance data all come from one team.
For brands consolidating regional agencies or evaluating a replacement for an incumbent, the pilot should answer one question: should this partnership scale?
Book a strategy call with us to review your baseline, size the pilot budget and draft the one-page agreement.
FAQs
Does a 90-day agency pilot still work when the sales cycle runs 60 days or more?
Yes, with adjusted gates. For B2B or high-consideration products, score the pilot on leading indicators such as qualified leads, cost per sales-accepted lead and pipeline created.
Then extend the performance read by another quarter. The learning and delivery gates still apply unchanged at day 91.
How do you keep the test fair when the incumbent still runs the rest of the account?
Ring-fence the test inside the ad accounts. Exclude the challenger's market or audience from every campaign the incumbent still runs, and freeze the incumbent's budget in that market for the pilot period.
Without those exclusions, retargeting and broad campaigns bleed into the test and both agencies claim the same conversions. Keep live account access on both sides so either team can check.
Should pilot creative come from the agency or from the brand?
From the agency, at least for the test cells. If the brand supplies all the creative, the pilot only grades media buying, and half of what drives paid social performance goes untested. Brand guidelines and approvals still apply.
What happens to learnings and creative if you don't convert?
They stay with the brand under the contract terms above. The step most teams miss is exporting before partner access is removed: pull the ad account's historical data, the test log and any dashboards the agency built in its own tools, because those often don't live in the brand's accounts.
Can an agency pilot run on Advantage+ or Performance Max when the algorithm controls most of the levers?
Yes. Automated campaigns shift the agency's value to the inputs the algorithm can't generate itself: creative variety, conversion signal quality and measurement design. Grade those.
Bid and targeting tweaks matter less in automated campaigns, so they shouldn't carry much of the score.







