Agency Pilot Projects: How to Scope a 90-Day Pilot Before Committing to a Global Partner

Ioana Cozma
Published:
October 5, 2026
|
Updated:

A paid media partnership can become one of the largest commitments a marketing team makes, and one of the hardest to reverse. It gets harder when the partner will run several markets at once. Swapping agencies later means a new search, a new onboarding and months of lost learning.

The search alone is expensive. The ANA and the 4A's put the average client cost of an agency pitch at $408,500.

This guide shows how to scope an agency pilot project that answers the real question before the long contract starts. It covers scope, spend floor, pricing, success gates, pilot markets and contract terms.

P.S. If you'd rather start with a scoped test than a year-long contract, our paid media team runs 90-day pilots with the success gates written in before launch.

TL;DR

  • A strong pilot usually runs for about 90 days, focuses on one or two channels, and limits the initial market scope.
  • Set the budget high enough to generate enough conversions for the platform and test design to produce useful evidence.
  • Pay a fixed pilot fee, kept separate from media, with no auto-renewal.
  • Agree delivery, performance, learning and relationship success gates before launch.
  • Keep every ad account, pixel and creative file in the brand's name.

Before the mechanics, one principle matters. We treat media buying as a laboratory, so ad spend should produce either performance or useful learning. An agency pilot applies the same principle to the agency itself.

The main purpose is to see how a partner thinks, tests and reports when creative, media and data sit in the same room, over a window long enough to learn something real. A pilot judged only on CPA is grading the market. A pilot judged on the learning system is grading the agency.

What Is an Agency Pilot Project?

An agency pilot project is a paid, fixed-scope engagement with an agency, usually about 90 days on one or two channels. It has written success gates and a set decision date, and a brand runs it before signing an agency-of-record contract.

Some teams call it a marketing agency pilot, a trial period or a test project. The label matters less than the structure: a defined scope, a defined budget and a written answer to "what would make us sign?"

How Does an Agency Pilot Differ From an Agency-of-Record Contract?

An agency-of-record (AOR) contract hands an agency ongoing responsibility for a channel or a brand, usually for a year or more. A pilot, on the other hand, is a bounded test of whether that agency deserves it.

The table below places the pilot next to the engagement types it usually gets confused with.

Agency pilot vs other engagement types

Engagement

Typical length

What it proves

What it can't prove

Exit terms

Account audit

2-4 weeks

Quality of diagnosis

Execution

None

Agency pilot

About 90 days

Execution, learning system, fit

Long-run scale economics

No notice period, no auto-renewal

Project engagement

Varies

One deliverable

Ongoing management

Ends at delivery

AOR or retainer

12-24 months initially

Compounding performance

Nothing it hasn't already earned

60-90 days' notice

The AOR terms reflect what Everything-PR reports as typical: 12 to 24 months for an initial engagement, with 60 or 90 days' notice to exit. If you only need a diagnosis, a paid media account audit is faster and cheaper than a pilot.

Why Run a Marketing Agency Pilot Before Signing an Agency of Record?

A pilot caps the downside of one of the most expensive procurement decisions in marketing. You find out how an agency actually works before a multi-year relationship makes switching slow and costly.

The Cost of a Wrong Agency Choice

A wrong pick is expensive to make and slow to undo. According to the ANA and 4A's 2025 tenure report that we shared above, the average client-agency AOR relationship now lasts about 7 years. Media-only agencies average 3.7 years.

The same report shows a clear difference in tenure between review approaches. Brands that review their agencies frequently average as little as 3.8 years, while those without mandatory reviews average 8.1 years. An early exit can also mean another selection process, another onboarding period and additional disruption.

Why Agency Pilots Are Gaining Ground

Brands are switching agencies more frequently while putting agency spend under greater scrutiny. In 2025, incumbents kept just 21% of the media spend put up for review, according to COMvergence data, the lowest retention rate in eight years. Gartner's CMO survey found 39% of CMOs plan to cut agency budgets.

When more accounts move and budgets are under review, proving fit before committing becomes standard practice.

A pilot makes the most sense when:

  • You're entering a new region and need a partner with real local execution.
  • An incumbent's results are declining and you want evidence before switching.
  • You're consolidating several regional agencies into one global partner.

If you're still building the shortlist, our guide on how to choose the right growth agency covers that earlier stage.

Horizontal bar chart of average agency-of-record tenure: integrated and independent agencies 7.3 years, holding-company agencies 5.8, media-only agencies 3.7, no mandatory reviews 8.1 and frequent reviews 3.8.
Media-only agencies and frequently reviewed accounts have the shortest AOR tenure. Source: ANA and 4A's, Client-Agency Relationship Tenure Report, 2025

How to Pilot an Agency: Scope a 90-Day Pilot in 8 Steps

To pilot an agency, write one testable hypothesis, lock the baseline, narrow the scope, then put market, budget, team, gates and fee on a single page both sides sign.

The eight steps below run in order, and each one depends on the step before it.

  1. Write the pilot hypothesis in one sentence. Use this format: "If {agency} runs {channel} in {market}, we expect {metric} to move from {baseline} to {target} by the end of the pilot."

For example: "If the agency runs Meta prospecting in Germany, we expect cost per first purchase to beat our trailing-quarter baseline by week 12." One sentence forces both sides to agree on the channel, the market and the number that matters. Our ad testing service starts every engagement with a 90-day testing roadmap tied to ROAS, CPA or CAC, which is the same discipline applied to creative.

  1. Lock the baseline and a metric dictionary in the first two weeks. The baseline is your trailing 90 days of CAC, CPA and MER. The metric dictionary defines what counts as a conversion, which attribution window applies, and which system is the source of truth: the platform, GA4 or your MMM. Once it's signed, nobody gets to redefine success halfway through. Our marketing agency metrics guide covers which metrics deserve a place in it.
  2. Pick one or two channels and exclude everything else in writing. Site rebuilds, brand campaigns and SEO stay out of scope. Every extra workstream adds a variable, and a pilot with too many variables produces a result nobody can read.
  3. Choose the pilot market. Select a market large enough to generate reliable data and representative enough to reflect the planned rollout. Market selection is covered in detail below.
  4. Set the media budget. Size the budget so each test cell can exit the platform's learning phase. The math is in the ad spend section below.
  5. Name the team. Ask for a named senior strategist, media buyer and creative lead, with hours per week for each, plus a clause that keeps them on the account for the whole pilot. The people who pitch are not always the people who run the work, and a pilot is your one chance to test the real team.
  6. Agree the success gates and kill criteria before launch. Write down what passes, what fails and what ends the pilot early before the first dollar is spent. The scorecard is below.
  7. Put the pilot on one page. Seven lines are enough: scope, dates, deliverables, success gates, fee, account ownership and the decision date. This one-pager becomes the document both sides sign, and later the starting point for the AOR contract.

Pro tip: We recommend asking each finalist to walk through its week-two test plan during the pitch. How specific that plan is tells you more about the pilot than the credentials deck does.

One-page agency pilot agreement with seven numbered lines: scope, dates, deliverables, success gates, fee, account ownership and decision date, plus brand and agency signature lines.
The one-page agency pilot agreement: seven lines both sides sign before launch. Source: inBeat Agency

What a 90-Day Agency Pilot Timeline Looks Like

A 90-day agency pilot runs in five phases: two weeks of setup, two weeks of launch, a month of iteration, a month of scaling and scoring, then a final review in week 13.

90-day agency pilot timeline checkpoints

Weeks

Checkpoint

What to verify

1-2

Plan approved

Baseline, tracking and test plan agreed

3-4

Campaigns live

Launch happened within the agreed scope

5-8

First learning read

Tests are producing documented insights

9-12

Gates scored

Performance and learning evidence is complete

13

Final review (day-91 decision)

Decision-makers have the full pilot readout

Let’s see what the agency should be working on in each phase:

  • Weeks 1-2: Complete the account and tracking audit, finalize the written test plan and agree on the metric dictionary.
  • Weeks 3-4: Launch the first campaigns and creative tests within the agreed scope. Confirm that tracking works and each test has enough variation to produce a useful comparison.
  • Weeks 5-8: Run iteration rounds and produce the first written test readouts, including what worked, what failed and what should change next.
  • Weeks 9-12: Scale stronger performers, review holdout results where applicable and refresh creative as creative fatigue starts to affect performance.
  • Week 13: Deliver the final readout, the scale plan and the recommendation for the next stage.

Add one mid-point review in week 6. Ask for a written "on track" or "at risk" verdict against each success gate, so the final review holds no surprises.

How Much Ad Spend Does an Agency Pilot Need?

For Meta and TikTok, budget should be large enough to generate roughly 50 optimization events within the platform's learning window. Google requires a different approach based on conversion volume and bidding strategy.

Below those thresholds, the algorithm never stabilizes, and the pilot ends up measuring noise.

Here's what each platform says:

For Meta and TikTok, that gives a working formula: minimum weekly spend per ad set is roughly 50 × your target CPA. At a $40 CPA, one ad set needs about $2,000 a week.

A pilot testing three ad sets at once needs about $6,000 a week, or roughly $26,000 a month in media. This figure is derived from the platform rule, so treat it as a planning floor.

Platform learning thresholds and pilot budget implications

Platform

Learning threshold

Pilot budget implication

Meta

About 50 results in the week after the last significant edit

Budget about 50 × CPA per ad set per week, and avoid edits that reset learning

TikTok

50 conversions

Same formula; consolidate ad groups to reach it faster

Google (Smart Bidding)

1-2 conversion cycles; 15+ conversions in 30 days for Target ROAS

Allow two to four weeks before judging bid performance

Bar chart of minimum weekly Meta spend per ad set by target CPA: $1,000 at $20, $2,000 at $40, $4,000 at $80 and $7,500 at $150.
Minimum weekly spend per Meta ad set to reach about 50 optimization events, at CPAs of $20, $40, $80 and $150. Source: inBeat calculation based on Meta's learning-phase guidance

Our Meta advertising and TikTok advertising teams size every pilot around this floor before anything launches.

What to Do When the Pilot Budget Can't Reach the Learning Floor

Adjust the test design before reducing the budget further. Three options can preserve the quality of the evidence:

  • Optimize for a higher-funnel event. Add-to-cart or lead events happen more frequently than purchases, so they reach 50 a week at a lower spend.
  • Cut the number of ad sets. Two well-funded test cells beat five starved ones.
  • Extend the pilot. A 120-day pilot with a readable result is worth more than a 90-day pilot with an unreadable one.

If the pilot includes a geographic holdout, both markets need enough budget to generate a usable comparison. The market-selection section below covers that structure.

How Should an Agency Pilot Be Priced?

Pay a fixed pilot fee, kept separate from media, at or slightly above the agency's normal retainer rate. Free pilots can create weak incentives and make senior staffing harder to secure.

A pilot costs the agency more per month than a steady retainer does, because the audit, tracking fixes and account restructuring all land in the first few weeks. Heavy discounting can also make it harder to allocate senior staff to this setup work.

These are the five models you'll usually see:

Agency pilot pricing models

Model

How it works

Best for

Risk to brand

Risk to agency

Flat pilot fee

Fixed fee for the pilot period

Most pilots

Low

Absorbs setup cost

Fee credited to AOR

Pilot fee is deducted from the first retainer months if you convert

Brands likely to convert

Low

Low

Fee-at-risk

A share of the fee, typically 10-20% in our experience, tied to hitting the gates

Brands with a mature baseline

Low

Medium

Percentage of media spend

Fee scales with spend

Rarely suits a pilot

Rewards spending over learning

Low

Free or spec work

No fee

Avoid

Weak incentives; senior staff harder to secure

High

Before signing, we suggest checking these four things in the pilot quote:

  1. Is media billed separately from the agency fee?
  2. Are setup fees refundable?
  3. Is the fee fixed for the whole pilot, with no overage charges?
  4. Will the pilot fee be credited against the retainer if you sign an AOR?

Which Success Gates and Kill Criteria Should an Agency Pilot Use?

Score four success gates, performance, learning, delivery and relationship, from 1 to 5 at the go/no-go review. Then set five kill criteria that end the pilot early.

This scorecard evaluates a marketing agency using evidence from your own account. The success gates answer "Did this agency earn the contract?"

The kill criteria answer "Should we stop now?" Write both down before launch so the evaluation criteria do not change in week 12.

Below is the weighting our team typically uses as a starting framework for agency pilots:

Agency pilot success gates and weights

Success gate

Weight

What a 5 looks like

What a 1 looks like

Performance vs. baseline

35%

Beats baseline CPA, and holdout lift is positive

Worse than baseline, with no diagnosis

Learning system

25%

Documented tests, clear winners and losers, next briefs written

No test log

Delivery

20%

Every launch on time, reporting on schedule

Two or more missed launches

Relationship and transparency

20%

Full account access, senior team present, replies within the SLA

Screenshots only, team changed

How to Set Performance Gates for an Agency Pilot

Set performance gates as staged thresholds against the baseline you locked in step 2. A single end-of-pilot number doesn't work, because paid media rarely improves in a straight line.

A workable pattern is "within 10% of baseline CPA by week 8, and beating it by week 12". Where you can, add incrementality: compare the pilot market with its holdout so platform-reported gains don't carry the whole verdict. If measurement isn't set up in-house, a marketing measurement agency can help you set up the comparison before launch.

How to Set Learning Gates for an Agency Pilot

Learning gates measure whether the agency leaves you smarter than it found you. Ask for three artifacts by the final review:

  • A test log: Every hypothesis tested, the result and the decision it drove.
  • A count of hypotheses tested: A pilot that tested three ideas in 90 days hasn't built a learning loop.
  • A creative insights brief: Which angles, hooks and personas won, written so your team keeps the learning even if you don't sign.

In our experience, this is the gate that best predicts the next twelve months. Market conditions can change CPA, while a strong creative testing system keeps producing evidence for the next round of decisions. .

Which Kill Criteria Should End an Agency Pilot Early?

Kill criteria are the few signals serious enough to stop the pilot before the final review:

Agency pilot kill criteria

What you see

What it likely means

Action

Two consecutive missed launches

The account is understaffed

Stop at the next checkpoint

Metric definitions changed without sign-off

Reporting can't be relied on

Stop

No access to the live ad account

Lack of transparency

Stop

Named lead replaced without agreement

Staffing doesn't match the pitch

Escalate, then stop

Ad sets still learning-limited after week 6, with no plan

Poor account structure

Fix within 7 days, or stop

Our guide to marketing agency reporting covers what weekly reporting should include.

How to Choose Pilot Markets for a Global Agency Partner

Pilot in one representative mid-tier market, hold out a comparable market, and test the global operating model alongside local results. This combination tells you whether the agency can perform and whether it can repeat that performance across markets.

1. Pick a Representative Pilot Market

Choose the market that looks most like the rest of the rollout. The largest market may carry too much revenue risk for an initial test, while the smallest may not generate enough volume to produce a useful read.

A good pilot market has three things:

  • A known baseline, with at least 90 days of clean data.
  • Enough spend potential to clear the platform learning thresholds.
  • The language, creative and regulatory load typical of the markets that follow.
Pilot market options compared

Market option

Learning speed

Risk

How well it predicts the rollout

Verdict

Largest market

Fast

High

High

Avoid for pilots

Mid-tier representative market

Medium

Medium

High

Recommended

Smallest market

Slow, often below the learning floor

Low

Low

Avoid

Two markets (test plus holdout)

Medium

Medium

High

Best if budget allows

2. Use a Matched Holdout Market

Pair the pilot market with a comparable market the agency doesn't touch. The holdout keeps running as it did before, so seasonality, pricing changes and category shifts show up in both markets. What's left over is the agency's effect.

Match the two markets on size, seasonality and channel mix, and agree the comparison method before launch. Our guide to incrementality testing walks through geo holdouts and the calculation behind them.

3. Test a Global Agency's Operating Model

A global partner has to show that its performance travels. Local results only tell you what happened in one market. While the pilot runs, check how the agency handles:

  • Consolidated reporting: One dashboard, one currency, one metric dictionary across markets.
  • Localization workflow: How briefs, creative and copy move from one market to the next, and how long it takes.
  • Time-zone coverage: Who answers when a campaign breaks at 9 a.m. in your pilot market.
  • Local vs. hub staffing: Whether a central hub or an in-market team runs your account.
  • Network vs. bespoke structure: In H1 2026, 26% of reviewed media spend went to bespoke units run by the big holding companies. Ask which structure you're actually buying.
  • Continuity after rollout: Whether the people running the pilot will stay on the account once it expands.

Localization demand rises quickly as more markets are added . Our partnership with NielsenIQ involved 250+ content creators across 19 countries and content in 15+ languages. A pilot should show whether the agency's workflow can absorb that kind of volume before you sign it up for all of it.

How to Structure Responsibilities, Ownership, and Contract Terms in an Agency Pilot

Give the agency ownership of the test plan and execution, keep every account and asset in the brand's name, and sign a pilot contract with no auto-renewal and no notice period.

1. Define Brand and Agency Responsibilities

Brand-side delays distort pilot results as often as agency mistakes do, so write both sides' duties down:

Brand and agency responsibilities in a pilot

Area

Brand

Agency

Access

Grants partner access on day 1

Confirms tracking works

Approvals

Approves creative and budget changes within the agreed SLA

Submits changes with a clear rationale

Data

Provides baseline and finance data

Maintains the metric dictionary

Test planning

Reviews and signs off

Writes and runs the test plan and log

Budget decisions

Approves

Recommends

Reporting

Attends weekly reviews and makes the final decision

Reports weekly and presents the end-of-pilot readout

Add one clause to the one-pager: if a brand-side approval runs past the SLA, the pilot clock pauses. Otherwise the agency gets graded on weeks it couldn't use.

2. Keep Accounts and Assets Brand-Owned

Everything the pilot creates belongs to the brand from day one:

  • Ad accounts, pixels, Conversions API setup, GA4 properties and tag containers, all registered in the brand's own Business Manager or Google Ads manager account. The agency works through partner access.
  • Creative source files, including editable project files.
  • Creator content and its licences, with usage rights that survive the end of the pilot.

3. Set the Pilot Contract Terms

Five terms protect the brand in an agency pilot:

  • No auto-renewal. The pilot ends on its date unless both sides sign something new.
  • No notice period during the pilot. AOR contracts typically carry 60 to 90 days' notice. A pilot that needs notice to end has stopped being a pilot.
  • Media billed at cost. Any rebates, volume discounts or platform credits are disclosed.
  • Learnings transfer. The test log, insights brief and reports belong to the brand, signed or not.
  • A written decision date. Both sides know when the verdict happens and who makes it.

Pricing and fee credits sit with the pricing models above, so the contract only needs to reference them.

Should You Pilot Two Agencies or Test Against an Incumbent?

Split agencies by market or channel, and never put two agencies on the same audience on the same platform. Two agencies bidding into one auction raise each other's costs and argue over attribution.

The table below compares the main pilot structures and the trade-offs each one creates.

Pilot structures compared

Structure

How it works

Fair test?

Main risk

Use when

Parallel, same channel and market

Two agencies share one audience

No

Auction overlap and attribution disputes

Never

Parallel, split markets

Each agency gets a matched market

Yes

Market differences

Choosing between two finalists

Challenger vs. incumbent

The challenger takes one market or channel

Mostly

Different incentives during the test

Replacing a declining incumbent

Sequential

One agency runs, then the other

Weak

Seasonality between the two windows

Budget too small for a parallel test

Testing a challenger against the incumbent is the most common setup and the easiest to get wrong. The incumbent knows it's being tested, and its effort can shift in either direction. Give it a written brief covering the same 90 days and score both agencies on the same scorecard, so the comparison rests on identical rules.

How to Convert a Successful Pilot Into an Agency of Record (AOR)

At the day-91 decision, write the pilot's artifacts straight into the AOR: the metric dictionary, test log, named team, rate card and expansion triggers. The pilot already proved how the partnership works, so the contract should lock that in.

Five conversion terms carry the most weight:

  1. A 12-month rate lock at the rates quoted during the pilot.
  2. The pilot fee credited against the first retainer months, if you agreed that in the pricing model.
  3. Named-team continuity: the people who ran the pilot stay on the account.
  4. Market expansion triggers: for example, "add market B once market A holds its target CAC for 60 days".
  5. A performance review at month 6, scored against the same gates as the pilot.

A clean conversion also removes the need for repeated re-pitches. As the tenure data earlier in this article shows, brands that skip mandatory review cycles keep their agencies the longest.

When an Agency Pilot Is the Wrong Choice

A pilot is the wrong tool in four situations, and each has a better alternative:

  • The budget can't reach the learning floor. Run an account audit, or plan a longer pilot at a lower weekly spend.
  • The 90 days overlap your peak season. Move the window, or peak-season noise will swamp the result.
  • The work is brand or creative-platform work with a long payback. Scope it as a defined project with creative deliverables.
  • A full rebrand or replatform is underway. Wait until tracking and the site are stable, or start with an audit.

Scope Your Agency Pilot With inBeat's Paid Media Team

A good agency pilot sizes its budget to the learning floor, scores the partner on four gates, and reads results against a holdout market. This gives you evidence about how the agency thinks, tests and reports, which is the thing a credentials deck can't show you.

That's how our paid media team runs pilots. Performance creative, media buying and measurement sit together from week one, so the test log, the creative insights and the performance data all come from one team.

For brands consolidating regional agencies or evaluating a replacement for an incumbent, the pilot should answer one question: should this partnership scale?

Book a strategy call with us to review your baseline, size the pilot budget and draft the one-page agreement.

FAQs

Does a 90-day agency pilot still work when the sales cycle runs 60 days or more?

Yes, with adjusted gates. For B2B or high-consideration products, score the pilot on leading indicators such as qualified leads, cost per sales-accepted lead and pipeline created.

Then extend the performance read by another quarter. The learning and delivery gates still apply unchanged at day 91.

How do you keep the test fair when the incumbent still runs the rest of the account?

Ring-fence the test inside the ad accounts. Exclude the challenger's market or audience from every campaign the incumbent still runs, and freeze the incumbent's budget in that market for the pilot period.

Without those exclusions, retargeting and broad campaigns bleed into the test and both agencies claim the same conversions. Keep live account access on both sides so either team can check.

Should pilot creative come from the agency or from the brand?

From the agency, at least for the test cells. If the brand supplies all the creative, the pilot only grades media buying, and half of what drives paid social performance goes untested. Brand guidelines and approvals still apply.

What happens to learnings and creative if you don't convert?

They stay with the brand under the contract terms above. The step most teams miss is exporting before partner access is removed: pull the ad account's historical data, the test log and any dashboards the agency built in its own tools, because those often don't live in the brand's accounts.

Can an agency pilot run on Advantage+ or Performance Max when the algorithm controls most of the levers?

Yes. Automated campaigns shift the agency's value to the inputs the algorithm can't generate itself: creative variety, conversion signal quality and measurement design. Grade those.

Bid and targeting tweaks matter less in automated campaigns, so they shouldn't carry much of the score.

Ioana Cozma
Content Strategist & SEO Specialist

Ioana writes about growth marketing, paid media, influencer marketing, UGC, and content strategy—turning research and industry data into practical guidance for brands focused on customer acquisition, performance, and search visibility.

View LinkedIn Profile

Table of contents

Make your media budget work harder.

Certified partners

Paid Media