Blog / How Long to Test Facebook Ads

How Long to Test Facebook Ads: Most of Yours Can't Be Judged Yet

Poster: How long to test a Facebook ad? Schematic ladder of the four data floors an ad must clear before it can be judged.

Test a Facebook ad until it clears four floors: 5 delivery days, 1,000 impressions, 100 link clicks and 15 to 20 purchases. Before that, the only honest verdict is "wait". In two Shopify ad accounts we looked at in September 2026, most ads never got there. One brand had 6 of its top 50 ads by spend ready to grade. The other had 3. In the first account, the ungradable ads held about 40% of spend. Every "winner" or "loser" call on those ads was a guess.

Most advice on how long to test a Facebook ad counts days: three, seven, fourteen. Days matter less than data. A $20-a-day ad and a $500-a-day ad reach a verdict at very different speeds. So set floors for the data, not the calendar.

The numbers below come from two brands that ran a creative analysis in Drew, DataDrew's AI ads agent. Both stay anonymous. All figures are rounded.

TL;DR

  • Judge an ad only after it clears all four floors: 5+ delivery days, 1,000+ impressions, 100+ link clicks, 15–20+ purchases.
  • Most ads never clear them. One brand: 6 of its top 50 ads were gradable. Another, during a sale week: 3 of 50.
  • About 40% of spend sat in "wait" in the first account. That money is not winning or losing. It is unjudged.
  • Five verdicts, not two: Wait, Scale, Healthy, Iterate, Kill. The table is below.
  • Tired is not dead. The graded ad with the best ROAS was also the one fatiguing. Refresh it. Don't kill it.
  • One early kill is allowed: an ad that has spent 3x your target CPA with zero purchases.
The four floors before an ad gets a verdict Illustrative, no real store data. Five horizontal bars stacked top to bottom, each narrower: all ads launched, 5 or more delivery days, 1,000 or more impressions, 100 or more link clicks, 15 to 20 or more purchases. The last bar is small and blue and labeled gradable. The space removed at each step is labeled wait. Four floors before any verdict Share of ads still standing after each floor · illustrative All ads that spent this month 5+ delivery days ← wait 1,000+ impressions (CTR is readable) ← wait 100+ link clicks ← wait 15–20+ purchases = gradable In two live accounts: about 1 in 8, and about 1 in 16 Everything that drops out is "wait", not "lost"
Each floor unlocks one metric you can trust: days for delivery, impressions for CTR, clicks for landing-page conversion, purchases for CPA and ROAS. Only ads that clear all four get a verdict. In the two accounts in this post, that was a small minority of the top 50.

How long should you test a Facebook ad?

Test it until it clears four floors. Each one unlocks a metric.

Floor Minimum What it lets you judge
Delivery days5+Day-of-week swings and the 48–72 hour attribution lag have washed out
Impressions1,000+CTR and hook rate
Link clicks100+Landing-page conversion rate
Purchases15–20+CPA and ROAS

Days are the floor people know. They are the weakest one. Five days at low spend can still leave you with 400 impressions and one sale.

Example: Two ads, both with a true CPA of $40. One spends $20 a day, the other $200. The numbers are illustrative.

$20/day ad $200/day ad
Purchases per day0.55
Days to 15 purchases303
Days to 20 purchases404 (then waits for day 5)
Spend to reach 20 purchases$800$800

Both ads need the same $800 to earn a verdict. Only the calendar differs. After a "7-day test" at $20 a day, you have 3 or 4 purchases. That is not a result. Budget in purchases, not days: 20 purchases times your CPA is the real cost of a test.

Which purchase number, 15 or 20? We saw both. One account graded at 15 purchases, the other at 20. Use 15 as the minimum to read an ad at all. Use 20 before you move real budget onto it. Creative Intelligence also adds a spend floor, about $500 in account currency, so tiny tests don't get crowned.

Meta's own rule sits next to these. An ad set needs about 50 optimization events in 7 days to exit the learning phase. That rule is about the ad set's delivery. The floors above are about whether one ad's numbers mean anything.

Why isn't 5 purchases enough?

Because counts of rare events are noisy. The typical swing on a count is about its square root.

  • 5 purchases: swing of about 2. Your CPA could be off by around 45%.
  • 15 purchases: swing of about 4. Around 25%.
  • 20 purchases: swing of about 4.5. Around 22%.

At 5 purchases, a 2x ROAS ad and a 3x ROAS ad can look the same. At 20, you can start to tell them apart. Most "this ad is our winner" calls happen closer to 5.

The same ad, measured at 3, 10 and 20 purchases Illustrative, no real store data. Three horizontal bands on a CPA axis from 0 to 100 dollars. A dashed vertical line marks the true CPA of 40 dollars. The band for 3 purchases is wide and orange, the band for 10 purchases is narrower and grey, the band for 20 purchases is narrowest and blue. More purchases, narrower range Likely range of measured CPA for one ad · illustrative true CPA $40 3 purchases about $25 to $95 10 purchases about $30 to $58 20 purchases about $33 to $52 $0 $20 $40 $60 $80 $100 Measured CPA. The swing on a count is about its square root.
Same ad, same true CPA. At 3 purchases the number you see could land anywhere from about $25 to $95, so a cheap-looking ad and an expensive one are indistinguishable. By 20 purchases the range is tight enough to compare against a target.

Example: Which ad is better? Both are made up.

  • Ad A: 4 days, 3 purchases, $30 CPA ($90 spent).
  • Ad B: 7 days, 22 purchases, $38 CPA ($836 spent).

A looks cheaper. But 3 purchases can swing by about 2 either way. On $90 of spend, that puts A's real CPA anywhere from about $20 to about $70. B's 22 purchases swing by about 5, so its real CPA sits around $31 to $48.

B is the one you can act on. A goes in Wait: it misses the 5-day floor and the purchase floor. Don't pause B to fund A.

How many ads in a real account can be judged?

Fewer than you think. We call this the sufficiency rate: the share of your ads that have enough data for a verdict.

The first brand asked Drew to grade its top 50 Meta ads by 30-day spend.

  • 6 of 50 cleared the floors. About 1 in 8.
  • Of those 6: 2 Scale, 3 Healthy, 1 fatiguing, 0 dead.
  • 44 sat in Wait. They held roughly 40% of the month's spend.

That last line is the finding. Four in every ten dollars went to ads nobody could rank, scale or cut yet.

What an ad account looks like through the floors Illustrative, not a real account. A 10 by 5 grid of squares, one per ad, sorted by spend with the biggest spender top left. Five blue squares are gradable. Forty-five grey squares are Wait. Below, a spend bar split into a blue part of 55 percent and a grey part of 45 percent. Your top 50 ads, through the floors One square per ad, biggest spender top left · illustrative, not a real account 5 gradable cleared all four floors 45 in Wait no verdict yet, either way Where the month's spend went 5 gradable ads: 55% 45 ads in Wait: 45% Nearly half the budget buys ads nobody can rank, scale or cut yet
Look at the account, not one ad. A handful of big spenders earn a verdict. The long tail each gets too little money to clear the floors, but together it can hold close to half the budget. That share is what you are really deciding about.

The graded ads had one more thing in common. All six ran the same script and the same offer, across both video and static. One working message, several formats. The format wasn't the winner. The message was.

What should you do with a tired ad that still has the best ROAS?

Refresh it. Don't kill it.

The fatiguing ad in that account had the best ROAS of all six graded ads. It was also worn out. Its CTR was down about 45% from its own 7-day peak, and frequency was above 3.

Those two facts don't conflict. CTR leads, ROAS lags. The audience has stopped clicking, but people who clicked last week are still buying. Next week's ROAS will follow the CTR down.

When a tired ad becomes a fatigue call Illustrative, no real store data. Blue CTR line peaks at day 7 and slides. Orange frequency line rises. Dashed blue line marks CTR 30 percent below peak. Dashed orange line marks frequency 2.5. The area after day 18, where both lines are past their thresholds, is shaded and labeled iterate zone. Fatigue needs both signals CTR Frequency · illustrative Iterate zone if ROAS still holds CTR 30% below peak frequency 2.5 7-day peak Day 0 7 14 21 28 Days since launch CTR falling alone can be noise. Frequency rising alone is normal. Both together is fatigue.
CTR peaks early and slides while frequency climbs. The ad only counts as fatiguing once both have tripped: CTR 30%+ below its own 7-day peak and frequency 2.5+. In that zone, ROAS decides the verdict. If it holds, brief a new cut. If not, kill it.

The concept is proven. The execution is tired. Brief a new hook, a new first three seconds, or a new format for the same idea. For the signals in detail, see how to spot ad creative fatigue and what to check when a winning ad stops working.

Example: Ad C, made up. Account-average ROAS is 2.5x.

Ad C Day 7 (peak) Day 21 (now)
CTR2.0%1.1%
Frequency1.63.2
ROAS4.0x3.5x

Run the rules. CTR is down 45% from its peak, past the 30% line. Frequency is above 2.5. So it's fatiguing. The kill bar is the higher of 1.0x and 0.75 × 2.5x, which is about 1.9x. ROAS of 3.5x clears it easily. Verdict: Iterate.

So keep C running and brief two new hooks on the same message. Move budget to the new cut once it clears the floors. Don't raise C's budget: more spend on a tired ad pushes frequency up faster.

One more surprise: the two Scale verdicts went to ads with lower ROAS than the tired one. They earned Scale because they beat the account average with frequency still low and CTR holding. They had room to grow. The tired ad didn't.

Does a sale week change the answer?

It makes it worse. Sales squeeze time, so even fewer ads reach the floors.

The second brand asked which creatives worked best during a national-holiday sale week. Sales that week ran about 2.5x a normal week.

  • 3 of 50 ads cleared the floors. About 1 in 16.
  • The best sale creative had the best ROAS in the window. It was also fatiguing mid-sale: CTR down about 40% from peak, frequency near 4. Verdict: Iterate.
  • The one clean Scale verdict went to an evergreen product video. Not a sale ad.

Two lessons. A sale creative can wear out before the sale ends, so have a second cut ready by day three. And your evergreen winners often do the heavy lifting in a sale. Don't pause them to make room.

The verdict table: Wait, Scale, Healthy, Iterate, Kill

Once an ad clears the floors, it gets exactly one verdict. These are the rules Drew uses. Each ad is graded against its own history, not the account average.

Verdict Entry rule What to do
WaitBelow any of the four floors (or under ~$500 spend)Nothing. Let it run or let it go quietly. Don't call it a loser.
ScaleSpend at least 10x the account's median ad spend, ROAS at or above the account average, frequency under 2.5, CTR not fallingRaise budget in steps while there's headroom
HealthyNothing has moved enough, or signals are mixedLeave it alone
IterateFatiguing (CTR down 30%+ from its own 7-day peak and frequency 2.5+) but ROAS still holdsBrief a new execution of the same idea
KillFatiguing and ROAS below 1.0x or 0.75x the account average, whichever is higherMove the money
How an ad gets its verdict Illustrative decision tree, no store data. Top box: clears all four floors. No arrow to an orange Wait box. Yes arrow to a box asking fatiguing. Fatiguing yes leads to a ROAS check that splits into Iterate and Kill. Fatiguing no leads to a scale-rule check that splits into Scale and Healthy. How an ad gets its verdict Clears all four floors? no WAIT yes Fatiguing? CTR down 30%+ from own 7-day peak AND frequency 2.5+ yes no ROAS still holds? at or above max(1.0x, 0.75x account avg) Meets the scale rule? 10x median spend, ROAS ≥ avg, freq < 2.5 yes no yes no ITERATE KILL SCALE HEALTHY Fatigue decides "is the execution tired". ROAS decides "does the idea deserve another cut".
The first question is always about data, not performance. Only ads that clear the floors reach the fatigue check. After that, fatigue says whether the execution is tired, and ROAS decides between a new cut (Iterate) and moving the money (Kill).

Notice what's missing: there is no "kill because it looked bad on day two". Kill needs fatigue and weak economics, both on enough data.

When can you kill a Facebook ad early?

When it has spent 3x your target CPA and sold nothing.

That's the one exception to the floors, and the math backs it. If an ad's true CPA were at your target, it would average one purchase per target CPA of spend. After 3x that spend, the odds of still seeing zero purchases are about 5%. So zero sales at 3x target CPA means you can be about 95% sure the ad is worse than your target. Cut it.

Example: Your target CPA is $40. Three made-up ads, none of them near the floors.

  • $80 spent, 0 purchases. A $40-CPA ad would average 2 sales by now. The chance it shows zero anyway is about 14%, or 1 in 7. Too early to cut.
  • $120 spent, 0 purchases. That's 3x target. Expected: 3 sales. The chance of zero is about 5%, or 1 in 20. Cut it.
  • $120 spent, 1 purchase ($120 CPA). Looks awful, but a $40-CPA ad shows 1 or fewer sales at this point about 20% of the time. Not an early kill. Let it run to the floors, or cap its spend.

Everything else waits for the floors. An ad at 1x target CPA with no sale yet is normal. An ad at 2 purchases with a bad CPA is two data points.

What should you do with the money in "wait"?

Don't leave 40% of spend in limbo by accident. Decide how much you want there.

  1. Count your sufficiency rate. Take your top 50 ads by 30-day spend. How many clear all four floors? That's your number.
  2. Check the spend in Wait. If it's over a third of budget, you're testing more ads than your budget can judge.
  3. Launch fewer ads, with more budget each. Five ads that reach 20 purchases teach you more than twenty ads that reach 5.
  4. Test messages, not formats. In the first account, one message won across video and static. Put the next test budget into a new message.
  5. Grade weekly, act on verdicts only. Scale the Scales. Brief new cuts for the Iterates. Cut the Kills. Leave the rest.

For where to move money once you have verdicts, see cut or feed: ad budget allocation. For whether today's dip needs a reaction at all, see signal or noise: when to touch your ad account.

The sufficiency check (copy this)

Run it on any ad before you call it a winner or a loser.

  • Delivered on 5 or more days
  • 1,000 or more impressions
  • 100 or more link clicks
  • 15 or more purchases (20 before you raise budget)
  • Not a zero-purchase ad past 3x target CPA (if it is, cut it now)

All four boxes ticked? Grade it with the verdict table. Any box empty? The verdict is Wait.

How to run this in Drew

Doing this by hand means exporting ad-level data, filtering four columns, and computing each ad's own 7-day CTR peak. It takes about an hour, so most teams skip it.

Drew runs it as one question. Paste this:

Grade my top 50 Meta ads by 30-day spend. Only give a verdict to ads with 5+ delivery days, 1,000+ impressions, 100+ link clicks and 15+ purchases. Put the rest in Wait. Tell me how many ads are gradable, what share of spend sits in Wait, and which fatiguing ads still have the best ROAS.

The Creative Intelligence view shows the same verdicts every day, with the numbers behind each one. ROAS there is Meta-reported: a signal of creative strength, not proof of revenue.

FAQ

How long should I test a Facebook ad before turning it off?
Until it clears 5 delivery days, 1,000 impressions, 100 link clicks and 15 to 20 purchases. The one exception: turn it off early if it has spent 3x your target CPA with zero purchases.

How many conversions do I need before scaling a Facebook ad?
At least 20 purchases on the ad, plus ROAS at or above your account average, frequency under 2.5 and CTR not falling from its peak. Meta also wants about 50 optimization events in 7 days at the ad set level to exit learning.

When should I kill a Facebook ad and when should I scale it?
Kill it when it is fatiguing and its ROAS is below 1.0x or 0.75x your account average, whichever is higher. Scale it when it has taken 10x your median ad spend, beats your account-average ROAS, and still has frequency under 2.5. Below the data floors, do neither.

How do I scale Facebook ads without killing efficiency?
Scale only ads with a Scale verdict, in 10 to 15% budget steps every 2 to 3 days. Keep testing in a separate campaign. Watch frequency: once it passes 2.5 and CTR starts falling, the ad moves to Iterate, not more budget.

Is 1,000 impressions enough to judge a Facebook ad?
It's enough to judge CTR. It's not enough to judge ROAS. For that you need purchases, and 15 to 20 is the minimum for a number you can act on.

What is a sufficiency rate?
The share of your ads that have enough data to get a verdict. In two accounts we looked at, it was about 1 in 8 and about 1 in 16 of the top 50 ads by spend. A low rate means you're running more ads than your budget can judge.

Should I kill an ad with the best ROAS if it's fatiguing?
No. High ROAS plus falling CTR means the idea works and the execution is tired. Brief a new hook or format for the same message and let the new cut take over.

DD
John Abhishek Head of Growth @ Datadrew.

Take CONTROL of your Shopify growth today