Most new Meta ad creatives receive very little delivery. That is not unique to your account, and it is not evidence that the system has taken a dislike to one particular ad. Spend concentrates into a small part of the creative portfolio, and most of what a team launches stays close to zero.
Two independent 2026 datasets, built from different samples with different definitions, found the same broad shape.
What the two datasets observed
| Study | Scope | Selected delivery finding |
|---|---|---|
| Motion, Creative Benchmarks 2026 | 578,750 creatives, 6,015 accounts, $1.29B Meta spend | Roughly half received no or minimal spend; around 6% carried the majority within accounts |
| Interconnections, Creative Fatigue Benchmark 2026 | 368 ads, 15 DTC brands, 48,450,354 impressions, $685,791 spend | 80% stayed below 100,000 lifetime impressions; the top 10% carried 68% of spend |
The two studies use different samples, windows, definitions and business mixes, so they should not be pooled into one benchmark. Their shared value is the distribution pattern, not any single percentage. Both publish their methodology: Motion Creative Benchmarks 2026 and Interconnections: The 2026 Meta Creative Fatigue Benchmark.
The smaller study shows the long tail clearly
Interconnections tracked every included creative from its first delivery, which makes the tail visible in a way account averages never are.
| Percentile | Active days | Lifetime impressions |
|---|---|---|
| 25th | 6 | 524 |
| 50th | 18 | 10,665 |
| 75th | 46 | 74,418 |
| 90th | 76 | 293,876 |
The same sample produced four summary figures worth keeping:
- 80% of creatives never reached 100,000 impressions, and 89% never reached 250,000.
- The top 10% of creatives took 68% of spend.
- The top 20% took 84%.
- The entire bottom half accounted for about 1% of spend.
These are observations from 15 managed DTC brands between April 8 and August 11, 2026. They are not a platform-wide law, and they may not transfer to B2B, apps, lead generation or a different budget tier.
Why concentration is plausible in Meta's system
Meta's own engineering material describes a large recommendation stack that predicts click and conversion outcomes from many signals at once. Its GEM foundation model carries trillions of sparse parameters and billions of dense ones, and learns from ad creative representations alongside user and engagement features. The Adaptive Ranking Model then routes requests to models of different complexity depending on context and intent, all inside sub-second latency.
None of that explains why one advertiser's ad took $7 while another took $700. What it does establish is that "Meta evenly tests every creative" is the wrong mental model. Delivery is a ranking and allocation process, not a laboratory that guarantees equal sample sizes. The Meta ads auction explainer covers the mechanics underneath.
Sources: Meta Engineering on GEM and Meta Engineering on the Adaptive Ranking Model.
Low delivery is a symptom, not a diagnosis
| Observation | Plausible issue | What to inspect next |
|---|---|---|
| Almost no impressions | Eligibility, bid, budget, audience, schedule, learning signal, competition | Delivery status, diagnostics, structure, optimization event |
| Impressions but weak attention | Hook, visual, relevance, placement fit | Placement-level views, CTR, message clarity |
| Clicks but weak landing behaviour | Message mismatch, page speed, offer, traffic quality | Landing-page views, engagement, page CVR |
| Conversions but weak economics | AOV, margin, lead quality, refunds, attribution | Backend revenue and contribution |
| One ad takes nearly all spend | Predicted performance or allocation concentration | Compare business outcomes; do not assume equal testing |
The right next action depends on which stage actually has evidence behind it. Answering an impression problem with a creative brief is the most common way a team spends a month on the wrong thing.
Five habits that make starvation worse
1. Launching correlated variants
Ten ads built from the same hook, proof and visual compete as one idea family, not ten ideas. Real difference means a different problem, promise, mechanism or proof - not a different crop, colour or adjective.
2. Optimizing for an event with too little signal
If the chosen optimization event rarely occurs, the account gives the system very little feedback to act on. That is not an argument for automatically switching to a shallower event: the event still has to represent business value. It is an argument for knowing how many of them you actually generate per week before you judge a creative on them.
3. Fragmenting the budget
Many campaigns, ad sets, audiences, countries and creatives can divide a modest budget into cells that never become informative. Every extra cell is another place for evidence to fail to accumulate.
4. Editing before evidence accumulates
Frequent mid-flight changes destroy the comparison window and make cause attribution close to impossible afterwards. Keep a change log and pre-declare review points instead.
5. Reading low spend as a creative verdict
A low-delivery ad may genuinely be weak. It may also simply be unevaluated. Collapsing those two cases is how teams end up rewriting ads that were never seen.
A probability model for creative portfolios
Suppose, for planning only, that each genuinely distinct concept has a constant 5% chance of becoming a winner and that outcomes are independent. The probability of at least one winner after n concepts is 1 minus 0.95 to the power of n.
| Distinct concepts tested | Modeled probability of at least one winner | Expected winners |
|---|---|---|
| 5 | 22.6% | 0.25 |
| 10 | 40.1% | 0.50 |
| 20 | 64.2% | 1.00 |
| 32 | 80.6% | 1.60 |
| 45 | 90.1% | 2.25 |
| 59 | 95.2% | 2.95 |
Both assumptions are wrong in practice. Concepts are correlated, allocation is adaptive, and winner definitions differ by account. Correlation makes the table optimistic whenever the "different" ads are small edits of one idea. The model is not a prescription to run 45 ads. It explains why rare outcomes combined with low concept volume produce long stretches with no outlier - and why that stretch is not, on its own, proof that anything is broken. The creative testing budget calculator puts a defensible cost against each of those tests.
What to report each week
Report delivery by launch cohort rather than by account average.
| Metric | Definition |
|---|---|
| Launch cohort | Creatives first launched in the same week |
| Delivery coverage | Share crossing the account's minimum impression or spend threshold |
| Spend concentration | Share of cohort spend held by the top 10% and top 20% |
| Concept diversity | Distinct hypotheses / total creatives |
| Qualified hit rate | Economics-qualified winners / evaluable concepts |
| Unevaluated rate | Creatives that never crossed the evidence threshold / launched concepts |
That single view separates creative failure from evaluation failure, which is the distinction most weekly reports collapse into one number. Connected-account analytics can support the review; the interpretation and the decision stay with the advertiser.
Related Reading
Once an ad has earned real delivery and then declines, the question changes: creative fatigue or creative failure separates wear-out from an ad that never worked. For what the concentrated winners actually looked like, see what 578,750 creatives reveal about hooks and formats. To size the portfolio in the first place, how many Meta ad creatives you need works from budget rather than from a round number. And the production side of the same problem is covered in the creative operations benchmark.
