An AI media buyer should not be judged by whether its dashboard sounds confident. Fluency is the easiest thing for a language model to produce and the least correlated with being right about an ad account.
The things worth testing are narrower: does it identify verified facts, does it separate evidence from hypothesis, does it catch measurement problems before blaming creative, does it quantify the business impact, does it recommend a reversible next action, and does it admit when the data cannot support a conclusion. Those six behaviors can be graded. Below is a protocol for grading them.
One thing to state before the design: the 216 grades describe a proposed benchmark, not a completed AdRiseLab evaluation. Nothing here should be read as a published result for any system, ours included.
Why this category needs a harder test
Meta's advertising system is not a formula an outside tool can reconstruct. Meta describes its GEM ads foundation model as a hybrid architecture with trillions of sparse parameters and billions of dense parameters, drawing on ad content, engagement, activity history, location and creative-representation signals. Its Adaptive Ranking Model varies inference complexity according to user context and intent; Meta reported a 3% conversion increase and a 5% click-through increase for targeted users after that model launched on Instagram in Q4 2025.
Those are platform-level descriptions. They do not let any external system reconstruct Meta's internal decision for one ad on one impression. An honest analysis tool has to reason from the advertiser's observable data and then say plainly what remains unknowable - which is precisely the behavior a benchmark has to be built to detect, because it is also the behavior that sounds least impressive in a demo. Sources: Meta Engineering on GEM and Meta Engineering on the Adaptive Ranking Model.
Why an AI judge is not enough
The Creativity Benchmark research preprint used 678 practicing creatives and 11,012 anonymous pairwise comparisons. Its automated judge setups showed weak, inconsistent alignment with the human rankings, along with judge-specific biases.
That study examined marketing creativity rather than account diagnosis, so the finding does not transfer directly. The methodological lesson does: do not let one model both produce and certify the answer. Use human domain graders, and report how much they agreed with each other. Source: Creativity Benchmark research preprint.
The 24-case design
| Case family | Cases | Required variation |
|---|---|---|
| Delivery and auction | 6 | Low delivery, fragmentation, audience or placement shifts |
| Creative performance | 6 | Early failure, stable winner, spend concentration, suspected fatigue |
| Funnel and measurement | 6 | Event duplication, value mismatch, page-CVR decline, attribution conflict |
| Business economics | 6 | Margin pressure, AOV change, lead quality, refund or sales-cycle effects |
Vary spend levels and business models across the set, but give every system the identical packet for a given case. And do not build the benchmark only from obvious failures. It has to include cases where no action is the correct answer, cases with missing data that require abstention, cases with several plausible explanations, cases where platform metrics improve while business profit worsens, and cases where the apparent creative problem is really a measurement or funnel failure.
That last category is the one that separates systems. A tool that reaches for a creative refresh whenever ROAS falls will pass a benchmark made of creative problems and fail an account.
The standard input packet
Each case should carry fourteen complete recent days and the fourteen complete days before them; daily campaign, ad-set, ad, country and placement metrics; the optimization event and attribution setting; event definitions and known measurement limitations; backend orders or qualified-lead outcomes; revenue, contribution margin, refunds or lead-value assumptions; creative launch dates with a concept taxonomy; a change log covering budget, bid, audience, creative, landing page and offer; and an explicit list of the fields that are missing.
Strip personal data and advertiser identity. Keep metric labels and definitions intact, because half of real diagnosis is noticing that two metrics do not mean what their names suggest.
The three systems and one output template
For each case, compare an AI media-buyer system, an experienced human analyst and a deterministic rules baseline. Normalize every output to the same template: verified facts, diagnosis, alternative explanations, recommended action, expected signal, risk and rollback, confidence, missing information.
Randomize and relabel the outputs before grading. Perfect blinding is probably impossible - writing styles differ, and graders who work in this field will guess sometimes - so disclose that limitation rather than claiming a clean blind.
The 100-point rubric
| Dimension | Weight | What earns full credit |
|---|---|---|
| Factual and arithmetic accuracy | 25 | Every cited number and calculation matches the packet |
| Causal discipline | 20 | Observations, hypotheses and causal claims are separated |
| Recommendation specificity | 20 | Action, scope, owner, timing and decision signal are clear |
| Measurement and economics | 15 | Backend value, attribution, margin and event quality are considered |
| Safety and reversibility | 10 | No uncontrolled spend or destructive account action |
| Clarity and prioritization | 10 | The most material issue appears first |
Then add fatal-error gates that override the total. Regardless of score, fail any output that invents a metric or account fact, recommends uncontrolled budget or bid changes outside the test authority, ignores a known duplicated purchase event, treats correlation as proven causation, exposes personal or confidential information, or claims guaranteed performance.
The gates exist because weighted averages forgive the errors that matter most. An output can score 88 on presentation and still have hallucinated the CPM it built its recommendation on.
Pre-declare the pass criteria
These are proposed thresholds, not existing industry benchmarks:
| Criterion | Proposed threshold |
|---|---|
| Median overall score | At least 75 out of 100 |
| Material arithmetic error rate | Below 2% of factual claims |
| Unsafe recommendation count | 0 |
| Hallucinated account facts | 0 |
| Correct abstention on insufficient-data cases | At least 80% |
| Recommendation agreement with adjudicated action class | At least 80% |
Publish the thresholds before anyone reads a result. A threshold set afterwards will be set wherever the preferred system happened to land, and everyone involved will believe they chose it honestly.
Worked case: a ROAS decline
This case is modeled, not a customer result.
| Driver | Prior period | Recent period | Change |
|---|---|---|---|
| CPM | $12 | $14 | +16.7% |
| CTR | 1.20% | 1.00% | -16.7% |
| Landing-page CVR | 3.00% | 2.70% | -10.0% |
| AOV | $80 | $76 | -5.0% |
| Modeled ROAS | 2.40 | 1.47 | -38.9% |
A weak diagnosis reads: ROAS fell, so refresh the creative. It is fluent, it is actionable, and it is wrong about three of the four drivers.
A stronger one decomposes the change: the CTR decline is the largest demand-side contributor; the higher CPM adds a substantial auction-cost drag; the lower CVR and AOV show the problem is not confined to the ad at all. Creative work may address CTR, but no creative refresh repairs checkout CVR or product mix. The recommendation that follows is to validate tracking and mix shifts first, then assign separate owners and separate tests to creative, landing experience and economics.
The second answer is better not because it is longer but because it refuses a one-cause story that the arithmetic does not support. The full decomposition method covers the equation behind this case.
Measure calibration, not only correctness
Ask each system to state a confidence from 0% to 100%, then group answers by confidence band and compare stated confidence against adjudicated correctness. A system that says it is 95% certain on genuinely ambiguous cases is less trustworthy than one that abstains, even if their raw accuracy matches.
Report accuracy by confidence band, overconfidence rate, correct abstention rate, error severity and grader agreement. For ordinal rubric scores, publish an appropriate inter-rater reliability statistic alongside the raw distribution. One mean hides the disagreement, and the disagreement is usually the most informative part of the result.
What a product claim may say afterwards
Defensible after a completed test: that in a preregistered 24-case evaluation the system's median blind score was X under the published rubric; that it abstained correctly in X of Y insufficient-data cases; that the benchmark rested on 216 independent grade sheets.
Not defensible, before or after: that the system always finds the cause, that it replaces media buyers, that it guarantees lower CPA, or that it understands Meta's algorithm. No blind test can establish any of those, and a vendor willing to claim them has told you what its benchmark was for.
Where AdRiseLab sits in this
AdRiseLab's AI Media Buyer analyzes connected Meta account data, supports creative and competitor research, monitors fatigue and organizes reporting. It can apply supported budget or status changes after explicit approval, and it creates campaigns in Meta once the user confirms the launch. Budget, bidding and campaign-structure decisions stay with the user, and the product is Meta-only - Facebook and Instagram through a connected ad account.
That approval boundary belongs inside the benchmark rather than outside it. A system that cannot act without confirmation should be graded on the quality of what it proposes and the reversibility of what it asks for, which is exactly what the safety dimension and the fatal-error gates above are measuring.
Related Reading
For the honest version of the category question, can AI replace your media buyer in 2026 works through what is and is not automatable. What is an AI performance marketer defines the category and its current limits, and the 12-point Meta ads audit is the human framework these cases are built to test against.
