Creative testing · Part 1 of 3

How do you set up Meta ads for creative testing?

An optional install screening stage for teams with more creative variants than they can afford to validate on purchases. In my accounts a version needs about 10 purchases to show potential and about 50 for a verdict, and at a normal cost per purchase that adds up fast across many variants. Use it to nominate candidates, then validate their purchase performance. Smaller batches can go straight to purchase testing.

Where this sits in the flow

Validation happens in the purchase campaign, scaling in the main campaign. Screening on cheap signal and validating on purchases is an established mobile UA method.12

Stage 1this part

Elimination

Install campaign, new
  • Close versions of the same concept compete here
  • Signal arrives within days: CPI and early behavior
  • Goal: nominate variants for purchase validation
the ads that pull away advance, usually 1 to 2
Stage 2

Validation

Purchase campaign
  • Screened candidates or a small batch of new ads enter
  • Set a budget, evaluation window and purchase cost target
  • Judge evidence against the decision, not a universal event count
sufficient purchase evidence + cost below your target
Stage 3

Scaling

Main campaign
  • Ads below the cost line move to the main campaign
  • Set the cost line separately for your campaign
1
Prep
tool + platform + market
2
Setup
structure + budget
3
Launch
synced race
4
Read
watch, hands off
5
Exit
ads that pull away go to validation

Campaign structure

Example: 17 creatives, 4 concepts, max 5 versions per ad set. Budgets assume $1 CPI. Compare versions within each concept to nominate candidates for purchase validation. Spending across ad sets describes install delivery; it does not rank concepts by purchase potential.

INSTALL CAMPAIGNCBO · e.g. $80/dayConcept Amin $10/day5 versionsConcept Bmin $10/day4 versionsConcept Cmin $10/day5 versionsConcept Dmin $10/day3 versionsEXAMPLE4 ad sets × min $10 = $40campaign = $40 × 2 = $80

Budget math

Enter your own numbers. The ad set count is a lower bound. If you need more sets, calculate each concept's groups of up to 5 separately. My planning defaults are 10 installs per day per ad set and a campaign budget at 2x the sum of minimums. These are operating heuristics, not evidence thresholds. Installs and spend are not guaranteed for each ad. Round cost excludes purchase validation and creative production; set a separate total spending limit before launch.

Ad sets (max 5 versions per set)4
17 / 5, rounded up
Min ad set budget$10.00
$1.00 CPI × 10 installs
Campaign budget (daily)$80.00
2 × 4 × $10.00
Estimated round cost~$240
$80.00 × 3 days

Prep

Connect the bulk upload tool
Get a trial at adsuploader.com and follow the tutorial to connect ads manager. Skip uploading videos one by one.
Pick the platform
If you run web2app you can test directly on iOS. Without web2app on iOS, run the tests on Android. Meta's iOS 14+ limits get in the way here: up to 24 campaigns per app, one ad set per Advantage+ app campaign with no manual setup, and a 24-72 hour postback delay.1314 Android has no such ceiling and the data lands the same day.
Fix the market
Pick one T2 country where English is spoken and use the same country every round so rounds stay comparable. With several countries in one ad set, CBO pushes budget to the cheapest CPM country, so you end up comparing markets instead of creatives. This round only cuts weak versions. What survives gets validated in the markets you plan to scale in, usually Tier 1.3111

Setup

Open an install campaign
A new install campaign with campaign budget optimization (CBO). Meta routes the budget in real time to the ad sets it sees the best opportunity in.3
1 concept = 1 ad set
Each ad set holds at most 5 versions of a single concept. 17 creatives from 4 concepts means 4 ad sets (5+4+5+3). Meta's own guidance also says 6 or fewer ads per ad set.45
Keep targeting broad
No interest or lookalike constraints on the test ad sets. Let the creative decide who sees it, since adding constraints means you are testing targeting, not creative.153
Leave related media empty
Do not add related media to the ad. When Meta blends in related or recommended assets alongside your version, delivery splits between your creative and the blend, and the ad set's spend no longer tells you which version is actually winning.
Name by convention
The ad set carries the concept name (e.g. Duet-0926). Ad name: ConceptName_v2_XY_0926, meaning concept, variant number, creator initials, month and year.6
Set ad set minimums
Plan an ad set minimum around 10 installs a day at your expected CPI: $5 at $0.50 CPI, or $10 at $1 CPI. This is my delivery heuristic, not a minimum sample for a reliable verdict. Ad set minimums constrain allocation and do not guarantee installs or equal delivery to individual ads. Check actual delivery.3
Set the campaign budget
My starting budget is 2x the sum of ad set minimums. Four $10 minimums total $40; an $80 daily campaign budget leaves another $40 for flexible allocation. The minimums sit inside the $80, not on top. The multiplier is a planning choice, not a Meta requirement.

Launch

Upload
Open the campaign and one ad set manually, then upload the ads with adsuploader.
Set the sync rule
In Ads Manager select the 4 ad sets, then Rules → Create a new rule → Custom rule. Apply rule to: Active ads in these ad sets. Action: Turn off ads. Condition: Spent > $0.10, time range Maximum (Lifetime). Schedule: Continuously. Meta runs rules roughly every 30 minutes, so some ads stop a few cents past $0.10, which is fine, the goal is sync.78
Restart everything at once
When every ad has spent $0.10 and paused, delete the rule and reactivate all paused ads at the same time. The race starts equal for everyone. The ad approved first grabbing the early budget is a known problem, and a simultaneous start is the known fix. You will not find this equalizer in Meta's official docs, it is settled field practice.910

Read

Watch without touching
After the restart, wait for the spend split to separate without touching budgets or ads.
Use a review window and a spending limit
For the example below, review after 3 full days of delivery and stop screening at $240 total spend, whichever comes first. Choose your own limits before launch and monitor actual spend. These are operating limits, not statistical thresholds. At the limit, record advance, inconclusive or deprioritize; do not keep spending just to force a winner.
Spend nominates a candidate
Read allocation within each concept. Sustained spend concentration across daily reads can nominate 1-2 variants for purchase validation. It describes delivery under install optimization, not proven purchase value or a controlled A/B result. Little spend means limited evidence. If nothing separates, record inconclusive; do not eliminate the concept.3
CPI is a sanity check only
Check CPI and a relevant early event, such as starting the game or completing onboarding, using comparable reporting windows. CTR and CPM sit outside this read. Both swing with placement and country inside the same campaign, so they compare noise more than creative. A low CPI with weak activation can signal a mismatch between the ad and product. Hold the decision if quality is poor or reporting is incomplete. Strong installs still require purchase validation.

Reading happens at two levels

Illustrative allocation after 3 days, $240 total. Concept A received the most install spend; that does not make it the best purchase concept. Within A, v1 is a candidate for validation. The cumulative chart alone does not show whether the pattern persisted across daily reads.

1. Concept level
4 ad sets, cumulative spend · dashed line: planned minimum of $30 per set over 3 days
$0$55$110$100Concept A$63Concept B$44Concept C$33Concept D
2. Version level
inside Concept A, 5 versions
$0$30$60$55v1CPI $0.31$16v2CPI $0.36$12v3CPI $0.34$9.50v4CPI $0.41$7.50v5CPI $0.52

Exceptions

  • There can be exceptions. Cheap CPI does not always mean the best purchases, or a new concept may not be able to race the veteran winner variations you have optimized for weeks. In those cases judge each ad set's spend within itself; the ones that stand out can still go to validation, where the final call is made.121
  • An ad can get stuck in review. If the rest have stopped at $0.10 while one is still pending, whether to wait is the team's call; you can restart without it and roll that ad into the next round.
  • An ad can get approved and start spending before the sync rule catches it. Since the rule checks roughly every 30 minutes, not continuously, it can land at $5 or $10 instead of stopping at $0.10. You can stop it manually if you're watching, or trust the rule to shut it off once it runs.

Exit

The ads that pull away go to validation
Advance nominated candidates into purchase validation in the intended market and platform. Log each ad's spend, installs, early behavior, decision and reason. Keep inconclusive ads eligible for a smaller retest or direct purchase validation. Deprioritized means paused for now, not proven unable to sell.
Round timing
When to start a new round is your call, open one when the creatives are ready.

Step 2, Validation: evaluate nominated candidates on purchases against your cost target, with a defined budget and reporting window. Covered in a separate part.

Sources

Background on platform mechanics and creative testing. The budgets, review windows and decision examples above are operating heuristics, not thresholds established by these sources.