A creative testing system for mobile app and game UA is the loop that turns a batch of new ads into a winner you can trust, then turns that winner into the next batch. I build and run that loop inside the campaigns I buy, so the person who reads the spend is the person who writes the next version.
This page is for a growth or UA lead whose creative testing feels random, or whose team ships a lot of ads and learns little. It says what the system is, what it needs, and where UGC and AI footage fit. The first stage of the method is written out, and ad production makes the ads; this page is the system that decides what gets made next.
Platform facts on this page were checked against Google, Meta and TikTok documentation on 23 and 24 September 2026.
Why does creative testing feel random?
Creative testing for mobile apps and games feels random when a team compares its ads inside a live campaign and the platform, not the test, decides who sees what. Meta, Google and TikTok route budget to the ads their models expect to perform, so the early leader takes the spend and the rest never get a fair sample. What comes back is a ranking of delivery, not of ideas. Change three things per version and nothing carries to the next round. The randomness is in the setup.
Gordon, Zettelmeyer, Bhargava and Chapsky compared fifteen Facebook experiments with observational estimates and found the estimates often failed to reproduce the experiment, so a dashboard winner is not automatically a real one. That is why the first read of a large batch is only a screen: the elimination round starts every ad together in a dedicated install campaign and reads spend, and the verdict waits for revenue. When I removed the five ads Meta kept scaling, the algorithm surfaced six new creatives above a target those five never reached.
Why are we not iterating fast enough on creative?
A mobile app or game UA team is usually not iterating fast enough on ad creative because the loop has a handoff in it. The person who sees an ad tire writes a brief, the brief waits for a studio, the studio guesses from a portfolio, and the next version arrives after the winner has faded. Creative supply is the constraint, and the bottleneck inside it is deciding what to make next, not production speed. Fast iteration means the person who reads the spend allocation writes the winner's next version, with no brief or account manager in between.
That is the shape I run: each round ends in a written read, the next ideas come from it, winners become versions and losers are not made again. Ad production makes test versions first and finishes only the winner. On Horse Racing Solitaire 175 ads went through one account in five months, fourteen under £1.00 per install. The chart below shows what Meta did with the budget when an outside studio's ads joined with a brief but not the account.
How do you build a creative testing system?
A creative testing system for app and game UA has five parts. Two levels of test: concepts that differ a lot, judged as whole ideas on purchases, and versions of one concept where one named variable changes against a control, so a win teaches something. A structure that gives every ad the same start: test versions in a dedicated install campaign, every ad launched together, spend read as the screening signal, budgets untouched. A decision rule written before launch: the metric, the spending limit, the review window, what counts as inconclusive. A record of every ad with the variable it tested and the decision. And a production arm that remakes only the winners.
The first three are how I run accounts with more versions than purchases can validate; smaller batches go straight to purchase testing, and the elimination round is written out with the budget math. Most teams skip the record: the round logs each ad's spend, installs, early behavior and decision, and the log also needs the variable that changed and the next hypothesis. The rule comes first: Johari, Koomen, Pekelis and Walsh showed that checking a test repeatedly and stopping at the first significant read inflates false positives. I read the verdict on revenue, since the same IPM and CPI can hide a 5% or a 20% install to purchase rate.
| Platform | What it offers | How to read it |
|---|---|---|
| Meta | A/B test: one variable, a measurable hypothesis, an audience no other campaign uses, 7 to 30 days | Delivery inside each group is still optimized: which ad performs better as Meta delivers it |
| TikTok | Split Test: hypothesis first, one meaningfully different variable, at least 80% estimated power, 7 to 30 days, no edits during the run | Neither Smart Creative nor Automate Creative supports split tests; Smart Creative pauses ads it classifies as fatigued |
| Google App campaigns | Directional experiment: one attribute per test for a clear read, limited to certain App campaign setups, default end date of 14 days, no fixed budget or duration rule | Assets are combined into ads and rated relative to each other, not isolated exposures |
Where do UGC and AIGC fit in a creative testing system?
UGC and AIGC are formats to test inside a creative testing system for apps and games, not answers to it. For an app the hook is usually someone naming the problem in their own words before anything gets shown. Neither wins by default, and the account read decides the mix.
At EatBetter, creator content run as paid was the engine behind $10K to $180K in monthly recurring revenue in two months. See UGC ads for apps and AI UGC ads. The table sets out the job each format does in the loop and what the evidence behind it measures.
| Format | Its job in the loop | Where it leads | What the evidence measures |
|---|---|---|---|
| AI footage | Finds the claim: with no shoot behind it, puts three opening lines in front of a real audience before anyone films | Testing a hook before a shoot | Early, and engagement only: personalized generative video raised it among one retailer's existing customers, without separating personalization from AI, and measured nothing about installs (Kapoor and Kumar) |
| Creator on camera | Films the claim that won on their own phone, which reads as a post rather than an ad | When the reaction sells the product | Engagement and purchase intent, lifted on average across 135 experiments and varying by creator, product and platform (Barari, Eisend and Jain) |
| Gameplay or app capture | The workhorse for a game, cut with captions and motion graphics | Wherever the product sells itself | Not covered by the studies cited here |
What does a specialist who builds creative testing systems do?
A specialist who builds creative testing systems for mobile apps and games sets up the loop, runs it, and leaves it running. That means writing concepts from what the account has rewarded, deciding which variable each version tests, structuring the test so every ad gets the same start, reading spend rather than vanity metrics, killing losers before they eat budget, and writing the read the next round starts from. It also means training the internal team, because a system in one head is not a system. Producing every asset is a different job.
I have done this inside a UA retainer, where the ads are tested in the campaigns I run; at Deutsche Telekom, where I built the internal creative testing process and trained the team to run UA independently; and through ad production, where the read comes back with the ads. Creative supply is the binding constraint in most accounts I open, which is why the growth audit reads what has been tested, what shipped untested and whether spend or vanity metrics are picking the winners.
Why buy creative testing plus performance marketing from one person?
Creative testing plus performance marketing from one person closes the handoff where most creative learning is lost. The person who reads the spend allocation writes the next concept. The person who set the conversion signal knows why a cheap install ad lost to a more expensive one that brought payers. Split those jobs across a studio, an agency and a buyer and each hands the others a document that is always older than the data. One owner is faster and the read is honest, because the same person answers for the ROAS.
On a retainer I buy the media, own the signal and run the creative loop, and my network makes the ads when production is in scope. On one subscription account, the concept and its variant explained 25.2% of the variation in ROAS across 917 ads, and language 4.8%, a read you only get when creative record and revenue sit with the same person. Signal engineering comes first, because a testing loop on broken data finds the wrong winners faster; paid UA for mobile games and for subscription apps say how the loop differs by product.
Who it is for
- Subscription apps and mobile games spending around $100K a month on paid UA, or funded to get there
- Teams that ship a lot of creative and cannot say which ad won or why
- Games whose top ad has fatigued and whose next batch has no test plan
- Teams that have used an external studio for creative and want to compare it against a testing system
What you get
- The testing loop set up in your ad account: test versions in a dedicated install campaign, every ad started together, spend deciding, winners remade properly
- Creative briefs and asset strategy owned by me, with the ads made by my network when you add creative production
- A written read at the end of every round: what won, what lost, what to make next
- Weekly numbers against targets, and creative iteration rate as one of the measures I hold myself to
- Documentation your team keeps, and the training to run the process without me
Questions people ask
Why does our creative testing feel random?
Creative testing feels random when you compare ads inside a campaign the platform is optimizing, so the early leader takes the spend and the rest never get a fair sample. Fix the setup: a dedicated install campaign, every ad started together, spend as the screening signal, the decision rule written before launch. The elimination round is that setup.
Why are we not iterating fast enough on ad creative?
A UA team usually iterates slowly because a handoff sits between whoever sees the numbers and whoever makes the next ad, and the brief ages while it waits. I write the next version from the last round's read, with no brief or account manager in between. 175 ads through one account in five months is that pace.
How do you build a creative testing system?
Separate concept tests from version tests, test in a structure that gives every ad the same start, write the metric and the spending limit before launch, log every ad with the variable it tested and the decision, and remake only the winners. The platforms' test tools help; none replaces a hypothesis or a verdict read on revenue.
Do UGC and AIGC belong in a creative testing system?
Yes, as formats to test, not as answers. AI footage tests which claim lands before anyone films, a creator on camera tends to scale when the reaction sells the product, and gameplay or app capture leads where the product sells itself. What AI UGC ads are good for is its own page.
What does a specialist who builds creative testing systems do?
Reads the account, writes concepts from what it has rewarded, decides what each version tests, structures the test, reads spend, kills losers, writes the read the next round starts from, and trains your team to run it. At Deutsche Telekom that meant building the internal creative testing process and handing it over.
Should one person do creative testing plus performance marketing?
At the spend levels I work at, yes. The spend allocation, the conversion signal and the budget decision feed the next concept, and splitting them across a studio, an agency and a buyer puts a document between each. What a fractional Head of UA owns lists channels, measurement and creative testing as one job.
How long before you can call a creative winner?
The elimination round reads after a few days of delivery or at a spending limit set before launch, whichever comes first. Whether the winner brings payers takes as long as your revenue takes to mature, and how many purchases a ROAS needs before it means anything depends on the account.
Sources
- What are best practices for A/B tests for Meta ads? Meta Business Help Center. Test one variable, start from a measurable hypothesis, use an audience no other running campaign uses, and run at least 7 days and at most 30. Read in a browser on 24 September 2026.
- Split Test Best Practices TikTok Business Help Center, last updated January 2026. Hypothesis first, one meaningfully different variable, at least 80% estimated power, 7 to 30 days, no edits during the run. Product recommendations, not statistical constants. Checked 23 September 2026.
- About Smart Creative TikTok Business Help Center, last updated November 2025. Pauses creatives it classifies as fatigued and does not support split tests. The fatigue classification is not independently validated. Checked 23 September 2026.
- About Automate Creative TikTok Business Help Center, last updated September 2026. Symphony AI variations of your assets; split testing not supported. Checked 23 September 2026.
- Set up a directional experiment for App campaigns Google Ads Help. One attribute per test when a clear read is wanted, a default end date of 14 days, no strict universal budget or duration rule, and the leading group can change while the test runs. Currently scoped to particular App campaign and video setups. No update date shown. Checked 23 September 2026.
- About asset reporting for App campaigns Google Ads Help. An ad can contain several assets and the ratings are relative, so asset rows are not isolated exposure tests. Checked 23 September 2026.
- A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook Gordon, Zettelmeyer, Bhargava and Chapsky, Marketing Science, 2019. Fifteen US Facebook experiments, 500 million user observations and 1.6 billion impressions. Two authors were Facebook employees; not app specific.
- Always Valid Inference: Continuous Monitoring of A/B Tests Johari, Koomen, Pekelis and Walsh, Operations Research, 2022. Checking a test with a fixed horizon repeatedly and stopping at significance inflates false positives. General web experiments; funded by Optimizely, which implemented the method.
- A meta-analysis of the effectiveness of social media influencers: Mechanisms and moderation Barari, Eisend and Jain, Journal of the Academy of Marketing Science, online 2025. 71 papers, 135 experimental studies, 571 effect sizes. Outcomes are engagement and purchase intent, not app installs or LTV, and effects vary by creator, product and platform.
- Frontiers: Generative AI and Personalized Video Advertisements Kapoor and Kumar, Marketing Science, 2025. A WhatsApp field experiment with 21,328 existing customers of one Indian retailer; personalized generative video raised engagement by roughly 6 to 9 percentage points. Personalization, video and AI authorship are not separately identified, and the outcome is engagement, not installs or revenue.