Take two install weeks from the same campaign, $10,000 of spend and 2,000 installs each. The older week shows 78% ROAS to date and the younger one 57%. A dashboard that stops there reads the younger week as the one to cut. At every age both weeks have reached, the younger one is about 30% ahead. Revenue to date told you how old each cohort was. It did not tell you which one was better.
I read paid cohorts in three ways. At the same age. On proceeds, after the store, tax and refunds. In cells with enough purchases to mean something. The account figures below come from studies I have published, one account each. Every number in the worked example is arithmetic on a labeled assumption.
Why does revenue to date mislead?
Revenue accumulates with time. A cohort installed 56 days ago has had 56 days to pay, renew and watch ads. One installed 14 days ago has had 14. Compare their totals and you have added age to performance. The fix is to index every cohort by days since install and compare only at ages both have reached.
The tooling already works this way, which is why I distrust any report that does not. Apple’s App Store Connect Analytics reports Day 1, Day 7 and Day 35 download to paid for each download cohort. AppsFlyer’s cohort dashboard groups users by acquisition date, counts that date as day 0, and bases cohort days on calendar days, not on the install timestamp. A user who installs at 23:50 therefore has ten minutes of day 0. Two tools can disagree on D0 revenue for that reason alone.
The checkpoints are conventions. Apple picked 35 days, SKAN ends at 35, many teams use 28. What is not a convention is the rule under them. Same age, or no comparison. My own routine is D0 and D7 for the early read and D28 for the decision, in the MMP cohort report. D28 is the age I checked day 0 against in the $606K study, where day 0 ROAS ranked 38 Meta campaigns almost the same way day 28 did, with a rank correlation of +0.96, some of which is built in, since day 0 revenue is part of day 28 revenue. So I trust D0 to rank cohorts against each other, and I wait for D28 for the absolute number that goes against the target.
What does a comparison at the same age look like?
At the same age, the younger install week leads by about 30% at every checkpoint it has reached, even though its revenue to date looks worse. Every number in this section is illustrative. Two install weeks from the same campaign, $10,000 of spend and 2,000 installs each, so a $5.00 CPI. Week A is 56 days old today and Week B is 14 days old. The cells are cumulative proceeds per install, after store commission, tax and refunds.
| Install week | D0 | D7 | D14 | D28 | D56 | To date |
|---|---|---|---|---|---|---|
| Week A, 56 days old | $0.90 | $1.60 | $2.20 | $3.00 | $3.90 | $7,800, 78% ROAS |
| Week B, 14 days old | $1.20 | $2.10 | $2.85 | $5,700, 57% ROAS |
The wrong read is the last column. Week A has returned $7,800 against Week B’s $5,700, so Week A looks like the better targeting or the better creative, and Week B looks like a candidate for a cut. Week A has had four times as long to earn.
The right read is any column both weeks have filled. At D0, B is 33% ahead. At D7, 31% ahead. At D14, $2.85 against $2.20, 30% ahead, which is 57% ROAS against 44%. Week B is the better cohort at every age it has reached, and the naive read had the verdict backwards.
Illustrative cohorts from the table above, $5.00 CPI. The left chart adds age to performance. The right chart compares performance only.
Week B is 14 days old and my decision age is 28, so there is a second question. What do I do with its spend for two weeks? I project it from older cohorts. Week A grew from $2.20 at D14 to $3.00 at D28, a multiplier of 1.36. Applied to Week B’s $2.85 that gives about $3.89 per install at D28, or 78% ROAS. One older cohort is a guess, so I take the range of that D14 to D28 multiplier across older cohorts. Use cohorts with a similar country and creative mix. Say the range runs from 1.25 to 1.45. Week B then lands between $3.56 and $4.13, between 71% and 83% ROAS at D28.
Whether that band decides anything depends on the target. If the D28 target is 70%, both ends clear it and Week B can scale now. If it is 75%, the band straddles the line, Week B stays on watch at steady spend, and I read it again at D28. The multipliers hold only while product, pricing, paywall and channel mix stay close to the older cohorts. After a paywall test or a new market, start the range again.
Which revenue is the money?
Measure cohorts on proceeds for any profitability or payback decision, meaning what reaches you after the store commission, tax and refunds. For money, I use RevenueCat proceeds. For the split by campaign I use MMP revenue, and I treat it as a gross number for ranking, never as cash.
Apple defines Sales as the total amount billed to customers and Proceeds as the amount you receive, warns that proceeds in Sales and Trends are not final, and estimates US dollar amounts with the previous month’s exchange rates. The gap between the two is the commission and the tax where the price includes it, and refunds come off both. I set out how store commission, tax and refunds each move proceeds in the CPA ceiling article. The commission is 30%, 26%, 15%, or 10% plus a 5% billing fee, depending on store, program, region and subscriber tenure, and Google Play’s 2026 fee table made the spread wider.
RevenueCat keeps the layers apart. Its revenue chart subtracts refunds first and then estimates tax and commission to get to proceeds, and its reconciliation guide maps Apple’s customer price to its revenue figure and Apple’s partner share to its proceeds. Two caveats. The tax and commission are RevenueCat’s estimates, not the store’s statement. And Charts v3 subtracts a refund on the day it occurs, while a cohort read needs it assigned back to the purchase it reversed, so move refunds to their original cohort before you compare weeks.
The MMP number is a different thing. Adjust and AppsFlyer report whatever revenue the SDK or the server sent, gross or net depending on your setup, and I treat it as gross. They report it inside their own attribution rules and windows. Comparing a RevenueCat LTV with an Adjust ROAS compares two bases at once. I wrote up the three comparisons between Meta, RevenueCat and Adjust and the checks that make a gap trustworthy. For this article the rule is one basis, one currency, one calendar, on every level of the analysis.
How small can a cell be before the number means nothing?
A cell needs 50 purchases before I read it, because every cut you add multiplies the cells and divides the users in each one. A campaign with 2,000 installs across 10 countries, 2 stores and 5 creatives is 100 cells of 20 installs. At a 3% payer rate that is fewer than one payer per cell.
I took the number from one account I measured, $1.2M over 22 weeks, where an ad’s ROAS in one week predicted the next week’s ROAS best once the ad had 50 or more purchases that week, and even then the correlation was 0.44. Below 50 I do not read the cell. I roll it up, merging weeks, or merging countries into tiers, until it clears the line.
The statistics say the same thing in a different language. A payer rate is a proportion, and the standard formula for its margin of error misbehaves at small counts, which is why Brown, Cai and DasGupta recommend the Wilson interval instead. A 3% payer rate on 1,000 installs, 30 payers, has a 95% Wilson interval of about 2.1% to 4.3%. The same 3% on 100 installs, 3 payers, runs from about 1.0% to 8.5%. The second cell cannot tell a strong country from a weak one, and no amount of staring at it will change that.
Two more habits keep small cells from lying to you. Decide the decision age before you look, because checking daily and acting the first time a cohort crosses the line is the optional stopping problem that A/B testers already know. And treat a surprising winner in a small cell as a hypothesis. Gelman and Loken call the problem the garden of forking paths. With enough cuts, some cell will look great by chance, and a different dataset would have made a different cell look great. A subgroup result counts when it repeats in the next independent cohort week.
How do I do cohort analysis properly?
Start with the whole business by install week, then drill down only as far as the cell still has 50 or more purchases.
- The whole business, proceeds by install week, same age. Reconcile with the store and RevenueCat totals first.
- By store. Fee schedules, attribution regimes and refund rules differ, so read iOS and Android before you blend them.
- By channel, on the MMP basis, tracking the ratio of the platform number to the MMP number as a calibration factor.
- By campaign within channel.
- By country within campaign, only where the country clears the gate.
- By creative, the noisiest level. Read it as direction and pool across weeks.
How do I compare ROAS across countries fairly?
Compare countries at the same cohort age, on proceeds per install, inside one store, with the payer rate and its interval next to it, because countries differ in what a payer is worth before any difference in how well the ads work. Apple sets comparable prices across 175 storefronts with tax and exchange rates built in, and developers can override any of them by hand. Where the price includes VAT the commission is calculated after tax. The store fee can differ by region. The plan mix differs, so proceeds per payer differ even at the same price.
Then there is the mix. Two campaigns can rank one way inside every country and the other way blended, which is Simpson’s paradox, after Edward Simpson’s 1951 paper, applied to a media plan. Illustrative: Campaign X runs 70% ROAS in Tier 1 on $8,000 and 130% in Tier 3 on $2,000, blending to 82%. Campaign Y runs 60% in Tier 1 on $2,000 and 120% in Tier 3 on $8,000, blending to 108%. X beats Y in every tier and loses blended, purely on where the spend went. The right comparison depends on which mix you can actually buy at scale, and you only find out by looking inside.
In my audit of $383K of Meta tROAS spend, one country had run at 28% ROAS for three months inside a campaign that printed 70% blended, and nothing at campaign level moved. I scored every country on D0 and on All ROAS against the campaign’s own target, waited three months before excluding anything, and set a spend floor first, because a country with $80 behind it has a noisy ROAS, not a verdict. That audit shows the mix problem at country level. The same age rule and the gate are what this article adds.
What about the revenue the MMP calls organic?
Some of it is paid revenue the MMP could not trace. Under ATT a share of iOS paid revenue has no click to trace, so it lands in organic, and a cohort table shows only the revenue you could attribute. In one account’s organic and paid iOS installs over 103 days, on days with 100 more paid iOS installs about 28 more organic iOS installs showed up, a slope of 0.277 on iOS against 0.059 on Android, where the install referrer still works. Paid iOS cohorts in that account looked like a floor, and the gap between the two slopes is the part to size with a holdout.
SKAN makes the floor lower and later. Under SKAN 4 a campaign gets at most three postbacks, covering days 0 to 2, 3 to 7 and 8 to 35, the fine conversion value arrives only in the first, and at the lowest crowd anonymity tier no value arrives at all. Any revenue past day 35 that a dashboard assigns to an iOS campaign is modeled. Meta’s reporting runs on its own windows, and its developer documentation says that Facebook generally has a larger attribution window than most mobile measurement partners. So keep three layers apart instead of collapsing them into one ROAS. Deterministically attributed proceeds, the floor for paid. Modeled or probabilistic proceeds, labeled with the method. Unattributed proceeds, reported on their own and assigned to paid only through a stated assumption, such as a measured halo or a holdout result.
The direction of the organic interaction is not settled either, and I would not assume it in your favor. At eBay, Blake, Nosko and Tadelis found returns to paid search that were a fraction of what the attribution said, with brand keyword ads showing no measurable benefit in the short term. At one large US mobile game developer, Ju, Zhao and Aral found the opposite sign. Switching its ads off worldwide cut organic installs by 20% to 30%, and every $100 of spend went with about 32 paid installs and 2 organic ones, which they trace to paid installs lifting store rankings. It is a preprint from a single company, revised in July 2026, and the figures moved between versions. Different channel, different product, different answer. Your account has its own number, and the two reports in my iOS organic article are how I start looking for it.
When is a cohort old enough to judge?
Three tests, and a cohort has to pass all of them.
The windows have closed. The SKAN window of 35 days plus postback delay, the platform’s attribution window, the main refund exposure, and the store’s finalization of proceeds. Before that, part of the number is still arriving.
The revenue events have had time to happen. A yearly plan is readable soon after the trial converts, and the next unknown is the renewal a year out. A monthly plan will not have renewed by D28, so I take that revenue from the curve of older cohorts and let the band below decide. Games with purchases and ads keep earning, so maturity is a judgment, and the third test below is where I make it. How long you can wait is the ceiling’s job, and the ceiling is set before the cohort exists.
Later revenue would not change the decision. The band from the worked example. If the low and high multipliers from older cohorts both land the young cohort on the same side of the target, the decision is made, and waiting adds precision to a verdict that will not change. If they straddle it, the cohort stays on watch.
What can a cohort table not tell you?
It cannot tell you what the spend caused. Attribution is a ledger of who was matched to what, under one system’s rules. Gordon, Zettelmeyer, Bhargava and Chapsky compared 15 Facebook advertising experiments, 1.6 billion impressions, with the observational methods teams use every day, and found that those methods often failed to recover what the experiments measured, even with extensive conditioning. Two of the authors worked at Facebook, the data are Facebook’s, and the experiments were 2015 US Facebook campaigns, not app installs, which is a reason to read it carefully, not a reason to dismiss it. Only a holdout or a geo experiment answers the causal question, and even those are noisy. Lewis and Rao showed across 25 field experiments that individual purchases are so volatile relative to ad cost that informative intervals need enormous samples. I set out how I would run the test for one network in the AppLovin incrementality piece.
It cannot tell you the source of SKAN revenue past day 35 or below the anonymity threshold. It cannot tell you what a cohort will do after a price change, a new paywall, a new market or a new fee table, because the multipliers assume the future looks like the older cohorts. And reconciliation between systems explains the gaps between them without telling you which system’s attribution is right. Each is right for a different question: the store for cash, RevenueCat for subscription state, the MMP for credit across networks, the ad platform for its own optimization signal. And if the question is why a cohort got worse instead of which cohort is better, that is my diagnostic for rising CPI and stalled ROAS, and it starts from the change log, not from the cohort table.
What should you check before you call a cohort profitable?
- Same age. Every comparison at a day both cohorts have reached, D0 and D7 to rank, D28 to decide.
- Proceeds. After commission, tax and refunds, with refunds moved back to the cohort they came from, in one currency and one calendar.
- The gate. 50 purchases in the cell, or roll it up.
- Store before blend. iOS and Android read separately before any combined number.
- The mix check. Does the winner still win inside each country tier?
- Replication. A surprising cell has to repeat in the next cohort week.
- The band. For young cohorts, a low and a high multiplier from older cohorts, and a decision only when both land on the same side of the target.
- The organic layer. Deterministic, modeled and unattributed proceeds kept apart.
If you cannot fill in those eight lines for your account, the growth audit is where I read the cohorts and apply the gates for an account, and you can book it on its own at any spend level.
Sources and scope
I had each page checked on 5 October 2026. Platform documentation changes without notice, so check the date before you quote a window or a fee. The worked example, the Simpson’s paradox illustration and the Wilson intervals are arithmetic on labeled assumptions, not an account. The account figures I link to are one account each, observed and not experimental.
- Apple, Analytics dashboard, App Store Connect Analytics Help
- Apple, View units, proceeds, sales, and pre-orders, App Store Connect Help
- Apple, Manage pricing for auto renewable subscriptions, App Store Connect Help
- AppsFlyer, Cohort and retention dashboard, AppsFlyer Help Center
- Adjust, How SKAdNetwork 4 works, Adjust Help Center
- Apple, Changes for apps in the European Union, updated 18 August 2026, for the 26% rate inside the App Store
- Google, Service fees, Play Console Help
- Meta, App Events API, Meta for Developers
- RevenueCat, Revenue chart, Charts v3 and Reconciling with App Store Financial Reports, RevenueCat Docs
- Brown, Cai and DasGupta, Interval Estimation for a Binomial Proportion, Statistical Science 16(2), 2001
- Gelman and Loken, The Statistical Crisis in Science, American Scientist 102(6), 2014
- Simpson, The Interpretation of Interaction in Contingency Tables, Journal of the Royal Statistical Society, Series B 13(2), 1951
- Stanford Encyclopedia of Philosophy, Simpson’s Paradox, as the explainer
- Gordon, Zettelmeyer, Bhargava and Chapsky, A Comparison of Approaches to Advertising Measurement, Marketing Science 38(2), 2019, 15 US Facebook experiments, two authors at Facebook
- Blake, Nosko and Tadelis, Consumer Heterogeneity and Paid Search Effectiveness, Econometrica 83(1), 2015, eBay, authors affiliated with eBay at the time
- Ju, Zhao and Aral, Advertising Spillovers in Mobile Apps: Evidence from Ad Shutoffs and Store Rankings, arXiv preprint, submitted April 2025 and revised July 2026, one US game developer
- Lewis and Rao, The Unfavorable Economics of Measuring the Returns to Advertising, Quarterly Journal of Economics 130(4), 2015
Questions people ask
How do I tell which of my paid app install cohorts are actually paying back?
Line every cohort up by days since install and compare them only at ages both have reached, on proceeds after the store commission, tax and refunds. Revenue to date rewards the older cohort for being older. I read D0 and D7 early, decide at D28, and only in cells with 50 or more purchases. Below that I merge weeks or merge countries into tiers.
How do I do cohort analysis properly for paid user acquisition?
Start with the whole business by install week, then split by store, then by channel, then by campaign, and go down to country and creative only where the cell still has 50 or more purchases. Compare cohorts only at the same age, keep one revenue basis, one currency and one calendar on every level, and treat a surprising winner in a small cell as a hypothesis to retest on the next cohort week.
Should I measure cohort ROAS on gross revenue or on proceeds?
Measure cohort ROAS on proceeds for any profitability or payback decision. Apple and Google take their commission after tax where the price includes it, and refunds take back what reached you. MMP and ad platform revenue is gross or net depending on what the SDK sent, so I treat it as gross, use RevenueCat proceeds as the money and MMP revenue only to split that money by campaign.
How do I measure ROAS across countries without fooling myself?
Compare countries at the same cohort age, on proceeds, inside one store before blending across stores, and only where the country clears the purchase gate. Prices, tax, store fee and plan mix all differ, so a blended number hides the spread. In my audit of $383K one country ran at 28% ROAS for three months inside a campaign that printed 70%.
When is a cohort old enough to judge?
A cohort is old enough to judge when its reporting windows have closed, the revenue events that matter have had time to happen, and the range of plausible outcomes no longer straddles your target. For a yearly plan that is soon after the trial converts. For a monthly plan the first renewal falls after D28, so that revenue comes from the curve of older cohorts and the range it gives has to clear the target. My decision age is D28, with D0 used to rank, since D0 ranked 38 campaigns almost the same way D28 did in my $606K study.