Method
Every teardown on this site is scored against one fixed rubric instead of a fresh judgment call each time. This page is where that rubric lives.
v1 · Dated 2026-08-03
This is a first draft of the rubric, written to be edited, not a finished framework. It carries no claim about how many times it has been run. It will be revised as it is actually used, and this page will note what changed and when that happens.
Whether the numbers can be trusted at all. Every other stage in this rubric is scored against data that has to be right first: a 5/5 store listing or funnel measured against broken event data is a 5/5 on noise. See why revenue numbers don't match for the five specific causes point 4 below checks against.
| Point 1 | A written event taxonomy exists for every purchase-adjacent event (trial start, trial conversion, renewal, refund, billing-retry recovery), and every tool that logs one of these uses the same name and the same trigger condition for it. |
|---|---|
| Point 2 | Revenue reaching the analytics tool is validated against Apple, through a subscription platform, a direct App Store Server Notifications feed, or receipt verification wired into the client SDK, not a bare client-fired event with nothing behind it. |
| Point 3 | One canonical subscriber identity resolves the same person across every tool in use (analytics, subscription platform, ad-attribution SDK), and anonymous-to-identified merges are logged, not silent. |
| Point 4 | RevenueCat (or equivalent), the analytics tool, and App Store Connect have actually been reconciled for the most recently closed month, and every point of disagreement has a named cause. |
| Point 5 | Someone on the team can state, without checking, which single number is the source of truth for a finance question and which is the source of truth for a same-week product question. |
Failing (score 0 or 1 of 5): nobody has compared the three dashboards against each other, revenue tracking is entirely client-fired with no server-side validation anywhere, and there is no written definition distinguishing a trial conversion from an ordinary renewal.
Passing (score 4 or 5 of 5): the reconciliation is a routine step with an owner, not a one-off investigation, and every number anyone reports traces back to a documented definition someone else on the team could independently check.
Whether the listing is doing its job before a dollar of paid traffic touches it. Checked directly against the live listing and against App Store Connect's own field rules, not against taste.
| Point 1 | The app name uses its full budget, up to Apple’s 30-character limit (Apple Developer, App Store Connect Help, "App information" reference, checked 2026-08-03), for a real descriptive or keyword-bearing phrase, not padded or left short. |
|---|---|
| Point 2 | The subtitle (also capped at 30 characters) carries keyword or positioning value distinct from the app name; the two fields are not repeating the same words. |
| Point 3 | The keyword field is filled close to its 100-character limit, contains no word already present in the app name, subtitle, or category, and has been reviewed since the last name or subtitle change. |
| Point 4 | Screenshots exist for every supported device size class Apple requires an image for, and none is more than one major OS or UI redesign out of date. |
| Point 5 | Every locale the app is actually live in on the App Store has a complete, localized screenshot and metadata set; no live locale is silently falling back to translated text under untranslated screenshots. |
Failing (score 0 or 1 of 5): the app name and subtitle repeat each other, the keyword field is mostly empty or full of duplicate words, and at least one live locale has no localized screenshots at all.
Passing (score 4 or 5 of 5): every text field is used to its stated limit with no duplication across fields, screenshots are current, and localization coverage matches the app’s actual live storefront list exactly, not just its major markets.
One funnel, several rates, and each rate needs the right denominator or it does not mean anything. Apple’s own App Analytics defines Conversion Rate as total downloads and pre-orders divided by unique device impressions (Apple Developer, App Store Connect Help, "App metrics" reference, checked 2026-08-03). Everything below applies that same discipline, one denominator per stage, past the download and into the subscription.
| Point 1 | Impression-to-product-page-view rate is tracked on its own, not folded into a single blended conversion number, isolating how well the icon, name and category pull attention into the listing. |
|---|---|
| Point 2 | Product-page-view-to-install rate is tracked separately from impression-to-install, since a listing can be well-optimized while the page itself fails to convert, or the reverse. |
| Point 3 | Install-to-trial-start rate is measured against first opens, not total downloads, so redownloads and reinstalls by existing users do not dilute the denominator. |
| Point 4 | Trial-to-paid conversion rate is measured against trial starts in the same cohort, and read only after the full trial length plus a settlement buffer has elapsed for that cohort, not against the current month’s in-progress trial starts. |
| Point 5 | Every rate above is attached to a specific app version, locale and traffic source, not reported as one blended number: a blended rate cannot distinguish a store-surface problem (stage two) from a paywall problem, and the two need different fixes. |
Failing (score 0 or 1 of 5): the team has one number, "conversion rate," with no visibility into which stage it describes, and trial-to-paid is being read before the cohort’s trial period has finished.
Passing (score 4 or 5 of 5): every stage of the funnel has its own tracked rate, with the correct denominator, broken out by version and locale, and trial-to-paid is read only on matured cohorts.
Whether this team can actually learn something from a test, decided before the test runs, not after it is inconclusive. This is the stage the free subscription A/B test sample size calculator exists to support directly; every formula referenced below is cited in full on that page.
| Point 1 | Sample size was calculated before the test started, against the actual baseline rate and a stated minimum detectable effect, not a rule of thumb like "500 conversions per variant" that ignores the app’s own baseline. |
|---|---|
| Point 2 | If the test targets one stage of a multi-stage funnel, the sample size is computed on the resulting overall rate, not the targeted stage’s own rate in isolation: an absolute effect on one stage shrinks once it passes through the downstream stages. |
| Point 3 | The test has a fixed, pre-declared read date or sample size, and nobody is checking results daily and stopping at the first apparent win: continuous monitoring with early stopping inflates a nominal 5% false-positive rate to roughly 26% at a check every 500 or so visitors. |
| Point 4 | The read date accounts for trial length and a settlement buffer, not just enrollment volume: the last cohort enrolled is not measurable as trial-to-paid until its full trial period plus a few days for billing retries has passed. |
| Point 5 | If the primary metric is revenue per install rather than a conversion rate, the team is treating it as the heavy-tailed, high-variance metric it is, not applying a binary-proportion formula meant for conversion rates to it. |
Failing (score 0 or 1 of 5): sample size is a round number picked without a formula, the team checks the dashboard daily and calls the test the moment it looks good, and a paywall-only test’s power was calculated on the paywall’s own conversion rate instead of the diluted overall rate.
Passing (score 4 or 5 of 5): sample size and read date are both fixed before the test starts, calculated against the actual funnel shape, and nobody looks at results until the pre-declared date arrives.
Four stages, five points each: twenty points total. Read the score stage by stage, not as one blended number, for the same reason stage three scores a funnel one denominator at a time. A low score on stage one, measurement integrity, makes every other stage's score unreliable, since it was measured against numbers that have not themselves been checked.
See it applied, once it exists, at /teardowns.