Pull up the difficulty score for any keyword on any major ASO tool. The number is real. It is also the same shape for every app that searches it.
That is the global-model assumption — that a single calibrated weight set, derived across the whole catalog of apps in the database, can rank keyword opportunities or experiment designs for your app. It is the assumption that produces the "Most ASO tools use 'difficulty' as a black box score" complaint that surfaces in every practitioner conversation we have been part of for the last six months.
The complaint is not that the score is wrong. It is that it is generic. A global model, by construction, cannot tell a casual gaming title from a healthtech app from a B2B fintech tool. The shape of what each one is testing is genuinely different. The shape of what works is genuinely different. The shape of what counts as evidence is genuinely different. A score that does not see those differences cannot be defended in the room where the experiment gets approved.
This piece is the long version of why ASOLOOP is built without one.
What "global model" actually means in ASO tools
Most ASO intelligence tools train their models on aggregate behavior across a large catalog. The number you see — keyword difficulty, predicted CVR, opportunity score — is some function of what works on average across all apps the tool observes. The output is then displayed against your app, but the calibration was never against your app. It was against the average.
Three things follow from that:
- Your verticals are not in the average. A casual gaming title runs in a market where icon style, character expression, color palette, and FOMO copy carry most of the install signal. A healthtech app runs in a market where trust, regulatory framing, and screenshot order carry most of the signal. A fintech app runs in a market where security cues and feature-list specificity carry most of the signal. Each of these is a different space of what experiments measure. A global average compresses them into one.
- Your audience is not in the average. Even within a vertical, your app's specific audience — paid acquisition cohort, geo mix, language, prior install graph — is a particular slice. The global model does not see that slice. It sees the population.
- Your prior tests are not in the average. Your last 12 experiments produced signals about which dimensions move CVR for this app. The global model does not have access to your tests; it cannot weight what you have already learned.
A practitioner's instinct that this is wrong is correct. A score that ignores all three of those layers is structurally generic, regardless of how sophisticated the model behind it.
What "per app, per audience" looks like as an architectural commitment
ASOLOOP's signal model is per app, per audience by design. There is no global ASO model. Every signal contribution is scoped to the app it was learned from; the ranking layer that decides what to test next reads only that app's accumulated evidence.
That is a commitment, not a feature flag. It rules out a class of marketing copy ASOLOOP will never write — "in our database, apps like yours saw a 14% lift on icon experiments". We do not know what apps like yours did. We know what your app did. The catalog of what other apps did is interesting only as a proof point that compounding works in general; it is not the basis for a recommendation about your next test.
The trade is real. We give up the easy comparative claim — "benchmarked against 50,000 apps" — and we give back something a practitioner can actually defend in a leadership review: "this recommendation is built from your app's last 12 experiments, weighted by precision tier and recency, with every signal traceable to the experiment that produced it." Two different shapes of credibility. The first wins on cover-letter language. The second wins when the VP asks the second question.
The five-value experiment-type taxonomy
A consequence of per-app design is that ASOLOOP has to ask the operator something at create time that global-model tools never do: what kind of experiment is this, exactly?
The five values:
native_ppo— Apple PPO (custom product page, native experimentation surface).native_sle— Google Store Listing Experiment (Play Console native experimentation). The Android default experiment type: up to 5 concurrent localized experiments × 3 treatments; Custom Store Listings (CSL) are Google's second organic lane.cpp_paid— Custom Product Page surfaced via paid traffic (Apple Search Ads creative-set assignment).csl_paid— Custom Store Listing surfaced via paid traffic (Google Custom Store Listing routing).owned_routed— owned-channel traffic routed via OneLink, Adjust, Branch, or equivalent, where the variant assignment is carried by the click ID.
The taxonomy is not a UX preference. Each value implies a different signal precision class for what the experiment can measure post-install. native_ppo and native_sle produce cohort-window approximations because the store does not expose variant assignment for organic traffic. cpp_paid, csl_paid, and owned_routed produce per-variant deterministic measurements because the variant ID rides with the install.
When you tell ASOLOOP at create time that this is a native_ppo test, the system already knows the experiment will fire at the cohort-window tier. It surfaces that to you before you launch — which means you make the decision "is cohort-window approximation precise enough for what I need to learn?" with the answer in front of you, not retroactively.
The global-model alternative is a single experiment-type input that all five conditions get squeezed into, and a single confidence number that mixes the precision tiers silently. ASOLOOP refuses that compression. The taxonomy is the price of refusing it.
The five outcome categories
The second create-time question ASOLOOP asks: which 1–2 outcomes does this experiment care about?
The five values:
activation— first meaningful action post-install (signed up, completed onboarding, fired the first event the product cares about).subscription_start— first paid subscription event.purchase— first transactional purchase event.retention_d7— return on day 7 post-install.retention_d30— return on day 30 post-install.
You pick one or two. Not all five. The opt-in is deliberate.
A subscription healthtech app and a transactional commerce app and a retention-driven gaming app are testing fundamentally different post-install signals. ASOLOOP does not silently track everything and let the system pick what looked good. The operator names what the experiment is for, and the signal contribution is computed against that named outcome. If you said the experiment was about subscription_start and the data shows installs went up but subscription starts did not, the system reports that honestly — it does not retreat to the install metric to manufacture a positive verdict.
This is what honest precision looks like at the data-model level. The outcome the experiment was designed for is the outcome the experiment is judged on.
Why this matters for your vertical, specifically
The way the per-app commitment shows up in practice depends on the kind of app you run.
Gaming. Your icon and screenshot work moves CVR more than your subtitle work. Your audiences split sharply by genre — a casual word game and a hardcore strategy game share almost nothing on the things that matter. ASOLOOP's per-app signal accumulation means the system that recommends your next icon test has the last 12 icon iterations of your game in its weight set, not the average of someone else's game. The recommendation is genre-aware because it is app-aware.
Healthtech. Your CVR ceiling is gated by trust signals — regulatory framing, screenshot order, "as seen in" cues, professional review markers. Your monetization model (often subscription) makes subscription_start the right outcome category, not activation. ASOLOOP's per-app model lets you weight subscription_start consistently across iterations, and the claim-safety validator on generated copy refuses regulated-claim language without an evidence object. Both behaviors are direct outputs of treating healthtech as its own context, not a row in an average.
Fintech. Your cohort skews older, your trust signals are different again (security cues, FDIC-equivalent disclosures, transparency about fees). Your retention curves are steep and the right outcome is often retention_d30, not activation. The global model does not have a "for fintech apps with this audience composition, weight retention more than activation" parameter. A per-app system has it implicitly — your prior signals carry that weighting forward.
B2B / vertical SaaS apps. Your traffic is small, your acquisition cycles are long, your statistical significance bar is structurally hard to clear. ASOLOOP's tier-disclosure surface is the unlock here — knowing that a small-traffic test is going to fire at cohort-window or insufficient-sample tier before running it changes the decision about whether to run it as a routed CPP instead.
The pattern is consistent. The work each kind of app does is different in a way the global model cannot represent. A per-app, per-audience system represents it by carrying your app's signals forward and weighing them against your app's outcomes.
What you give up
Per-app learning has a cold start. The first three or four experiments on a new app run in Learning mode — the recommendation surface shows three to four ranked hypotheses with explicit thin-evidence disclaimers and signal-link traceability per card. The system is honest that the signal base is being built. It does not pretend to know more than it does.
That is the structural cost of refusing the global-model shortcut. There is no synthetic baseline for your new app. The system catches up after a handful of experiments. By the time you cross the conclusive-experiment threshold, the recommendations move into Classic mode — same shape (three to four ranked candidates with full traceability), confident framing, no disclaimers.
The cost is real and the trade is intentional. A loud confident recommendation on day one of a new app is a global model talking. A quiet honest recommendation that gets sharper across the next four tests is a per-app model building.
The honest read
A practitioner once described the right posture better than we could:
Most ASO tools use 'difficulty' as a black box score.
The honest version of that complaint is not that scores are bad. It is that scores divorced from the app they are supposed to apply to are structurally untrustworthy. A practitioner who has been burned by a black-box recommendation will not act on the next one without proof of work. Per-app, per-audience design is the proof of work — every signal traces to your app's experiments, every ranking weight is your app's evidence, every recommendation can be questioned at the level where it was learned.
Generic is convenient marketing. Specific is what survives a leadership meeting.