All posts
ASOMay 8, 20267 min read

What 'inconclusive' actually means: a primer on signal precision

Apple says 'inconclusive.' Your tool nods along. But three completely different conditions can produce that same word — and only one of them is a dead end. Here is how to tell which is which.

A

ASOLOOP team

Field notes

There is a sentence I have heard from four different ASO managers in the last two months, almost word for word:

I've run a couple of experiments and they all were inconclusive… Only big numbers can say something.

The implied next move is always the same — wait, accumulate more traffic, try again later. Sometimes that is the correct read. Often it is not. The problem is that the word inconclusive covers at least three different conditions, and each one wants a different next move.

This is a primer on what inconclusive actually means once you stop treating it as a single thing.


The platform's "inconclusive" is a binary

Apple's PPO surfaces a verdict against an internal confidence threshold: either the platform thinks one variant is convincingly better, or it does not. Below the threshold, you see inconclusive. The same is true of most other platform-reported significance signals.

A binary verdict is a useful summary. It is also the entire reason "inconclusive" gets treated as one condition. Your dashboard shows the same word every time. Your spreadsheet records the same status. Your retro lists "four inconclusive tests this quarter" with no further structure.

The platform is not wrong. It is just compressed. Underneath that one word are three distinct precision regimes, and only one of them tells you the experiment was a dead end.


Three precision tiers

ASOLOOP labels every post-install signal with one of three precision tiers. The tiers describe what the experiment was actually capable of measuring, before any verdict was reached.

Per-variant deterministic. Each install carries the variant assignment because the traffic was routed — Apple Search Ads creative IDs, UTM-tagged Google Ads, OneLink / Adjust / Branch trackers for owned media. The post-install metric you observe is genuinely per arm. If this experiment came back inconclusive, the effect size was small relative to the traffic you put through it.

Per-variant modeled. Same routing types, but on ATT-denied iOS. SKAdNetwork delivers aggregate, delayed postbacks. The variant assignment can be reconstructed but the per-install precision is modeled, not measured. ASOLOOP surfaces this with an explicit modeled, not measured tag — practitioners working on iOS know this regime, but most tools mix it silently with the deterministic tier.

Cohort-window approximation. This is the regime that catches most native PPO, Google Store Listing Experiment, and Custom Store Listing tests. The store does not expose variant assignment for organic traffic. The cohort during the experiment window includes installs from both arms in the platform-assigned split, so the post-install delta is a population-level estimate, not per-arm. The signal is real but the precision class is fundamentally lower than the routed-traffic tiers.

Every metric contribution ASOLOOP writes carries one of these tier labels. The reasoning layer surfaces the tier in the why-text on every hypothesis card, so when ASOLOOP weights a signal in ranking, the operator can see the precision class behind the weight. Every signal can be inspected, questioned, or revoked. That is what bounded means in practice — not a slogan, an explicit data field on every contribution.


Why the tier changes the next move

The same word — inconclusive — wants a different action depending on which tier produced it.

If your last test was a cohort-window approximation, the answer is rarely "more traffic." Adding traffic to an organic native PPO does not turn a population-level estimate into a per-variant one. The structural ceiling is the cohort. The honest next move might be re-running the same hypothesis as a CPP routed via Apple Search Ads, where the same install volume buys you a fundamentally higher precision class.

If your last test was per-variant modeled (ATT-denied iOS), the modeling uncertainty is stacking with the effect-size uncertainty. The next move might be a parallel Android leg if cross-platform parity holds for that hypothesis class — Android attribution is more permissive and the cohort window matches more cleanly. Or it might be accepting that for ATT-denied surfaces, the modeled tier is the ceiling and adjusting your hypothesis bar accordingly.

If your last test was per-variant deterministic and still inconclusive, the platform is telling you something accurate. The effect was real-or-zero at the precision level you measured at, and the result did not separate from noise. Here, "we need more traffic" is sometimes the correct diagnosis. Sometimes the correct diagnosis is that the variant did not actually move the metric and the next iteration should test something with a larger expected lift.

Three different inconclusive results. Three different next moves. None of them visible if your tooling treats inconclusive as one thing.


What changes when you see the tier before you launch

Most operators learn the tier of an experiment retrospectively, after they have already spent four weeks of traffic on it. ASOLOOP puts the measurement basis in front of you at create time — every subtype row in the wizard states what it is Measured by, in plain words, before you activate anything. What it does not do yet is name a precision tier on that screen: the tier itself is resolved at read time from consent coverage and cohort size, and it surfaces with the result rather than in the picker.

That single piece of information rearranges the experimentation calendar.

A team running a healthtech app in iOS-heavy markets, with most users in ATT-denied populations, would see at create time that a native PPO icon test will fire at the cohort-window tier. With that information up front, the team can decide: is cohort-window approximation the precision class my decision actually needs? If yes, run it. If not, restructure as a CPP routed via paid traffic before committing the four weeks.

This is what "verify early and often" means at the planning layer. (That phrase is practitioner shorthand we keep hearing — it is the right posture for AI in this category, and it is the right posture for experiment design.)


The honest tradeoff

Bounded AI is more demanding at the front end than the typical "click run" experience competitors offer. The user picks the experiment type from a five-value taxonomy at create time. The user selects one or two outcome categories from five — activation, subscription_start, purchase, retention_d7, retention_d30 — instead of letting the system silently track everything. The user sees the precision tier before launch.

That is more upfront input than the dashboard era trained operators to expect. The trade is intentional. You give up the fiction that the system can give a universal verdict, and you get back a system that tells you precisely what the next test can prove. Every signal it weights carries a traceable tier label. Every recommendation it surfaces can be questioned at the precision-class level. Nothing is presented as conclusive that was actually approximate.

The line we keep coming back to is the practitioner's, not ours: "It's not your all-in-one solution but it helps… best to have it reviewed by a professional." That is the correct posture. Bounded AI does not replace your judgment. It surfaces enough of its own work that your judgment has something accurate to operate on.


Revisiting the four-inconclusive quarter

In the previous post we walked through a healthtech operator with four inconclusive tests over 14 weeks who was ready to call the quarter a wash. That story reads differently through the precision-tier lens.

Two of those four tests were native PPO icon variants — cohort-window tier. They came back inconclusive at the population-level aggregation, but a routed CPP test on the same hypothesis would have fired at the per-variant tier. With that visible at create time, two of the four would have been restructured before launch. Of the remaining two, one ran at per-variant deterministic and hit a competitor-launch confound mid-window. One had a per-variant deterministic conclusive signal in a sub-cohort that the aggregate buried.

"Four inconclusive tests, the quarter was a wash" is the spreadsheet read. The tier-aware read is closer to: two tests that should have been routed for the precision the decision needed; one test with an external confound; one test with a real signal in a sub-cohort that aggregate reporting hid. The quarter is not a wash. The next quarter has four pieces of methodology refinement to use.

The verdict is not always the most useful information from an experiment. The precision class behind the verdict often is.

Other tools advise · ASOLOOP operates

Stop running ASO by hand.

Point ASOLOOP at your app, set your autonomy level, and read the receipts. Seven days, your apps, real experiments running on both stores.

Credit card required to start · Not billed during the trial · Cancel anytime

Your store credentials and your own AI-model keys — inspectable, revocable any time.