The Honest Truth About ASO Data

No ASO tool knows your keyword's search volume on the App Store. Apple does not publish it, to anyone, at any price. So every volume number you have ever seen in an ASO dashboard is a model output rather than a measurement, and the good news is that this is fine as long as you know which one you are looking at.

The problem is that almost nothing in the category tells you.

What the stores actually publish

Apple gives developers no per-keyword search volume. App Store Connect's App Analytics reports impressions, product page views and downloads broken down by acquisition source, so you can see how much came from App Store Search as a channel. You cannot see which searches.

Google exposes no public Play Store search API. Google Ads Keyword Planner does publish volume, but it measures Google web search, which is a different product with different users.

That distinction gets waved away constantly and it should not be. Web search volume for "budget app" reflects people at a desk researching options, often on a laptop. Play search volume for "budget app" reflects people already inside a store with intent to install. Same string, different population, different intent, no reliable conversion factor between them.

Apple Search Ads is the one Apple surface that gives keyword-level signal, and it exists to help advertisers bid. It is not an organic volume report, and treating it as one imports the biases of an ad auction into an SEO decision.

That is the entire supply of first-party keyword data. Everything else is inference.

How the inference actually works

This is worth understanding, because a reader who knows the pipeline can use the outputs properly instead of either trusting them blindly or dismissing them entirely.

The observable inputs are real. Autocomplete suggestions and the order they appear in. Search result rankings, sampled repeatedly over time. Category chart positions. App-level download estimates. Panel and SDK data from devices whose owners agreed to be measured. Signals leaking out of ad auctions. Extrapolation from one store to the other.

Every one of those is a proxy. None of them is search volume.

A typical volume estimate is then built by stacking them. A model turns ranking movements into implied traffic. That rests on a download estimate for the apps involved. That download estimate rests on a panel sample, which rests on an assumption that the panel resembles the population. Each layer has an error term.

Stacked models multiply error, they do not average it away. A number produced through three inferential steps carries the uncertainty of all three, and the dashboard shows you none of it.

Why the numbers look so confident

Because confidence sells and ranges do not.

A keyword difficulty score of 42.7 looks like an instrument reading. "Somewhere between low and moderate, based on rankings we sampled last week" looks like hedging, even though it is the more accurate statement. Product design rewards the first one: it fits in a table cell, it sorts, it makes a nice chart, and nobody asks a follow-up question.

A number without an error bar is not more accurate than a number with one. It is less honest by exactly the width of the missing bar.

What inference is good for, and what it is not

Inferred numbers do real work inside their limits.

They are useful for relative comparison within a single tool at a single point in time. If keyword A scores higher than keyword B in the same dashboard on the same day, the model is at least applying a consistent method to both, and that ordering is often good enough to make a decision.

They are useful for direction of travel. A term trending up over eight weeks in a consistent model is meaningful even when the absolute figure is not.

They are useful for shortlisting. Going from four hundred candidate keywords to twelve is exactly the job a rough model should do.

They are not useful for absolute forecasting. "This keyword gets 12,000 monthly searches" cannot be verified by you, by the vendor, or by anyone except Apple.

They are not useful in a business case. If a revenue projection has an inferred search volume in the denominator chain, the projection inherits every error above and reports none of them.

And they cannot support a claim that a specific keyword will produce a specific number of installs. Nothing available to a developer can support that claim.

Why labelling inference beats a confident number

Not because honesty is nice. Because a labelled number tells you something an unlabelled one cannot: what decision it is allowed to carry.

If a figure says "modelled, from ranking samples", you know to use it for shortlisting and not for a forecast. That is actionable metadata about the number itself, and it is free to provide.

It tells you when to stop and go get ground truth. Knowing that a volume figure is a guess is what prompts you to run a listing experiment instead of arguing about the guess for a week.

It saves you real time. The most expensive ASO mistake is not picking a mediocre keyword, it is spending a fortnight optimising towards a number that a model invented, then having no way to work out why the result did not land.

And it preserves the signal in the measurements you do have. When a dashboard renders modelled volume, observed rank and your own download count in identical confident type, it destroys your ability to tell which is which. Uniform confidence is indistinguishable from no information about confidence at all.

What to do instead

You can run a real ASO programme without a single volume estimate.

Track ranks for terms you already rank on. Rank position is observable. Somebody can look at the search results page and see where you sit. This is measurement, not modelling.

Use your own acquisition data. App Analytics and Play Console are measuring you, first-party, no panel involved. They will not tell you which keyword, but they will tell you truthfully whether search traffic to your listing went up after you changed something.

Run deliberate experiments. Change one thing, wait a defined period, compare. Apple's product page optimisation tests and Play's store listing experiments are genuine measurement of your own listing against a control.

Read the search results page yourself. For any term you care about, look at what ranks. You will learn who you would have to displace, how strong their listings are, and whether the term is winnable. It costs nothing and it is direct observation.

Interrogate your vendor. Four questions, and they are fair ones: is this figure measured or modelled, what are the inputs, what is the error, and what changed the last time you updated the model. A vendor who answers those clearly is more useful than one with a bigger number, and the answers tell you which columns in their product to trust.

The position worth holding

Apple exposes no per-keyword search volume. Google exposes no Play search API. Nobody has solved this quietly. An industry has grown up around estimating it, some of that estimation is skilful, and none of it is measurement.

The useful response is not to abandon the tools. It is to insist that every number arrives with its provenance attached, and to treat any product that will not say where a figure came from as having answered the question.

Where AppSubmit shows an inferred figure it is labelled as inference rather than presented as measurement, and where a check cannot run against a store it is skipped rather than filled in with a plausible number. That is the same standard described above, and you can hold any tool to it.

Related

  • App Store character limits (app-store-character-limits)
  • Google Play character limits (google-play-character-limits)
  • App Store versus Google Play locale codes (app-store-vs-google-play-locale-codes)

Last reviewed 2026-08-17. Apple and Google change their rules without notice, so check anything decision-critical against their live documentation.