Back to articles
Two phones side by side showing different apps

Everything is 4.5 stars

Open either app store and scroll. Almost everything sits between 4.3 and 4.8.

A rating scale where nearly every entry lands in a half-star band is not measuring quality. It has stopped being a measurement at all and become a formality, like a restaurant score where anything under a certain number means "closed".

The interesting question is not whether ratings are inflated. They obviously are. It is why the inflation is structural, which is what makes it unfixable from the outside, and what to use instead.


Where the number comes from

The mechanism matters more than the result.

Almost nobody reviews an app unprompted

Left alone, review volume is close to zero and the people who do bother are the extremes: the delighted and the furious. That produces a bimodal distribution, lots of fives and ones, very little in between, because the merely satisfied majority never opens the dialog. The average of that distribution is not a measure of typical experience. It is an artefact of who self-selects.

Every store rating you see is a correction to that problem, applied by the developer.

So developers choose the moment

The store APIs let an app request a review, and the developer decides when. Nobody asks after a crash. They ask after something worked: a task completed, a goal hit, a purchase confirmed.

This is not a trick anyone is hiding. It is standard practice, discussed openly in the industry, and research suggests when and how you prompt can influence your store rating more than the quality of the app does. Read that again, because it is the whole article: the timing of the question outweighs the answer.

And many filter before asking

The more consequential version is the pre-prompt. Before triggering the store dialog, the app asks its own question: How are you enjoying the app?

Tap the happy face and you get the review dialog. Tap the unhappy one and you get a support form. Nothing was faked. Negative sentiment simply never reached the public rating. Both stores permit this, and reported lifts run to several tenths of a star.

A tenth of a star sounds trivial. Across a category where everything sits between 4.3 and 4.8, several tenths is the difference between the top of the list and the middle.

Which produces an arms race

Once enough apps prompt strategically, 4.5 becomes the floor rather than a good score. Anything below starts to look defective, so everyone else optimises their prompts to keep up, which raises the baseline again.

This is not hypothetical. A long-running study of an online labour marketplace found perfect five star ratings rose from roughly a third of all evaluations to about 85 percent over six years, at which point the score could no longer separate good work from mediocre work. Different marketplace, identical dynamic.


What a review actually costs to leave

Worth understanding the denominator, because it explains why the numbers are so easy to move.

Reviewing an app is unusually expensive in attention terms. You have to stop what you came to do, decide a number, often write something, and submit it. The reward is nothing. Nobody thanks you and the developer may never reply.

So the base rate is tiny. A minority of users ever leave a review unprompted, which means a store rating is built from a very small, very unrepresentative sample of the people who use the thing.

Two consequences follow, and both are counter-intuitive.

Small apps have noisier ratings than big ones. An app with 300 reviews can be moved half a star by a couple of dozen people. The score is not wrong exactly, it is just imprecise in a way the single displayed number hides completely.

Ratings lag reality badly. An app with 50,000 reviews carries years of accumulated history. If it got substantially worse six months ago, the average barely moves, because the recent reviews are drowned by the old ones. The number describes the app as it was, weighted by how popular it used to be.

This is why sorting by recent matters more than any other single step in the method below. The average is a historical artefact. The last three weeks of reviews are the current product.


What the number still tells you

Not nothing. Three things survive.

Below about 4.0 is meaningful. Given how hard the whole system pushes upward, an app that has landed under four has usually done something substantial: a botched migration, a hostile pricing change, sustained crashes. Worth investigating rather than avoiding outright, because the cause is often a single decision rather than general quality.

A very low review count is meaningful. Forty reviews is a small sample and easy to influence. Forty thousand is harder to distort.

A sudden drop is the most useful signal on the page. Ratings move slowly, so a visible fall means something recent and significant. The reviews from that period will tell you exactly what.

What the number cannot tell you is whether a 4.7 is better than a 4.5. That difference is prompt engineering, not product quality.


The five minute method

Ignore the average. This takes about as long as reading it would, and returns something real.

1. Sort recent negative reviews by date

Not "most helpful", which surfaces old reviews that may describe fixed problems. Most recent, filtered to one and two stars. Read fifteen.

You are looking for repetition, not outrage. One person furious about something irrelevant to you is noise. Five people in three weeks describing the same broken sync is a fact about the product.

Phrases worth weighting heavily:

  • "was great until the update that added…" — the clearest early signal of feature creep, covered in why apps get worse over time
  • "cancelled and was still charged" — a statement about how the company operates, not the software
  • the same bug reported across months — nobody is fixing it

2. Check the update history

On the store listing. Steady releases over twelve months or more is the strongest single indicator of a maintained product.

A six month gap means finished or abandoned, and the reviews tell you which. Constant releases with notes that only say "bug fixes and performance improvements" often mean SDK churn and experiments rather than product work.

3. Read the permissions against the features

Every request should map to something you can name. A checklist app wanting contacts has no plausible mapping, and the usual explanation is a bundled advertising or analytics SDK collecting on its own behalf. Detail in app permissions explained.

4. Work out how it makes money

The highest-signal question available, and the store listing answers it: "Contains ads", in-app purchase price ranges, subscription pricing, or nothing at all.

This predicts more about the next two years than any review. An app funded by attention has structural reasons to want more of yours. See how free apps actually make money.

5. Check you can leave

Search the app name plus "export". If your data cannot come out, a bad decision by that company later becomes your problem permanently, and you will keep using it out of captivity rather than preference.

Five checks. None involves the star rating, and together they answer the question the rating pretends to.

If you only have time for one, make it the first. Recent negative reviews are the highest-density source of truth on the entire listing, because they are the only part written by people with no incentive to be there.


When a bad rating is the honest one

The flip side deserves saying, because "ratings are inflated" leads people to the wrong conclusion about low scores.

Some of the best-behaved companies have worse ratings, and the reasons are structural.

Apps that do not prompt. Not asking means your rating is built almost entirely from the self-selecting extremes, which skews low. Choosing not to interrupt people costs real points.

Apps that made an unpopular but correct decision. Removing a feature that was harming the product, tightening a permission, ending a free tier that was never sustainable. Each generates a wave of one-star reviews that reflects one decision rather than overall quality. We wrote about the first of those in the app update that removes features, and the second in the end of the free tier.

Apps in categories where users arrive frustrated. Banking, insurance, utilities, anything you open because something has gone wrong. The app inherits the mood of the situation.

So a 3.9 is worth reading rather than avoiding. The reviews will tell you within two minutes whether it is a broken product or a company that did something unpopular and defensible. That distinction is invisible in the number and obvious in the text, which is the entire argument of this article in one example.


Why nobody fixes this

It is worth being clear that this is not a bug either store is likely to repair.

High ratings serve the platform. Stores want a catalogue that looks good. A marketplace where the average app scored 3.1 would convert worse.

They serve developers, obviously.

And they serve most users, most of the time. For a casual download the rating is a rough filter that mostly works, in the sense that genuinely broken apps do fall below four. The system is bad at fine distinctions and adequate at coarse ones.

So the incentives point the same way and nothing forces a correction. The rating is not going to start meaning more. What changes is only whether you rely on it.


Our own numbers, for the record

We build apps, so applying this to ourselves is the least we can do.

PackPilot and PawDex are both free on Google Play and both young, with small review counts. By point two of the method above, that means our ratings are not yet worth much as evidence, whatever they say on a given week. A handful of reviews can be moved by a handful of people.

Neither app uses a pre-prompt to filter who reaches the review dialog. Neither interrupts a task to ask. That costs us rating points against apps that do, which is a real cost and the honest reason to mention it: an article criticising rating manipulation from a company quietly doing it would be worthless.

Judge both apps by the five checks rather than the score. Update history, recent negative reviews, permissions against features, funding model, whether you can export. That is a harder standard than 4.6 stars and it is the one we would rather be held to.


The short version

  • Nearly every app sits at 4.5 because ratings are collected by prompts developers time, not by users volunteering.
  • Prompt design can matter more to a score than product quality. Pre-prompt filtering routes unhappy users away from public reviews, legitimately, on both stores.
  • Below 4.0, a very small review count, and a sudden drop are the only parts of the number worth reading.
  • Use instead: recent one and two star reviews, update history, permissions versus features, business model, export.
  • Five minutes, and it answers what the average only implies.

Related Reading

Frequently Asked Questions