How we rate running shoes

Every number on this site traces back to a runner filling in a form about a shoe they ran in. Here is exactly what happens to that data.

What we ask runners

A review covers an overall rating plus five dimensions — comfort, cushioning, durability, value and appearance — each on a five-point scale. We also ask context that changes how a rating should be read: how long the runner used the shoe, roughly how many miles they put in, their weekly mileage, their typical surface, whether they would buy the shoe again, and how it compared to the shoe they ran in before it.

That last question matters more than it looks. It gives us direct head-to-head signal between two specific shoes from someone who ran in both — something a star average can never tell you.

Why we do not just show an average

A raw average punishes well-reviewed shoes and flatters lightly-reviewed ones. A shoe with three glowing reviews would sit above a shoe with four hundred reviews averaging slightly lower, even though we know far less about the first one.

So we fit a statistical model that estimates each shoe's true rating while accounting for how much evidence there is. A shoe with few reviews is pulled toward the typical rating for shoes like it; as reviews accumulate, the estimate moves toward what its own reviewers actually said. This is why a shoe's score can shift as reviews come in, and why it shifts less the more reviews it already has.

Ranges, not false precision

Alongside each estimate we publish a plausible range — the band the true rating most likely falls in. A shoe with 200 reviews has a narrow band. A shoe with eight has a wide one.

When two shoes have overlapping ranges, the gap between their ranks is not meaningful, and we would rather show you that than pretend to a precision we do not have. If a ranking looks close, it probably is.

Percentiles, not stars

Scores are presented as percentiles against a comparison set — all shoes, shoes in the same category, shoes in the same product line, or the brand's own lineup. "82nd percentile for cushioning among daily trainers" is a more useful statement than "4.2 stars", because it tells you how the shoe stands relative to its actual alternatives.

What we exclude

  • Ratings from other websites. Retailer and third-party review scores are never folded into a BetterShoes score. Where we show them, they are labelled as coming from elsewhere.
  • Incomplete and unmoderated reviews. Only reviews that pass moderation and are marked complete count toward a rating or appear on a shoe page.
  • Suspected manipulation. Submissions are scored for fraud signals, and anything flagged is held for manual review before it can affect a rating.

How often scores update

Ratings are recomputed as new reviews arrive. A shoe's page shows when its data was last refreshed. Rankings therefore move over time — that is the system working, not a bug.

When we get it wrong

Specs come from a mixture of manufacturer data and retailer feeds, and they are sometimes wrong. If a weight, stack height or drop looks off, or a shoe is categorised badly, tell us at [email protected] and we will fix it. See our editorial policy for how corrections are handled.

Add a review for a shoe you have run in — every review makes the estimate for that shoe a little sharper.