Skip to main content

Reviewed 30 September 2026 · quarterly cycle

Why hoobuy spreadsheet shoes reddit threads disagree with each other

Case one: the high-rated seller who shipped a wrong size

The thread that started this piece was a footwear recommendation with a strong rating and eighty-odd replies. The buyer who posted it had ordered one pair, received a pair that matched the photograph, and rated the transaction highly. What the post did not say, because the buyer did not know it yet, was that the pair had arrived in a variant one size off from the one he had selected on the order form. His review was accurate about everything except the single fact a later reader would need, and there was no way for him to know that at the time he wrote it.

We traced that case because it repeated. A rating on a marketplace listing is compiled across every variant that listing sells, and a footwear listing almost always sells many: colours, sizes, sometimes two constructions under one title with a shared photograph. A buyer who receives a correct shoe in a slightly wrong tone rates the transaction on speed, packaging and whether the shoe looked genuine, all of which can be entirely satisfactory. The variant error is invisible at that moment and becomes visible a week later, by which point the review is written, the score is banked, and the average has absorbed an experience it does not describe.

The structure of the problem is worth stating plainly rather than deploring. A rating is an average of heterogeneous experiences that the rating format flattens into one number, and the flattening is the feature rather than a bug. The distribution underneath can be, and frequently is, bimodal, with a large satisfied group and a smaller group whose orders failed on a specific attribute that the score has no field to express. Reading that single number as a probability of your own satisfaction is therefore a category error rather than merely an optimistic reading.

The rating on that listing was not wrong, and the buyer who wrote the review was not careless. The score simply had nowhere to record the one attribute that mattered, and there is no version of a five-point scale that can carry a variant, a size and a fit at the same time. That is the whole argument of this piece in one sentence, and every case below is a different version of the same gap between what a signal measures and what a reader needs to know.

Case two: the thread that aged badly

The second case is a thread that was correct on the day it was written and wrong by the time we read it. It named a seller, described three purchases in detail, and included photographs of what arrived. Two seasons later the seller had changed the product line, the photographs no longer matched anything on the live listing, and the thread was still visible, still indexed, and still collecting agreement from readers who had no way to tell that its subject had moved.

Nothing about that thread was dishonest. Its author did the work and published the result, which is more than most recommendation posts manage. What failed is that a thread is a document with a creation date and no expiry, and a platform that surfaces documents by agreement surfaces old ones alongside new ones without distinguishing between them. In our own files, the strongest predictor of whether a thread is still usable is not its score but whether it names a date and a variant, and most of them name neither.

There is a mechanical reason the decay is invisible from inside the thread. A seller who changes their product line does not announce it, and the listing that the thread links to is frequently updated in place rather than replaced, so the address still resolves and the page still loads. The reader sees a working link and a confident recommendation and concludes that the two belong together. They may have belonged together two seasons ago and the link has simply outlived the relationship.

The practical consequence is that a reader cannot tell from the thread itself whether it describes a currently purchasable object. We treat any recommendation without a date as a hypothesis about a seller rather than as a fact about an item, and we test the hypothesis before spending anything. That test takes about ten minutes: open the seller store, look for the item or its successor, and compare the current photographs against the ones in the thread. A thread that survives that comparison is worth acting on. One that does not has still done you a service by naming a seller worth checking, which is more than a directory would have done.

Case three: the honest review nobody upvoted

The third case is the one we keep returning to. A buyer wrote a measured post about inconsistent sizing across three orders from the same seller. It was specific, it listed the sizes ordered and the measurements taken, and it ended with the observation that the third pair fitted differently from the first two in the same nominal size. It received almost no attention at all, and by the time we found it, it sat below a single-line post that said the seller was solid.

The reason looks obvious in hindsight. The post did not offer a purchase, a warning, or a resolution. It offered a distribution, and a distribution is not what anyone is looking for when they open a recommendation thread. Readers want a yes or a no with a link attached, which is precisely the format that hides the failure mode they are about to experience. The thread format rewards confidence and the underlying reality is a spread.

We are not arguing that the audience is at fault, and we are not arguing that the careful reviewer was unlucky. We are noting that the review containing the most durable information was the least rewarded, and that this is a stable property of how recommendation threads work rather than an accident of one forum on one afternoon. Any reading strategy that ignores it will keep rediscovering the same three cases under different names, and will keep being surprised by them in the same way.

The post also had a second disadvantage that is worth naming, because it is the one that stops people writing reviews at all. It contained no conclusion. It described a pattern across three orders and left the reader to decide what it meant, which is honest and unhelpful at the same time. A post that says the sizing was inconsistent across three orders is more useful with one added line stating which of the three fitted correctly, and that line is the difference between a caution and a finding.

What the three cases share

All three cases are the same failure at different latitudes: the visible signal and the underlying mechanism are misaligned. Ratings measure whether an order completed and shipped quickly. Thread age measures when a document was written. Upvotes measure agreement, relief or warning value. None of the three measures whether the item you are about to buy will arrive as the item you intended to buy, which is the only question you actually have, and that misalignment is not fixable by reading more carefully or by reading more threads.

It is fixable only by adding the fields the signal omits, which is why our own buying notes carry two columns that no thread format has: the exact variant ordered, and the date the order was placed. With those two fields a review becomes checkable against reality. Without them it is an impression with a score attached, and the score is measuring the wrong event entirely.

So we read threads for names and then verify independently. A thread tells us which sellers deserve a second look, and that is genuinely useful information that is hard to obtain any other way. It does not tell us what a seller will send next month, and treating it as though it does is how a buyer eventually contributes a case one of their own to somebody else thread.

The reading habit that follows is small and specific. When a recommendation is decisive for you, ask three questions of it before you spend. Does it name a date. Does it name a variant. Does it describe more than one order from the same seller. The third question is the one that matters most and the one most people skip, because it requires reading the whole post to answer. A single-order review tells you that one parcel arrived. A three-order review tells you what the seller does when something goes slightly wrong, which is the situation you are actually trying to predict.

What we measured ourselves

Across thirty high-vote footwear recommendation posts we logged in five communities, twenty-four named a seller without stating either a date or a variant identifier. In the six that stated both, the recommendation could be checked against a current listing; in the rest it could not be checked at all.

Basis: Editor log of publicly visible recommendation threads collected during 2026 Q3; each post scored for the presence of a date, a variant identifier and a seller identifier.

Where to go next

Outbound links to kabosheet.com may earn this desk a referral credit. It does not change what we write, what we measure, or what we mark as unverified.

Full disclosure