If someone complains all the time about everything, the people around them quickly learn to ignore it. Our systems should too.
If someone complains all the time about everything, the people around them quickly learn to ignore it. Our systems should too.
In other words, we successfully encode the difference between "wow, juhzy only complains when its really bad" vs "ignore juhzy, they complain about everything." This vs. the more naive and destructive "5 bad reviews and you're out" we seem to do too much of now.
I suppose it should work the other way too. Someone who always "5-stars" everything should be ignored as well.
The point is that the system right now seems to have a negative bias and is able to do so because of a glut of providers which it burns through like an expendable, renewable resource. (Which morally sucks because they are real people with real lives.)
Yelp has no idea how often I eat out at restaurants.*
* just kidding; of course they do; they subscribe to a feed of my location data published by some game on my phone.
When I say "seldom" I of course mean "uses the service often, but seldom complains". The unit of time in this exponential weighting is not chronological, but uses of the service.
This of course is harder with yelp than uber.
Some people can complain about anything. (When I'm in a grumpy mood, I certainly will. Puppies? Too wiggly. Sunshine? Too bright!) When I read a review, my goal is to find out what I would think of the place or thing. The reviewers I look for are balanced and thoughtful, able to point the good points about something they hate and the bad points of the things they love.
I have a very smart friend who apparently has never liked a movie. If one comes up, he will have a complaint about it. If you see one at a theater with him, before the credits have finished rolling, you'll be getting a list of his issues. It's exhausting. And it means I'll never listen to him on the topic of whether a movie is good.
It's the classic stopped clock: it's right twice a day, but you can't know which times, so why look at it?
From a perspective that is external to your inner monologue, I have no way of knowing. Even if your criticisms are valid, I know too many people who seem to never talk about the positive parts to know whether you're just perpetually unimpressed. Or, more likely, you're the type that never talks about the nice parts. I'm guessing this because of your comment elsewhere in the tree that specifically asks why one should compliment when something is doing what it's supposed to do. (The answer is that people usually look at reviews for confirmation that X is as advertised.)
Or there's actually nothing redeeming about the subject of your review.
So yes, taken in context I would end up throwing your opinion away. In practice, I don't check every reviewer's review history, but maybe that would be a useful signal to see.
A summary of the opinions of many strangers has some signal.
The facts of a stranger have some utility.
For the record, I rarely leave reviews. When I do, though, it's only for the exceptionally bad or the exceptionally good.
It's hard for a product or service to exceed its expected value by more than a fraction or small multiple, while it's easy to cause misery well in excess of a dozen or even a hundred times the expected value. I believe that's more a reflection of it being easier to destroy vs create rather than a reflection of psychology. Combine this distribution with a heuristic to only report deviations from expectation in excess of a minimum threshold and the "average review score" will reliably undershoot the actual quality.
That's only a problem if you want to interpret the average review score as an absolute measure of quality, though, and I don't think anyone really does. Most of us are more interested in communicating and informing our decision processes than in passing moral judgement, and if our goal is to optimize communication then we should expect negative reviews to dominate the discussion because they're inherently capable of more meaningful excursions from the mean.
I agree with your overall point: it would be fantastically useful to be able to contextualize reviews against reviewer psychology. That way I could ignore both shouting from negative-nellys and forced positivity from those who feel compelled to balance the universe :-)
In reality, however, it seems that positive reviews tend to dominate. Using Google Maps reviews as my barometer, I hardly ever see any place rated less than 4.5 stars. So, I tend to think to myself "4.5-5 stars: might be good. 4 stars: probably okay. Less than 4: maybe steer clear."
Though, in practice I disregard reviews, take a plunge, and then decide on my own. Often I find myself in conflict with the average majority opinion.
If I have a specific complaint for a place close to my heart, like a coffee shop or restaurant or local shop, I’ll talk to the manager, privately and calmly, and be on my way.
It's interesting that your main criteria is how you feel you were treated, though. Discussed in the article:
"Restaurant reviews in which people sound traumatized by perceived injustice don’t tend to comment much on the food — it’s usually the perception of being treated rudely or uncaringly that seems to have pushed people into processing by writing out their feelings in a public forum."
I honestly feel the same about poor service. I'm usually accommodating and understanding, but I'm no monk. Of course there are times when I feel either the poor food or poor service merit some mention.
Most of the time, however, I think "they're human, going through human things. No big deal."
Doesn't mean you have to, just means that it'd be more helpful to others if you write "it actually is what it claims to be" reviews.
Do you consider them equally attractive?
> If someone complains all the time about everything, the people around them quickly learn to ignore it. Our systems should too.
That assumes that "too much" negative criticism is false and merely the result of a disagreeable personality. I think that assumption is false.
I think that's false, people who frequently post negative reviews may just be the kind of people who have legitimately higher standards vs. people who are happy with even objectively crappy things.
I don't think the frequency of negative reviews, by itself, gives you any real information about the quality of the reviewer.
Amazon, as we all know, basically expresses quality as a raw average of stars plus a histogram, with no correction except for removing reviews. Meanwhile, BeerAdvocate is essentially just a labor-of-love beer tracking site, but it offers multiple secondary stats about ratings. Each beer gets a percentage deviation stat (pDev) to show how varied their rating are, each specific review lists its deviation (rDev) from the product's average review, and each reviewer's profile offers their frequency of reviewing above, below, and inside the average window (|rDev| - pDev > 0, then rDev - pDev).
I don't think anything is actively done with those stats, but even that's enough to spawn forum threads where reviewers discuss which beers are most controversial, which beers they gave outlier reviews to, and whether they're typically harsh or generous reviewers. The site also recommends rating sub-categories to get people thinking about different aspects of a problem, while Amazon is laden with reviews that are either myopic (e.g. that XKCD about a tornado tracker) or completely off-topic (e.g. about the seller instead of the product).
A site that wanted to go a little further could use the same stats I mentioned for spam detection (there are several papers on doing that effectively) and for score correction. (Basically, take a user's average deviation over all reviews and use that to scale or adjust their impact on averages.)
Given how easy all that is, why do Yelp and Amazon consistently have some of the least useful reviews and averages of any site I know? I suspect this article nails it - casual readers appreciate the simplicity more than depth.
The main drawback I see (assuming that you are seeking an honest picture of that provider's quality, and not a tool for punishment) is that you don't meaningfully capture people who only write reviews to flag serious problems.
Yes - reading the comments here I'm realizing a major problem with any reviewer-centered system is that people decide whether to review on hugely varied conditions.
An always-five-stars reviewer might just be easily impressed (or a fraud), but they equally might subscribe to "if you don't have anything nice to say...". And quite a lot of people write exclusively bad reviews, but it's not obvious how to discern grumpy reviewers from people who only speak up about major issues.
A partial fix might be available by analyzing how far a given review is off user's average difference from product average, which could discern a five-star bot from a person who only reviews great products. But even that doesn't solve the XKCD problem where a product has median-case appeal but a high rate of critical failures. In true "what can't meta-analysis fix?" style, this could be improved by looking at a user's average distance outside 1SD of the mean review, or perhaps by special-casing products with multimodal reviews.
Of course, it's deeply unclear how to convert this to an output. Scaling reviews based on reviews sounds like a nightmare, reviews shouldn't be differential equations and no one wants to see 4+ layers of statistics to buy a new lamp. Perhaps all of the indirect work could be done off raw ratings behind the scenes to produce a general "adjusted average" for display?
(More realistically, the serious-problem case only seems solvable by reading text reviews, and just devaluing outright fraud and always-angry cranks would be a massive improvement over existing systems.)