Ratings inflation on Uber and elsewhere
qz.com
qz.com
My uber ride either gets me to my destination safely, reliably, and timely, or it doesn't. It's just a yes or no, and should be rated as such. This would smooth out the subjective reviews and give drivers a much more fair rating. Support requests can take care of any specific complaints. Likewise tips can do the same for positive signals.
The trouble is that rides are not reproducible readily. Conditions change (congestion, path taken, mood of the client or driver)
Thirds means you have to accumulate a large amount of any kind of rating for it to maybe mean something.
Drivers could game the system by servicing "good drives".
And yes, conditions change, which is why the 5 star system is too vague and useless. What is the difference between 4 or 5 stars? And why would this mean the same thing to anyone else who gets a ride after you?
It's best turn this into a binary signal instead. Any system can be gamed, but it's much better when there are less random choices.
So let's say, if you have less than 30% or something like that, you get out of Uber. And if you have more than 85%, you can be part of Uber X VIP (at least here in Brazil, if you use a lot Uber, you earn one month of Uber X VIP, so just works with drivers with top ratings. If you use a lot again, then you'll earn another month and so on)
If Uber really cares about why the service was bad, they can always break out multiple yes-no questions. (They probably don't, though.)
A binary choice would result in most people actually saying the trip was ok, even if a little late, especially since they're also in the car and can see the situation for themselves.
Otherwise the trip is 'not ok' and contact support if there's a big problem.
There more it becomes clear to me that giving someone a not-ok rating threatens their job, the more that I don't want to take that nuclear option.
Besides, I don't want to be connected again to someone I ranked negatively once (even if other people gave them 5 stars). I don't think such blacklisting - even in mild form, eg. with some sort of expiration period - is implemented. [EDIT: apparently it is, although the app never explicitly states that].
This is to say that inflation isn't the only, or even main inefficiency of Uber's system; from my perspective.
1. Uber keeps track of driver's cancellation rates, and if they do it enough, they'll receive notices to timeouts to deactivation.
2. If you rate a driver/rider 3 stars or less, you will never be matched with them ever again.
Ad. 2. Thanks for the clarification!
I think I'd rather just not rate at all, having to take into account rumours, how other people rate, how I think Uber might use the data, etc, is just too much to bother with.
They should probably adopt a thumbs up/down rating system like Netflix switched to. It' much more obvious what the rating should be, allows them to remove bad drivers from the system, etc.
I would happily pay the minimum fare to give these drivers a 1 star review.
Just because someone has a beef with Uber doesn't mean it's a good excuse for making me late by trolling me into a transaction they are not willing to fulfill to begin with. I'm not Uber. It's like getting a job at McDonald's and sabotaging from the inside (because they made your hot-dog stand go bust).
If that's the level of social ethics they demonstrated at their prior reasonably good jobs, I'm not too surprised about their losing them.
These people are being sent into a race to the bottom by an outsider. An outsider who just happened to have attracted a big pile of money in a faraway country.
It's not the same thing at all. Strikes may delay people, and they do, but they don't involve deception.
Strikes, especially, though not exclusively, transport strikes, are announced - and announced in advance.
Even if - say - your local surgery strikes, surely they won't book you an appointment, then trick you into sitting in their waiting room, while noone is actually going to see you.
The rebellious acts of these people and the collateral inconveniences they cause are rather small compared to the moral injustice that is imposed on these people.
And yes, you, as a client of Uber are also guilty to some extent.
Perhaps ask yourself this question: what amount of money would be necessary to put me out of business? Imagine that someone stands up and brings together a bunch of VCs and collects that money. Imagine that you have invested a lot in your current business, and you have nowhere to go if you lost that business. That's the situation we're talking about here.
I meant a surgery as in primary care physicians or dentists, not surgeons. It could be a barber though, or whatever really. You're missing the point I laid out explicitly - strikes aren't based on deception, and are announced.
And it's for a reason: to maximize the impact it has on the employer, while at the same time minimizing it for the "guilty to some extent" average people. Which is the exact opposite of the strategy you described.
Uber does not fight fare, so therefore, neither should those fighting against them.
As others have stated, a single-dimension scale is difficult in Uber where it compounds the basics (get to my destination in a safe and hassle-free way) with the extra mile (some drivers are extra nice, offer you water and snacks and let you play your music etc).
If everyone who fulfills the basics are supposed to get five stars, how do we signal the extra mile except for tips?
My guess, though, is that the driver would prefer the tip.
>If everyone who fulfills the basics are supposed to get five stars, how do we signal the extra mile except for tips?
Which is asking for ways of showing appreciation without giving a tip (or perhaps even in addition to giving a tip)
I gave a useful answer (the 'complement' feature that both uber and lyft have) as well as a mildly snarky joke ("My guess, though, is that the driver would prefer the tip.")
I mean in the case of Lyft I give them money. Why should another adult care how I "rate" them?
Some cultures in America naturally tend to 'Fantastic' as their base line of descriptive language. This probably translates to ratings too.
I feel like it's really got to be an exceptional case to warrant "burdening" someone who is mostly an acquaintance with your drama.
For ratings, as the article discussed, anything less than 5 stars is a black mark for the driver, so those ratings are reserved for the exceptionally bad. And even then, exceptionally bad I blame on the driver (not on Lyft/Uber). Confused by GPS directions, stopping halfway down the block because they app says so, wild detour because of bad Line routing... these have no impact on my rating.
Being told in clipped, broken, heavily-accented sentences: "Go to Uber", "call Uber", "it's with Uber". But Uber where, exactly?
While I was calling all the nearby Uber locations, I kept thinking, "If this guy got 4.6 stars, do the stars mean anything? At all?"
I would retrieve my stuff, finally, when his daughter took the phone explaining that he accidentally took it and then dropped it off at his local Uber... two-or-three cities away.
All in all, an unpleasant eye-opener.
Edit: I think I see why the companies don't need to be too concerned. They only need to remove the worst actors. The ones that could get them into media (e.g, someone ends up with an Uber driver showing criminal tendencies). For that, this system probably works. Anyone who didn't make you feel like your life is at risk is a 5-star. This situation is unlikely to change unless there is competitive pressure on these companies. So it is very important to prevent these companies from monopolizing the sharing economy. Once older cabs and hotels are gone, and these companies have monopoly, we could be left with no option but to take an Uber/AirBnb with unreliable, terrible and probably even dangerous "service" and no legal recourse.
Uber has no incentive to keep bad drivers around; after all, it costs them little to fire them and onboard a new one, and there's plenty of people willing to drive a car.
Besides, it's not such a big problem. What ends up happening is that they raise the bar; e.g. an average of 4.6 becomes the equivalent of a 3. A friend of mine has a house on AirBNB with a rating of 3.9, and it's essentially invisible - you'll only find it if you set very specific filters.
Specifically, they can use percentile ranks to do the "translation": https://en.wikipedia.org/wiki/Percentile_rank
Hotels have had rating systems for a long, long time. Since before AirBnB existed. And most people don't think a star rating is a substitute for an actual vetting and licensing process.
https://daringfireball.net/linked/2017/03/18/youtube-thumbs-...
Maybe for you. The Netflix ratings were always a personalized rating (which is explained in one popup on sign-up but never mentioned again). I think people not understanding that was also one of the reasons they abandoned the star ratings and changed the language to "90% match" in the transition.
how am I as a customer supposed to make an
informed decision?
There are browser extensions that will add IMDB/rottentomatoes ratings to the Netflix homepage - if you think those are closer to what you want.(Once I saw them side-by-side I saw that a lot of the content on UK Netflix was poorly reviewed - so I can believe maxyme's claim Netflix wanted to boost the scores of poorly reviewed content.
Was your driver late?
Was your driver polite?
Did you feel the way they drove was safe?
Did you have troubles communicating with the driver?
(Yes, no, kind of).
ASK FOR FACTS, NOT OPINIONS.
I understand customers can't be bothered to complete a full-blown questionnaire, but the system could pick one or two questions at random after every ride. It could even incentivize answering two instead of one, or three instead of two (as in Google's "Local Guide").
Over time, data will accumulate.
I was pretty pissed off with a number of things, hence leaving the company. But the exit interview asked very specific questions, none of which addressed the reasons for me leaving.
Them: "We understand. On a rating of 1 to 5 stars, how many stars would you give this job? This will be part of your file here."
You: "Well obviously 5. Thanks for the job and good luck with your other candidates."
This would be a poor question. Any perceived lateness is because Uber underestimates the time-to-pickup.
In Manhattan they pretty consistently underestimate by 50% it feels like.
I wouldn't even the least bit surprised if they already figured out that people are more likely to take trips if they intentionally underestimate.
Without distinguishing only broad actions are possible. Like broad judgements where you don't know what it is supposed to say.
> Uber said in July that it would make riders add an explanation when they awarded a driver less than five stars.
This acknowledges the inability to distinguish what is ment with the judgement and one of their tries to fix this.
Eg. if somebody apologizes for keeping me waiting, it's a different experience already.
Also some drivers "arrive at the destination", but not the exact destination where you said you were waiting at, and the way we manage to handle this together is part of what I consider as being late or not, too.
Three buttons.
Acceptable
Exceptional. Why?
Below average. Why?
* Great
* Typical
* Horrible
(I intentionally avoided the word 'normal' here due to it implying an average instead of a median.)
And a proposal for what to do with these three inputs:
Have three internal counters:
* Good
* Neutral
* Bad
'Great' increments 'good' by 1/2, decrements 'neutral' by 1/3, and decrements 'bad' by 1/6.
'Typical' increments 'neutral' by 2/3 and decrements 'good' and 'bad' by 1/3 each.
'Horrible' increments 'bad' by 1/2, decrements 'good' by 1/3, and decrements 'neutral' by 1/6.
Show the end users and the drivers the rating graphically by just stacking the three counters next to each other horizontally.
Edit: Note that I indeed mean INCREMENTS, I do NOT mean "multiplies the current value by".
Three of the four examples you give - polite, safe, good communication - are mostly opinions, though.
I mean, there are exceptions, but at the individual contributor level, that's a pretty good heuristic.
(I think our profession is vulnerable to unadjusted people making the rules for the rest.)
It's a hard problem. And especially in cases like Uber, I don't think there's a technical fix: as long as the company puts the cutoff for continued employment above the second-highest rating, there's a high perceived moral cost to giving a rating less than the maximum.
You can just cap the max-value at a 5, to avoid someone getting a "astronomically great rating".
Rating normalization isn't meant to prevent rating-fraud, nor does it make the problem significantly better/worse than it currently is. But it will solve the problem of rating inflation.
Not to mention that once users know that their ratings are being normalized anyway, they won't feel the need to give everyone a 5-star rating.
And note that tipping also saw inflation. If you read old newspapers, the standard amount used to be 15%, and 10% before that. At least with stars it's capped at five.
4.5: Should be fine if there are many votes and the top ones don't look forged
4: average
3.5: Avoid at all cost
People are so bad at giving objective ratings. I often see people mostly rate their friends and the good time they had: "had a very good evening with my friends, blabla" Ok, what about the food, the service? seriously.
Cleanliness: 1-5; Reception: 1-5; Location: 1-5
There were about 6 or so of these and I gave a 1 for each (since zero wasn't an option). MY final score was 3.5.
That said, reviews on airbnb are generally not worth the paper they are written on.
If the reviews all say "Place was a bit messy but the price was right and the location was great" versus "too bad this beautiful home is so far from the train station," you'll know whether or not the place is right for you.
And people generally won't downvote for a negative they knew going in.
It’s what works in the academia, where anyone can get aribitrarily close to a 4.0/4.0 GPA, if the grades are inflated enough, but clearly only 10% of students can fit in the top 10% of the class, making this latter rating much more significant.
Simple example (I am sure there are better ways of doing so) after ten Uber rides you are given five stars, and you can distribute that to the drivers you liked the most.
You have to anchor the ratings or they have no grounding in reality. (Produce a few models of bad service and drivers and excellent service and drivers.)
That is, unless you really want to chase some sort of unattainable perfection.
GPA in practice has numerous problems on which I could write a whole other book, but where it fails it is often for social reasons which would equally apply to and be gamed by people under the "limited resource/grade to the curve" model.
You will see people start to game away and politicise how the "stars" are awarded and distributed in both. In the later specifically, you run the real risk of having to start to fail genuinely good people and pass genuinely bad people. And there is no universal algorithm used by people in their subjective distribution of stars that allows one to compare or interpret the distribution of stars after the fact or between different "ecosystems".
With sufficient training and effort 100% should be meeting or exceeding the minimum quality standard. It should not be forced into a normal distribution.
Now you have to decide does no rating mean good (nothing to complain about) or bad ("No comment.")
But, some people never give 5 star ratings. The way things are now? I think that what a driver's rating is depends more on if they get me or if they get a "nobody's perfect" rater than it depends on their actual driving ability.
If the ratings were adjusted by the rater, I think the end result would be a rating system that better reflected the things they are trying to rate than the luck of who they get rating them.
For those who don't know, when you get your car serviced you're asked to fill out a satisfaction survey. And they all, ALL, have this disclaimer at the top: if you can't give us a 9 or 10, talk to us first. Because they're pretty heavily penalized if they don't get an average rating of 9/10 or better.
It's this bizarre game - they could just say "we want to make you happy so talk to us if you have a problem" but corporate HQ doesn't trust local dealers to do this so they survey them. And of course local dealer game the surveys.
There's two answers to that question. If they should, you give them every maximum numeric rating. If they shouldn't, you file a customer service complaint.
So, it's yet another dark pattern.
I admit until I was asked to judge an event and received a short training session on judging I also fell into this category. I now realize how bad I was. There should be some sort of how to rank/judge/review things class that everyone should take in school.
I had a perfect 5 star rating for many years, then I visited a 3rd world country where I had some problem with the Uber driver that was supposed to pick me from the airport. He gave me a one star rating. Then Ubers from the same country all gave me bad ratings. I have no idea why. Perhaps because I didn't leave a tip? I don't know.
Then I came back to Austria, my rating was around 4.90, but then it continued to drop slowly, and it's still dropping. Now it's 4.62. I'm having trouble getting rides, presumably because drivers reject me. I used to be picked up by 4.9+ drivers, but now it takes up to 10 minutes to find a ride, and they are all 4.4 drivers.
I don't understand why my rating keep dropping and dropping, especially since it used to be a constant 5, in the same country, for many years.
I'm guessing that people (drivers) rate similarly to what the score is currently. If it's between a 4 and a 5, that's what you'll get.
I was thinking, I wonder what I've done or vibe I gave out to merit that? I was perfectly polite, I was at the pickup spot, I was very accommodating to where he wanted to drop off. We hadn't had any disagreement, I'd been perfectly pleasant though not talkative. I guessed in the end that an outstanding passenger to him probably gives him an extra tip.
I expect there's other info that could be used to calculate the score as well. Average tips compared to other drivers in the area comes to mind, although that could have negative consequences (mainly increased pressure to tip). Perhaps if it were combined with other inputs though. I'm sure there are other options too.
https://medium.com/@homakov/why-airbnb-reviews-are-flawed-b8...
Was your job/ride successful? 0-110% ? (and probably use intervals of 10).
Then most people for the average ride where it works will give 100% as the job was done. If the driver went above and beyond, you can give them 110% If you pick 110% an extra option asks, "What did the driver do that was above and beyond?" Which stresses that 110% is for exceptional service, and also collects what makes it exceptional.
For rides where the driver didn't do the job properly, like took a wrong turn, was late, etc. You can mark it accordingly, 90%, 80% etc
Actually making it even less options might make it simple and removing a number system all together.
Scale would be:
- Horrible/bad/did not do the job at all/etc - For when the job/ride did not completed or something truly awful happened. Can play with wording. When choosing this you get a big box to type in the issue and a customer support service should reach out and see what's wrong. - Did the job but slight issues - When choosing this you have to say what the slight issues were, so there is some friction still in choosing this option - Did the job, no issues - This will be the default, friction-free option. So if they pick this they don't have to type anything else. Quick and easy. - Exceptional service - If they pick this they are also forced to type in what the exceptional service was.
You could even socially tweak/test it with forcing people to tip if they pick exceptional service (so effectively putting their money where their mouth is, if it's exceptional you need to pay why). With countries where tipping is optional this is a good play (but in USA maybe not so).
The tipping part is just an idea. But the idea is to make choosing exceptional a bit more friction then just choosing job done. And also how to translate that culturally to other markets. By making them type in more for those situations it should funnel only users willing to type to give those ratings.
And an option is for users to see not only the ratings of drivers but what other people typed in the exceptional (or typed in the slight problems) categories.
4.5 - he (or she) is gonna be polite
4.6 - ooh, I (mildly) wonder why that is
4.7 - nothing
4.8 - nothing
4.9 - nothing
5.0 - oh gosh, they're still learning
I guess in Uber’s case this might be implemented as: randomly showing previous riders a rating for a driver they had, and asking if they agree.
Making the reviews private (as the study suggests) may solve the social pressure issue but it will damage trust (how do I know the 2 stars are legitimate and not a company's way to oust specific contractors/landlords/etc?)
I also think it's completely fine that people are rating apps and toasters more negatively, because most of the time, the source of dissatisfaction (e.g. a bug) can be weeded out once and for all, while being an Uber driver is emotional work 10 hours a day.