The Never-Ending War on Fake Reviews
newyorker.com
newyorker.com
My first project at Google was working on fake reviews on Maps. ML was certainly useful for some things, but at the end of the day sometimes you literally have no practical way of knowing if a review is fake or not. Was this written by a customer or by the business owner's brother in law? Or by his competitor down the street? Who knows? Which means it's pretty hard to get good, complete training data to train your model on.
Of course, there are classes of fake reviews that are easy enough to detect. But as the fake review-writing AIs get better, I don't know how the anti-fake review AIs can win.
Since humans are good at creating lies, and that we all get fooled at some moment in our life, it means the problem hasn't been solved for millennial and that it's not an AI problem.
Sure you can throw at the AI way more data than we used to have, but even totalitarian state didn't manage to silenced the opposition so I doubt a business bound by law can do anything about it.
Google protection measures are already becoming super invasive and instead of helping me, has locked me out of my account several times because it detected I was a fraud.
There is a limit on what you can do properly here.
Can't you check the account itself? If the account's behavior over time shows signs it's an actual human user instead of a bot.
This can't filter the owner's brother in law, but if bot accounts can be filtered then they cannot leave lots of reviews either, because the owner does not have hundreds of relatives.
The account would need a bit of history, e.g. one or two weeks, before allowing it to write reviews.
The user would see the review, others wouldn't, and after 1 or 2 weeks when the user activity confirms the user is not a bot, make the review visible for everyone.
Surely, you don't think that average users check their reviews from private browsing tabs to make sure it's there and stuff.
Do you not care about lying to them just because they are "average"?
It may not be the best solution, every solution has drawbacks, but do you have an idea for a better solution to filter out fake reviews? If so, please present it.
Account level signals are nice but insufficient. There's a very long tail of low activity accounts that are difficult to distinguish from a bot. You can't really classify them as bots without generating a ton of false positives, and in general you'd rather let a possibly fake review through than piss off a genuine first time reviewer.
Google does have a big advantage in verifying good accounts through other activity, but that advantage shrinks dramatically for new accounts. There's still some things you can do, but at the end of the day you're only willing to assign so much risk to a brand new account based on metadata.
A google account cannot really be low activity if its owner browses the web, watches youtube, etc. The tiny minority who logs out of google when browsing, etc. could not leave a review, but it may be a small price for more valid reviews.
The strategy at many studios was to use a barrage of imperfect solutions which raised the effective per/unit cost of botting to something most players weren't willing to accept, then hunt the remaining "big-time" offenders using a separate set of tactics, like identity "fingerprinting", correlating billing information etc.
I think in increasing the number of reviews from verified accounts (this account verified to be connected with human of name X), the rest of the anonymous or pseudonymous reviews may be looked at with more scrutiny.
Right now, let's assume that the default for social networks and other places of interaction on the internet is something like 10% verified accounts and 90% unverified. We know that within that 90% unverified, there are many real humans with their real names. So we spend a lot of effort to parse that 90% to find the real humans, the ones pretending to be others, the bots, the anonymous ones, and others. Companies are using ML, individuals like me just guess and try to only accept people that I've met in person and avoid platforms where bots seem to hang. However, with sites like Amazon, I have to hope that Amazon has filtered out the reviews from the fake accounts. Yes, there will be fake reviews from verified accounts, and that seems like another issue, and maybe less prevalent, I'm not sure.
If the percentage of verified accounts flips, from 10% verified / 90% unverified to 90% verified / 10% unverified, I think many of these fake reviews filter themselves out. We would trust that this specific person is saying this, and still leave room for pseudonyms and anonymity, but be more clear that one is using such measures to hide their identity.
How do you think this would impact fake reviews on Amazon? What secondary impacts do you see more verified accounts on Amazon and on other platforms having on discourse? I hope this post wasn't too long, again, I'm new to HN and hope I'm staying within bounds.
The bad players will find a way to get verified and now you're worse off than before (fewer legit reviews plus bad ones with a fraudulent stamp of approval).
Whatever the solution you have to take into account that bad players have a higher incentive to pass your test / jump through hoops than legitimate people do.
You buy something, and regardless of what it is, you get automated requests to review the purchase. Sometimes before the item has arrived. Practically always well before you have managed to establish whether the purchase was any good or not.
This is why I ignore pretty much every single please-review-me prod from Amazon and their retailers. Send me a request to review before I could have had proper time to evaluate the purchase in practice, and you get bucketed with other entitled f--kwits. You are also likely to lose my future business.
I believe Amazon could improve the S/N ratio of their reviews if they actually considered how long it takes to test out any of their purchases.
Hiring a few interns to sign up for all of these facebook group and track who is participating would go a long way. Doesn't seem like Amazon really cares much.
I'm reminded of videogamedunkey on YouTube and his video about game critics[0], where he goes into some detail about the preferences and integrity of the critic being just as important (if not more) than the review and rating.
Professional critics build a strong reputation around their taste and preference and they gain notoriety from doggedly sticking to those principles. So you'd likely trust the opinion of Roger Ebert or Mark Kermode if you share their taste in cinema, and even watch something you normally wouldn't if even they recommended it. You're very likely to subscribe to a publication whose critics align with your own tastes because they're effectively curating content for you in the form of recommendations.
None of that applies when you have reviews from a succession of total strangers - you're not going to research dozens or hundreds of commentators to establish a logical consistency in their point of view and decide whether or not their tastes align with yours. More often than not the reviews are low quality and low value, spread across a five-point scale but essentially treated as a binary like/didn't like system.
At that point, crowd-sourced reviews tell you nothing you didn't know already: some people enjoyed it for one reason and others didn't for a different reason. How do you know if those reasons are legit or authentic when they're not fake? How do you know if they weren't gamed or incentivised somehow to inflate expectations? Which ones can you trust?
I suppose they boost sales but just like advertising, that doesn't mean they're automatically beneficial to the consumer. It's just another attempt at manipulation.
Aggregating what they regard as reliable reviewers is pretty reasonable.
For general products it's much harder. There is less of a market. But perhaps it'll come. Aggregate Ars Technica, Tom's Hardware and some other hardware review sites maybe?
About the only thing you can do is find a particular writer or site that has a reputation you trust, but they might never review the specific product you're interested in. And when you're trying to buy something in an entirely new category you're not familiar with, it's hopeless.
In our capitalist society the easiest way for something to spread is a profit motive. Whoever spends the marketing resources to popularize it will have to recover their costs somehow. A lot of people will not pay for yet another subscription, and marketers will pay a lot to make their company look better.
Concentrated benefit and diffuse cost is how a minority can maintain a globally sub-optimal status quo.
Huh? I have bought hundreds of items the past year for my small business and returned probably a dozen. I've never had any problem. I can even drop off the item at my local Kohl's and they package the return and ship it at no cost, even if it's my fault for the return.
Edit: I'm talking about Amazon.
Because of a high percentage of the population being present, there is now substantial power to be had by influencing the discussions that take place.
https://plus.google.com/104092656004159577193/posts/RCyGi3HQ...
https://old.reddit.com/r/dredmorbius/comments/5wg0hp/when_ep...
A big part of the problem with review systems is the one-to-many nature of nearly all of them: when a person posts a review, that review and its score can be seen by everyone. This leverage makes it very efficient for businesses to game the system, as a small amount of fake information can "infect" the purchasing decisions of a large number of users.
So, one alternative might be a many-to-many review system where you only see reviews and ratings from your network of friends/follows (and maybe friends-of-friends, to increase coverage). So essentially Twitter, but with tools and UI that focus on reviews and ratings. That way, fake reviews could only affect a limited number of people, making the cost/benefit calculus much less attractive for would-be astroturfers and shills.
I did my dissertation on this topic - text only analysis. I used a dataset that was commonly used in the beginning, but there are some issues with it. I plan to extend this to real reviews, as in 80 million Amazon ones (when I get the time).
Text based features are useful, but non-text based ones are even more so. Even spamming groups can be detected; at least there has been research into that. Combining all the techniques in an ensemble would be productive - but is it really in Amazon's interest? My sense is whatever they do, they pick the low hanging fruit and trying to process every review that comes in would require a lot of CPUs perhaps. But stuff like floods of reviews for new products that are fairly similar should be easy to detect. Perhaps they are relying on Fakespot and reviewmeta to do the heavy lifting.
could start a reviewer guild and use crypto signatures to verify their guild membership is up to date and still valid. hoping the guild has the incentive to stay honest etc.
I'd argue that is exactly what The Wirecutter is. https://thewirecutter.com/ https://en.wikipedia.org/wiki/Wirecutter_(website)
They solve the reviewer problem by having in house staff write the recommendations (reviews). Those staff then find and bring in experts. Virtually every page has a "Why you should trust us" section listing experience of those who contributed to the article. It does mean you get credible opinions, but also that you don't get the "wisdom of the masses" such as at Amazon.
I'm generally a fan of the Wirecutter while remaining somewhat skeptical of the motives behind their reviews. My skepticism hasn't changed since NYTimes took over, but I do still visit the site when looking for a specific product. Nonetheless, they've made a good deal of affiliate $$ from me, and I haven't been severely disappointed yet.
e-commerce websites pay lip service to deleting fake reviews but higher product review scores result in more sales. Even if this doesn't result in intentionally ignoring fake reviews in pursuit of short term sales growth, note how Amazon.com like all other five star rating systems suffers from massive score inflation. Most have a 4.5/5.0 rating. 4.0 indicate some potential problems (or less sophisticated customers), and 3.5 means it's sub-par.
A lot of online stores simply have no customer review section because they have rationally determined that for them decreased conversions due to bad reviews and moderation costs exceed increase in conversions due to good reviews.
Indeed.
I liked how Goodreads had their original rating scale done: 3 was good, 4 was very good, and 5 was truly exceptional. You weren't supposed to give 5 to more than maybe 2% of the books you'd read.
Then the got bought out (by Amazon) and the scale got soon enough devalued.
I still think that one potential solution to the scale inflation would be to consider the grade distribution a person uses. If all they give out is ones or fives, their reviews should have near zero weight. If they give out a more balanced (on a ~Gaussian curve) reviews, then their rare extremes should be weighted much higher.
ReviewCoin could be redeemed at participating online retailers for goods, thus allowing reviewers to buy more products to review.
This is a silly idea but it makes me smile.
The customer ultimately determines their satisfaction level, even if they are complaining that a software product clearly described as Windows only is incompatible with their Mac. Or their "wireless" machine doesn't work without the power cord inserted.
Well, at least until this is perfected (and it will be):
https://www.theverge.com/tldr/2018/4/17/17247334/ai-fake-new...
The potential downside is that people can sometimes become overly trusting of truer to life higher bandwidth mediums. e.g. Fake IRS, grandson needs money for bail and baby (cue crying noise).
You could imagine a GAN-like ecosystem appearing, where you have loads of algorithms both writing and identifying fake reviews.
And with that, poor old humans like myself can surely not compete. You'd always have to have a healthy scepticism towards reviews. I already do, in an era where most of the fakes are probably (?) still human-authored.
Its turtles all the way down.
The power of affiliate links to warp reviews is underestimated even on Wirecutter, imo. If they have to choose between a product that has no affiliate links and one that does, it's pretty much impossible for that not to eventually affect the recommendations.
This would make it more difficult/obvious if one person were to submit many reviews (use face recognition), raise the barrier for fake reviews, and give a lot more ‘signal’ to people trying to determine if a review is fake.
Of course, with Deep Fakes and such this could be bypassed still, but it could still have an impact.
You might want to but you are forced to provide a one star rating instead. :)
1) Give a person a gift card to purchase product
2) They purchase product and review (following this procedure)
3) Pay them
4) Repeat
Correct answer is to have verifiable sources, citations.
What we used to call "journalism".
If a restaurant site removed all user reviews and replaced them with a food critic's opinions I would trust it less, not more, fake reviews notwithstanding.
I love getting reviews from locals and non-experts, I'm just really tired of having reviews, comments, and other interactions online with bots or people pretending to be other people.
~ Khayri R.R. Woulfe
~ Khayri R.R. Woulfe
And you have poor diplomacy skills. You lashed out, called someone pointless, a jerk, and a troll. This was directed toward someone who criticized you in a neutral tone.
// Contacting the mods via email to take a look at this issue.
~ Khayri R.R. Woulfe
Just CTRL + F for "sign."
More to the point, I'm sorry you feel like I was flaming/trolling you, I wasn't trying to do that. My memory of the guidelines is technically out of date, but in principle it still doesn't really make sense to me to sign your comments even if the explicit rule has been removed. If you had just done it once I wouldn't have said anything, but I looked at your comment history and noticed that you're relatively new to the community and have signed almost all of your comments.
I just figured I'd politely ask you not to do it since it is pretty redundant - your username is effectively being written twice for every comment you write, you're just adding the full last name explicitly. Sorry you felt attacked.
The excuse is also pretty lame. Judging by your profile, your account is new (less than 50 days old) and the Guideline you're referring to is wayback to 2016. There seems to be an inconsistency there. Whether you have older account or not, there is still a clear intention od flaming here. Using a two-year old archive of the guideline is pointless to justify your behavior. New accounts will naturally folllow the latest version of the Guidelines so again it is pointless to refer to old version of the Guidelines.
It is imprudent that you never tried to re-read the Guidelines since 2016, and that would be impossible either because any update to the site, it turned out, is properly published as a news in the front page.
So, again, the elements of trolling, flaming, pointlessness and dishonesty has been sustained.
I perceive this incident as an instance of how old users game the HN system by trolling new users using provocative behavior, virtue signalling and downvoting comments.
But I guess the problem lies in how HN fringe perceives "civility" and "diplomacy" which is at the level of a crude AI that doesn't get past beyond mere keyword bypasses and bowdlerizing techniques. Humans thinking and acting like machines.
~ Khayri R.R. Woulfe
Given the sheer potential for abuse I would advise great caution with measures to remedy. This problem is ancient given that it has been around for literal centuries at least. Just in the article a contemporary of Oscar Wilde used the technique!
https://trends.google.com/trends/explore?date=2015-05-01%202...