'Rater' is paid $10 hourly to teach Google's algorithm
finance.yahoo.com
finance.yahoo.com
Did you work from an office or remotely? Did you interact with other raters?
Was there an interview process? How about performance reviews? How did Google ensure quality?
Maybe that's why the search "X reddit" is so popular - it's got a bit of that social validation.
1. Do you click on the link in the results page?
2. After clicking on that link, do you go back to the search or not?
3. How long do you stay there before coming back?
Adding an explicit up/downvote would capture some additional information, but it would overlap with the above, and most users would ignore the button. I’d guess Google has at some point experimented with this and found it wasn’t worth keeping.
1. Clicking on a link is often a negative signal. If you're searching for "when was barack obama born?", hopefully you get the answer on the search results page itself, so you never have to click. Or, the caption for each search result should have a snippet containing the answer.
2. If you stay a long time on the search result you click on, is it because the information is hard to find or because the page is deep and useful?
I used to work on Search measurement at YouTube and Facebook, and run a human evaluation platform now. Here's an example what they can look like when you use high-quality raters: https://demos.surgehq.ai/google-vs-bing/
Sounds pretty ripe for abuse too.
civilized wrote (emphasis by me): "and Google enhances its ranking for me based on that".
I love stackoverflow content, but hate the surface, e.g. how it loads jquery from google (???) and if I disable it, then the content breaks, because obviously the only way to implement a toggleable div is when you send your users browsing history to google directly..
I would suggest that the big reason not to go down the avenue of social buttons is the inevitable gamification of it. Already we have have ratings farms for apps and amazon. Of course we will have similar if search implemented this.
Correct, SEO bots.
Additionally, people writing bots learn just like Google learns, and start generating what appears to be legitimate interactions. Same algorithms that fit to critique between 'bot' and 'non-bot' class can be used to create bots that fall into the 'non-bot' category.
It's a losing battle - but what we get is better than we would get if they collected less data.
Quality Raters don’t actually rank websites.
Quality Raters look at websites and answer questions about them. Google then uses that data as a small part of their algorithm to adjust some factors about those sites.
Quality Raters don’t actually rank pages on google.
> Quality Raters look at websites and answer questions about them.
Wrong. I was never asked any questions beyond to rank to site quality for a query
> Quality Raters don’t actually rank website
Wrong again. This is what we did. We assessed website rankings for specific queries.
> Quality Raters look at websites and answer questions about them.
Wrong. At least in my case. I was never asked any questions beyond to rank to site quality for a query
I heard it didn’t work very well for them. The results weren’t very good or helpful.
As someone else pointed out, this pays substantially more than AWS Mechanical Turk. To be honest, it's difficult for me to imagine an easier job than this one. I mean, most people who would do this job could easily now make more at McDonald's or Walmart or whatever. But I think jobs at McDonald's or Walmart are considerably harder, so I don't see why this job should pay the same.
Employee is generally used as the legal definition, but if you’re only job is to work for company X under their explicit micromanaged directions you fit the colloquial definition of employee even if your paycheck is signed by a different company.
This is some Kafka-esque shit. Unreal.
These "Google $title Experts" work as uncompensated customer support (a legal requirement that they treat as such, if the quality is any indicator) on various social media outlets like Reddit and Twitter. They don't just provide free labor, they also go to bat for Google's reputation, acting as marketing and PR liaisons to prevent backlash from users who've been failed by their official support processes.
One 'expert' I spoke with justified working for free by the "connections and access" they received. When I asked if Google would be flying them out to the events they have access to to meet with the people they have connections with, they said they were told they would have to cover their own travel and lodging. I suppose if Google doesn't pay them, they never have to worry about any Experts connecting and starting a union.
We're talking about a multi trillion dollar company. Absolutely ridiculous.
This is so true. When I try to find a fix for some broken behavior in a major Google app the most common place I end up is a support forum with zero solution, no comment from any Google employee, and a volunteer support person saying how the defect isn't actually a problem but that you should submit feedback within the app.
These volunteers can't be helping at this point. Google's disdain for users is pretty widely recognized by now and seeing a whole bunch of frustrated users being blown off yet again doesn't help.
First tier support is hosted primarily in India where they don't know who or how many are doing the work, and it's nearly impossible to get escalated as a customer to the correct support tier outside of their script. It's frustrating and mostly useless, but they can handle volume changes quickly at sub-human costs.
[1] https://en.wikipedia.org/wiki/AOL_Community_Leader_Program
I imagine there are Google executives in a room somewhere laughing their full heads off every time they think about how there are people who do free work for Google who think they are making the world a better place or something.
So when you go to a store that Google Maps says is open, and find it's closed, you refuse to correct the hours and go "MUHAHAHAHA...I refuse to help Google, take that, corporate scum!"
The thing is that business owners do update the info on Google Maps, and have an incentive to do so, but only volunteers update OSM.
OSM shines where businesses don't care but volunteers do: trails in a park, non-commercial POIs, etc.
I guess it's nice that ordinary people try to do these updates for Big G and Apple. But they're often wrong, or they presume a one-day change is a permanent change, or that it applies to every location on the planet.
Back when it was new and interesting, TripAdvisor used to send people hats if they reviewed a new or unusual hotel.
I wish I hadn't lost my hat before I had a chance to wear it while checking into a hotel.
“Your ratings will not directly affect how a particular webpage, website, or result appears in Google Search, nor will they cause specific webpages, websites, or results to move up or down on the search results page. Instead, your ratings will be used to measure how well search engine algorithms are performing for a broad range of searches.”
Is this accurate? Or is the guide bending the truth to discourage malicious raters from trying to influence site rankings? Google only uses the rater data to evaluate its search algorithm and doesn’t use the rater data to teach the algorithm? Seems like a missed opportunity.
https://static.googleusercontent.com/media/guidelines.raterh...
Obviously 2005 was a long time ago though and everything could have changed since then.
It was OK at first and then it started having me watch increasingly extreme porn. I think I made it about 3 months
[0] https://webcache.googleusercontent.com/search?q=cache:3HkKSZ...
There is nothing wrong with semi-supervised learning.
When you rate "accruacy" of information or "authorativeness" of a site, how do you know the site is "an authority"? are you guessing?
Sad story.
More like it matches some keywords like run-of-the-mill WordPress malware generates. I've seen plenty of both. Not to mention the similarity to the masses of biz/buzz/xyz spam websites.
It's not like you need a PhD to confirm if a search result sucks. Results are supposed to be good enough for the average user, so you need common sense and some on-job training to do a good job.
And having 10 common-sense guys churning through user feedback @$11/h is better than 2 geniuses @$55/h.
As one comment says, it probably beats fliping burgers.
They cannot rate all results anyway, so why not get the best possible evaluations?
Shouldn't quality data beat quantity if you extrapolate from the evaluated samples to the entire index?
Spammers target the lower part of the bell curve. If you let that part evaluate the quality of the results, do you think that they will recognize spam?
For example, here are some queries in my history - would the average person understand what I'm looking for and what a good search result would be?
"neovim worktrees"
"django models through"
"pandas econdb"
Standard Leapforce US English raters were paid $13.50/hr. If you were lucky enough to get promoted to "Preferred Agent" (PA), you made $17.40/hr. The criteria for PA promotion was basically to be in the top 5% or so on their quality metrics, plus various subjective quality requirements. Promotion rounds were irregular and random.
Leapforce's US English pay rates never changed over time. There were no raises, other than the aforementioned rare and seemingly luck-based promotion opportunity. I was one of the fortunate ones.
Leapforce raters were misclassified as 1099 contractors. There were no benefits. Training time was unpaid. Originally, raters could work as many hours as they wanted, subject to work availability. Around the end of 2013, they added a 40 hour/week limit because they were afraid a court or regulator might rule the "contractors" were employees and force them to pay retroactive overtime.
If a state dared to go after Leapforce or Google for misclassification, they would leave that state and fire the workers there. No severance, of course. There were numerous states they did not hire in, including their home state of California.
Task availability was usually very good for US English raters, but there were definitely periods (often around holidays) where that was not the case. There was never any communication about projected task availability. We certainly would have appreciated a quick heads-up of "hey, many of our engineers will be on vacation between mid-December and mid-January, so expect task availability to be light. Tasks will be added daily at 8AM PST, but probably won't last more than two or three hours."
I understand other locales were not as fortunate as we were for task availability.
Work availability for a particular rater was not necessarily as consistent. If you scored poorly on a monthly quality review or random quality audit, you could be abruptly suspended, fired, or limited to an hour or two a day until the next monthly review. There would be no prior warning of this. There were automated tripwires (which we called the "bot") that would flag and suspend you if your ratings were considered outliers, or if your task comments were too similar (even when there was a batch of tasks that were very similar), or if you worked too fast, or if you worked too slowly, or a million other reasons. If you were flagged by the bot, you had no work until the Leapforce quality people got around to reviewing you. There was no compensation for missed hours even if it turned out you hadn't done anything wrong.
Even more insultingly, Google loved to use quality reviews to informally implement guideline changes without actually changing the guidelines. If you weren't good at guessing their future direction, you could easily have your hours cut for the next month.
Around 2017, Google realized they (not just Leapforce or Lionbridge) could be sued for this misclassification nonsense, so they forced Leapforce and Lionbridge to convert their misclassified contractors to employees. Leapforce morphed into RaterLabs in what appeared to be some kind of liability-dodging scam. As employees, they cut us to 26 hours/week so they wouldn't have to pay for health insurance. There were still no benefits other than the minimum required by state law, if you were lucky enough to be in a state that required any. Nominal pay rates were cut in what seemed like a random and inconsistent (not even location based) way; depending on the cut, you got a few cents raise after accounting for the self-employment tax you no longer had to pay (which RaterLabs was quick to point out), or a pay cut.
There was a rater revolt around that time, but nothing seemed to come of it.
Appen bought Leapforce and RaterLabs later in 2017. Leapforce founders Daren Jackson and Caine Lai made off with a pile of cash. His employees weren't so fortunate. Now people are making $10/hr. for a job that paid $13.50/hr. in 2009, which would be the equivalent of $18.38/hr. in today's dollars.
The two people I knew who did Leapforce also did human powered answers for ChaCha(?) where they were paid per answer versus an hourly rate, and then migrated to Leapforce at a lower date. They were both in US and I would remember them complaining on AIM over task availability-- sometimes they'd have work, others they wouldn't. They would take breaks from playing Counterstrike to see if new tasks were available and disappear if there was work available. With Leapforce, were there different rigid time allotments per task? E.g. 30 seconds to classify an image vs 2 minutes for ranking a webpage to be shown as a search result? If you were flagged as a low performer, was there any recourse or remediation, or just getting disabled and hoping for a review?
AETs were usually reasonable until the Ruth Porat era. After Porat, they seemed to be getting pretty aggressive about reducing them as far as possible to the point of absurdity. We could release tasks for insufficient AET. I suspect they added automated tooling in the Porat era to "optimize" AET in their favor. With only a few exceptions I can remember, AETs only went down, not up.
If you were flagged by the bot, there was nothing you could do except wait for Leapforce to review you. That could take anywhere from a few hours to several days depending on how busy they were. I wouldn't say there was any meaningful appeal route; I mean, you could email their quality team, and maybe you could plead for leniency, but I don't know anyone had any real success with that. (I don't have personal experience with that; as I said, I was a PA/senior, which was a quality-based designation, so it follows that I tended to be pretty good at the job! But I always felt bad for other raters in chat who had to deal with that, seeing their livelihoods put at risk in such an arbitrary and robotic manner. If someone isn't a good fit for the job, then yes, it's not going to work out, but I do believe they should be given feedback and guidance on how to improve and a fair chance to do so.)
I assume they pay for all time worked under RaterLabs, due to employee wage and hour laws. Of course, they can still fire a rater for not meeting productivity requirements.
It doesn't work that way. Contrariwise.