Gmail really wants me to say yes
abe-winter.github.io
abe-winter.github.io
The model is trained to offer reply suggestions that have the highest chance of being accepted, from a large whitelist of the most common short replies. The whitelist contains many negative options. We're optimizing for click through rate. That's it. There's no editorial judgement, definitely not "‘no’ struck someone as too negative and they had to take it out."
We actually experimented with intentionally inserting more negative options to increase diversity. Doing this reliably causes a hit to our metrics.
Discovering this made me pretty happy about the world. Most people are generally pretty friendly to one another (at least over email!).
To be clearer on our metrics: In addition to qual there's also a precision/recall tradeoff, so saying ctr alone was an oversimplification. We have coverage targets, but that doesn't really impact the positivity of the suggestions.
https://www.youtube.com/watch?v=lLcpcytUnWU
The key line being: "I'm sure you believe everything you're saying, but what I'm saying is, if you believed something different you wouldn't be sitting where you're sitting."
> We're optimizing for click through rate. That's it. There's no editorial judgement...
Something that all of us as technologists need to learn is that this IS an editorial judgement. We do not get to disclaim responsibility just because we delegated that responsibility to an algorithm. It is we who delegated it, we who chose the algorithm and the metrics, and we who are responsible.
"We're optimizing for click through rate" is how we got the proliferation of misinformation on Facebook. It's how we got https://twitter.com/chrislhayes/status/1037831503101579264. It's how we got Pizzagate.
"We're optimizing for click through rate" is simply not good enough in 2019.
We could claim to be naive in 2000, perhaps even 2010. But today? We all know we're playing with fire now, and it doesn't much matter that of course we didn't MEAN to burn the house down. What matters is that we don't give lit matches to children and we know how to not set the house on fire.
Someone who thinks "We're optimizing for click through rate" is morally neutral is not qualified to be making these decisions, just as someone who thinks giving lit matches to children is morally neutral is not qualified to be a fire marshal.
Would you agree?
I'm afraid to ask your opinion of SwiftKey's autocomplete.
If you think that giving lit matches to children is morally neutral, you probably either don't understand how dangerous fire is or don't understand how unpredictable children can be. That's the point of the analogy.
And let's not lose sight of the fact that this is a tool that offers an automatic one to three-word email response for when you're too lazy to actually reply. The stakes are about as low as they come.
I'm sure you can think of many products that have fewer users, and features in those products that are less central to them than e-mail composition is to a mail client.
In the case of autocomplete suggestions, they can still cause harm even when they're statistically likely. What do you think Google would autosuggest for "Blacks are...", if it were around in 1950? What would the statistics on the completion of that sentence look like? And would the effect of someone seeing that list be 100% neutral - or would it subtly nudge them?
There's nothing obviously harmful with the specifics of GMail's auto-suggestions now. But the principle of 'we offer the suggestions that you've shown us you're most likely to want to use' is not morally neutral.
I'm not arguing against your principles necessarily, but deploying the argument where it's really not warranted is a form of crying wolf and turns people against it.
And as an aside, at my university the only time I ever saw any recruiting material for STEM, it was along the lines of 'Scholarships for women in STEM'.
https://news.ycombinator.com/item?id=19978313 https://news.ycombinator.com/item?id=19978300 https://news.ycombinator.com/item?id=19978171
But the point is, whether this particular case is a problem or not, the general attitude that led to its development is a problem, and we're using this opportunity to discuss that (even though the original article is not very coherent).
I would suggest that if you are frightened by what is essentially an auto-completing on-screen keyboard being optimized for usefulness over making some random person on Hacker News feel it was more nuanced than just trying to be useful... there's probably no way to implement a feature like this, and make it useful, which wouldn't have you shaking in your boots.
Don't be frightened by what is essentially a riff on T9.
It seems okay in this case though.
For example, suppose you did a controlled study and found that the new autocomplete UI was presenting users with only "yes" options, X% of the time. Is there a threshold beyond which you would unlaunch or adjust the feature?
What if the new UI was actually changing user behaviour, causing them to answer "yes" more often than they otherwise would have? Is there a delta (positive or negative) beyond which you would unlaunch or adjust the feature? If you had to adjust it, how would you know you were adjusting it correctly?
What sort of telemetry would you need to put in the product so that you really would know if this was happening?
I'm curious if your team thought about these questions before launching, and what ideas came up.
Maybe? Maybe not? There’s visible and tangible negative impact from Youtube algorithms (owned, incidentally by the same conpany). And Youtube (or Google? Or Alphabet?) rarely if ever acknowledges and fixes those problems because their algorithms are also optimised for click-through.
I'd apply that instead to their language autocomplete feature(the grayed out text that appears when you start typing and that you can just press tab to insert into your email) which is scary in the way it tries to shepherd your language
If they just set things up to go by click rates and assumed that whatever results occurred must necessarily be preferable, then that would be disclaiming responsibility in the way you describe, but GP's version doesn't sound like that.
(But this isn't blind -- they're picking the method while watching the output, so it may be a fancy way of picking the exact content you want plus other similar content plus adversarial noise).
It's easy to say that something isn't good enough if you don't elaborate on what the axis good-bad is.
Not every recommender is a moral agent. Auto-suggest and conspiracy videos aren't the same thing.
And technically, yes, there's not such thing as a non-editorialized recommender, but this feels like a pointless nit since there's a world of difference between trying to be a reflection vs intentionally not reflecting based on a human-made moral judgement.
What even is the moral editorialization in auto-suggest? It's obvious we don't want conspiracy videos pushed front and center of educational topics, but how on earth do you balance the number of "yes" vs "no"? That's clearly going to be a completely made up number. If you don't have a clear problem to fight, it's probably best to not editorialize just to say you did.
It almost sounds like human nature gets in the way of revenue, so it has to be manipulated into something more business friendly.
We won’t consciously make that choice but it often seems to be the consequence.
It's a question of capitalism optimizing for local optima versus global optima. What were seeing with the internet is, what happens when we apply this process not just to food, but culture itself?
What dangers do you see in this specific example. People will be too lazy to type no, if they don't agree?
I could imagine that globally the most common replies are three forms of "yes" but for any individual user the most common replies are two forms of "yes" and one "no", simply because there's less diversity in one user's replies than in the entire population.
That this the only other choice available is also significant, particularly in reference to other cases where they have chosen to make turning something off impossible.
Taking away the extremely popular "classic layout" is an example. It wasn't costing them anything, but getting people used to accepting arbitrary changes forced on them has worked out very well for Apple, and Google has to want some of that. Each passive acquiescence makes the next easier, until it becomes automatic.
The post is mainly about services that push users into accepting things (like data sharing or microphone access), but the only example raised is your "suggested replies" feature, which has no apparent connection to what the post criticizes.
Email prefill replies are a silly and relatively harmless example to make this case, but the point is every last button on every product we use is caked in this and we're realizing too late that metrics identifying what humans want often does not coincide with what humans want being good.
People may not say no as often, but one no can be very powerful. It's harder to say no. Billions of people nudged towards saying yes is not nothing.
> Doing this reliably causes a hit to our metrics.
I take this as one more piece of evidence that thinking in terms of rates and metrics make you blind to the user experience.
Another possibility is that modern wage laborers are forced to wear a mask of positivity and say "yes, absolutely!" or "awesome!" to everything, because they are terrified of losing their jobs and not being able to keep their necks above water, in a world where fewer and fewer people own most of everything.
Perhaps you yourself are wearing this mask of positivity, when you come here and say that this makes you "pretty happy about the world".
The road to hell is paved with good intentions (and blindly following simplistic metrics).
It's healthy to say no sometimes.
RSA Animate made a short animation about how a culture of forced positivity in white color jobs, and how it contributed to the late-2000s financial meltdown.
"Smile or Die" https://www.youtube.com/watch?v=u5um8QWWRvo
I think Google in general needs to start taking seriously the idea that this is an editorial judgement.
There’s nothing necessarily wrong with optimising for click-through rate, but doing that will skew the options you offer to your viewers in specific ways. Those outcomes are still Google’s responsibility, even if they are the result of choices which are not under Google’s direct control, because Google has made the editorial choice to use click-through rate as their primary metric of choice.
(It’s relatively innocuous in this particular case, but when Google uses click-through to choose between content offered up by other people, Google is setting up feedback loops between consumers, content producers and the algorithm that mediates between the two which can trap the system as a whole in pernicious local attractors in the phase space of possible choices that may be impossible to escape.)
Imagine a light switch in your house that doesn't go ON an OFF but has a bunch of different options every day depending on what the model is trained to do or who the PM for this "feature" is.
It's ultimately manipulation, a loss of control, an unwinable fight against a learning machine that gets what it wants.
While it's exceedingly commonplace to fire off a one-line e-mail such as "Yes I can" or "Sure that works" it's very uncommon to respond with just "No I can't" or "No that doesn't work". More likely you'd respond "That doesn't work but how about ..." or "That doesn't work because ..."
For that matter, everything about the Gmail redesign is terrible. It has gimmicky crap like this I don’t want, nothing new I do want, is slow as molasses, and is extremely buggy and unreliable. It is constantly making me wait (sometimes minutes on a slow connection) to do actions which used to be fast, constantly misinterpreting my inputs and weirdly scrolling my page around, it takes unreasonable amounts of CPU/memory, and I have lost text I was typing several times. I frequently end up wanting to punch my computer when using the current Gmail, something which never used to happen.
Sometimes I get fed up enough to use the plain html version. But that is not really satisfactory either.
Bringing back the Gmail of circa 10 years ago would be a huge improvement.
Edit: apparently the suggested responses feature can be disabled in settings. That’s something at least.
Gmail's "basic HTML" version still works. I use it myself.
I know it's always the first measurement that springs to mind but I can assure you there are others.
This way you can get other desired behaviours without negatively affecting your metrics.
To put it another way:
The choice of metric is also a design decision.
- "Let's meet sometime!"
- "Good idea, we really should meet in person!"
and then nothing happens.
I have very bad opinion on this feature overall, but now i am curious; how did you measure people wanting to say No and not finding a quick reply? And then how did you measure the quality of the No available on the test set?
(I'd guess they looked at peoples' manually typed responses too though?)
I don’t believe you are contributing to the downfall of civilization.
>We're optimizing for click through rate.
and so was Hitler in ~1923. Optimizing something that affects how other people act (same happens on YT Veritasium:My Video Went Viral https://www.youtube.com/watch?v=fHsa9DqmId8) leads to bad outcomes (radicalization, race to the bottom, catering to lowest denominator etc).
If you are saying yes, then that's basically all you need to say. If you're saying no, you usually also provide a reason, so you probably wouldn't use any of the short, generic "no" responses Gmail would come up with.
It's ironic coming from person with @gmail.com email. Fastmail's service is $5/month, but they chose the free one which harvests their data instead.
If even the privacy advocates are not using their own advice, I doubt freemium is going to be over anytime soon.
Anecdotally, I have my private mail at Fastmail and my work mail at GMail and I've never seen an ad related to work emails (or private emails ofc, but that was to be expected)
I tried all of them, fastmail, protonmail, tutanota, Runbox and a few others, and none of them came close to using gmail. So I figured, what the hell, g-suite is around the same price and it gets me 30gb of cloud storage as well.
I wish google was a little more clear on where the privacy starts and ends though. Like if you want to use google home with a g-suite, you have to enable a bunch of tracking. I know this has to do with the fact that google wants to separate business from personal use, but really, I don’t want two e-mails.
- I can have multiple inboxes displayed in Gmail at the same time, such as is:unread or is:starred
- I can compose an email while reading another in the same window/tab
That being said, there are a lot of things I like about Fastmail over Gmail, notably that its way faster. As an admin (our business uses both Fastmail and Gsuite), I also like how much simpler and more standards compliant Fastmail is as well.
I mark and report them spam, and yet gmail won't even pick up on them if the same mail comes through again later from the same sender.
The mobile client, I really like the gmail iOS client. I think fastmail was better than the native iOS mail client, but I just really like the gmail one.
Google docs, photos and drive are nice additions, but it’s mainly the first two.
The mobile client: I don't know the iOS client but urgh. The Android FM client is pretty lame :D Luckily, I rarely use it.
I use next cloud for all docs, images, etc. and neither has an integration so that's a toss-up for me ;)
Now I prefer the mix, I can certainly understand people who prefer the separation.
The false-positives are usually so absurd, that I have to assume they have no boundaries for their ML-whatever. If it says spam because it has decided that contact books are only worth 90% positive signal, then into spam it goes. ML > any hard-coded rational decisions.
Fastmail also based out of Australia. There are dimensions to hypocrisy, not all of which are relevant from every perspective. Product choice is complexly multi-dimensional. A better response might involve asking why such a proclaimed privacy-conscious consumer remains with Gmail.
That said, if you prefer US government to Australian government, I am sure there are also privacy-preserving email providers in US as well. And there is also Swiss Protonmail, outside of both US and UK jurisdictionm, but it is more expensive.
This is kind of meta because I’ve turned this
autocomplete feature off, I’m sure of it.
Did I just do it on my phone? Did my wifi
blip so the AJAX didn’t work? I certainly
didn’t turn it on.
This strike home hard for me, as a pervasive problem. So many tech companies conveniently "forget" about user preferences all the time.For example, on my kobo e-reader, I'm positive I've disabled auto-update. And yet, one day few weeks ago it auto-updated and the new version stopped displaying side-loaded .epub files (from project Guttenberg). No rollback, no appeal. Seller's 2-year warranty has recently expired. Now essentially I have a modestly expensive semi-brick that will only let me read two titles purchased via kobo store, and nothing else
There are a lot of things about the computers we use that we simply don't understand unless someone tells us (because we weren't involved and don't have access to the code) and yet I guess many people want to pretend that they know what's going on?
As a non-native English speaker these sentences are very confusing to me. I'm not used to English in an informal setting, and they seem overly informal answers with implied meanings.
"Do you want to come bowling with us on Friday?"
"I'm down"
However, it's not always interchangeable with "yes" - for example:
"Did you get the server migration finished over the weekend?"
"I'm down"
... Would be completely wrong.
Hanlon's razor.
This is exactly how 'enhanced' location access worked on android for years. It would actually grey out "no" option if you told it to save the choice!
But on the subject, there was actually a publication on the effort it took to get this system to not just return minor variations of the same answer.
That's what made me extract myself from Gmail. I created my new email address and forwarded my Gmail account to it, and then for every email coming to the Gmail address I updated the sender with my new address. It took me about six months before all my important mail was using my new address, but it actually wasn't that hard and I feel hugely less vulnerable to Google's whims.
Edit: At the same time I also started using a proper password manager, making me again much less vulnerable to losing an email account because I no longer rely on resetting passwords to get into my multifarious online accounts.
This does not solve the fact that your email archive is in gmail though.
The program to do this is called offlineimap.
I have my bank statements in an email account by a provider that is unlikely to ban people for questionable and unrelated offenses.
(You can totally get banned from Google's entire platform if you do something wrong on the Play store after e.g. being pressured to do so by your employer.)
Most of my private email communication (which has dwindled to a tiny number of messages per month) is on GMail.
I've also got a couple other addresses in addition to that. I also have email addresses on my own domain, but I don't know if I'm going to keep that domain forever, as it costs a lot of money.
Devil's advocate: that possibility strengthens your position when your employer tries to pressure you to do something neither you nor Google find acceptable.
This is too optimistic. Google has got a grip on consumers and they will certainly not be driven away by dark patterned prompts. Anyway it's definitely not 50/50.
Is it though? Or has this just become the polite way of saying "you don't know what you're doing but I - abe the artistic - do?"
I have this pet theory that the West has achieved safety and peace of such a degree that people have to box shadows to get that little bit of thrill in their lives. No, there's no scary dystopia coming. You've just got to chill.
This is like me panicking that Jetbrains wants me to print out secrets to stdout because I typed out sout<tab> while writing something handling a secret. My god! They must be moving us to a scary dystopia where everyone's plaintext password is in logs somewhere where Jetbrains can steal it!
What a ludicrous post. I can't believe anyone is even taking it seriously.
The reason it's not coming may be because there are people actively watching out. It's hard to tell what would happen if we didn't have them. Society gets better because of people willing to tackle hard questions. They are often inconvenient in how they ask you to reconsider life choices.
> You've just got to chill.
Until you get overlooked by the trained heuristics because they see how exposure to your thoughts cause lower click-through on ads. Then you scream and nobody hears you. And the metrics are all green.