Better password protections in Chrome
security.googleblog.com
security.googleblog.com
Here’s what happens. Your instance of Chrome hashes your username and password and encrypts that value with an ephemeral key which doesn’t leave Chrome, Kc
Google gets the encrypted cred-hash and applies a second round of encryption with a key only known to Google, Kg. They return the doubly-encrypted cred-hash, and also return 256 candidates encrypted with just Kg for you to compare against, those 256 candidates are selected based on a clear-text 3 byte prefix that Chrome also sends them.
When Chrome gets the results back, it decrypts the cred-hash using Kc which results in the cred-hash now being encrypted only with Kg; coincidentally this is the same exact key used to encrypt the candidates, allowing Chrome to do a simple byte-compare to see if there’s a match between your cred-hash and the encrypted breach candidates.
The cool part is the encryption operation is basically communicative. You can add your layer of encryption first, send the value to Google, they add their encryption and you get the double-encrypted blob back, and then you remove your encryption to end up with a plaintext encrypted just with Google’s key, even though Google never saw the plaintext, and you never saw their key.
This trick allows your browser, and only your browser, to compare your cred-hash against the breached cred-hashes while all the data you are comparing is actually encrypted with a secret key held by Google that never leaves Google’s server.
The only information leak is the 3-byte hash prefix to identify the slice of the comparison set. This is definitely not nothing, but perhaps it is a reasonable trade-off.
• I want to send my password (44) to Google but don't want them to know it's 44. So I multiply it by 37 (Kc) and send 1628.
• I also tell Google that the password is between 40 and 50. (Equivalent to sending three characters of cleartext.)
• Google multiplies 1628 by 78 (Kg) and sends me the result, 126984.
• Google also sends me every number between 40 and 50, multiplied by 78.
• I divide 126984 by Kc and get 3432. I see that 3432 is in the set of other numbers Google sent me (44x78). So I can infer that my password is in Google's set of breached passwords.
I believe the (first 3 bytes of the) hash of your username is used, as opposed to your plaintext password. I realize you were simplifying, but the difference between "plaintext password" and "hashed username" is rather essential in this case.
IIUC, their breach data consists of pairs of values - a username hash and the matching username-password hash (all encrypted ofc). This way they aren't storing a bunch of breached plain text credentials on their servers. This allows for a range search based on your username hash prefix, while the encrypted packages you send and receive are username-password hashes.
Edit: I believe I was incorrect - they only retain the _prefix_ of the username hash (as opposed to the entire thing). This has the benefit of effectively rendering the username unrecoverable - you'd be forced to attack the combined username-password hash. So it's not really a range based search, but instead a hash table. They just return the entire (24-bit addressed) bucket to you, which for 4 billion (!) sets of credentials amounts to ~240 on average.
You and I need to communicate via the National Postal Service. What you want to communicate is secret, so you don't want a postal worker or some random schmoe who finds your mail in the mailbox or a package on your stoop to be able to open it up and read what's inside, and you know from experience that any time you send something in the mail without securing it, it either gets read or (if it's not just text) stolen. We must live in a country with a really corrupt postal service, I guess -- kind of like the Internet.
We each have access to an arbitrary number of indestructible boxes, and an arbitrary number of indestructable key locks. Each box can have as many locks on it as you like. Unfortunately, each lock has only one key, and only the person who possesses the lock has the key. We can send things to each other in locked boxes, but of course the recipient doesn't have the key to the lock because the sender has it. Sending the key unsecured will get it stolen or, worse, copied. We also have no way to meet each other in person to exchange keys securely (and if we did, we could just skip the postal service altogether, anyway).
How can we arrange, with these resources, to communicate securely with each other?
[0] https://www.techrepublic.com/blog/it-security/the-key-exchan...
Alice attaches her lock.
Alice sends the locked package to Bob.
Bob attaches his lock.
Bob sends the twice locked package to Alice.
Alice removes her lock.
Alice sends the package, locked with only Bob's lock, to Bob.
Bob removes his lock and opens the package.
----
The scheme you discussed is really similar to this, but I guess with one key difference; in the scheme you describe the locks themselves contain information and are hidden by a second lock.
I don't think there is a good extension to the solution I posted above, but it's as if when Bob attaches his lock it is hidden until Alice removes her lock, and at that point Alice only cares about what Bob's lock looks like.
Your attack would work by removing Bob completely, you just play his part in the protocol.
This is true for any protocol though. If you don't know who you're talking to then you could be talking to anyone.
In practice most protocols get around this by either pre-sharing a private key, or using a public/private key pair with a challenge round to establish identities.
The key trick to what Google are doing is that Google never ends up knowing anything about what their users passwords are AND their users never end up knowing which passwords Google has discovered YET between them they arrange that the user knows if any of their passwords are known to Google. Magic!
With Pwned Passwords users can learn any arbitrary subset of the Pwned Passwords full password hashes they please for just one API call per prefix - with Password Checkup you have to start by guessing a username plus password combination, then do a bunch of fairly expensive operations, and you only get back a boolean. It's essentially never worth it for bad guys who are guessing.
This has two advantages, and doubtless both are attractive to Google although one advantage is better PR than the other
1. Bad Guys can't sieve the Google system for valuable data. For Pwned Passwords this isn't a big concern because Troy is using readily available password lists. Spending say $5M to get the raw passwords Troy used as input back by processing Troy's data makes no sense, just download a Torrent of it. But Google's Password Checkup doesn't just use public sources. Spending $5M (again just an example figure) to steal data from Google that may not be available to your criminal gang otherwise might be worth it for a big enough score. And it'd be bad PR for Google if a customer gets attacked by such crooks using data which those crooks couldn't have obtained at all without Google. Troy can always say "This data already existed, I didn't make the problem worse" but Google's non-public data can't make that claim.
2. Good guys can't either. Google has a valuable service here, locked into Google's infrastructure. They can choose to give it away, but they could also (as they have with some other products) later decide they'd prefer to monetize it. Either way, nobody else can duplicate it even though millions of people are using it, thanks to cryptography.
The costs are pretty enormous too. Pwned Passwords is mostly just a CDN, and clients are trivial. The upfront maths to make it work was hard, but that's a one-time thing. But Password Check incurs a considerable ongoing cost for Google to do hard maths AND each client needs lots of RAM (this is a memory hard problem on purpose, to exhaust bad guys who might try to exploit it)
Regarding costs, the one-time hash performed on intake is indeed computationally expensive. However, for queries the only cryptographic operation required is a single elliptic curve computation (using secp224r1) on the hashed record sent by the client. Computationally, that should be in the same ballpark as establishing a single TLS connection.
The primary expense (detailed in the paper) is bandwidth, on account of sending out an entire hash bucket with each response. The analysis in the paper assumes a 2 byte prefix, but they're using 3 bytes here to reduce bandwidth at the expense of privacy; for every query made, they have to send back ~250 values instead of just one.
This "But Google's Password Checkup doesn't just use public sources" is infering that Google's db is bigger and better. You don't know and this is example of magic lotion marketing.
"Our magic lotion for hair grow is muuuuch better, as it have secret formula".
Actually I suspect that Troy's db as for now is better/bigger.
Didn't you ever wonder when you see viable SQL injections and other attacks announced so frequently why Troy's data set is so small? Troy is taking the ethical high ground by not buying data. So his data is the tip of the iceberg.
0x0 - 0xF can be represented with 4 bits or half a byte.
Chrome has blocked autoplaying videos with sound for most sites except a small, hardcoded list that includes YouTube.
This means that if you create a YouTube competitor today, you are playing at a technical disadvantage. Or if you're just hosting videos on your own personal website!
What if Google's algorithms classify your new startup as "potential phishing", because users are re-using their own passwords on your site? How can you appeal? What recourse do you have against Big G's algorithm?
> What if Google's algorithms classify your new startup as "potential phishing", because users are re-using their own passwords on your site?
That's not how our phishing detection works. In fact, our internal studies show that a lot of users reuse their passwords often and while that's not the best password hygiene, it's the user's choice to make and we have to respect that and build protections with this in mind.
> How can you appeal?
Right from your search console.
> What recourse do you have against Big G's algorithm?
Ultimately, Google/Safe Browsing has a lot more to lose if their users stop trusting their product(s). I can tell you that we take false positives very seriously and try hard to provide a fair and speedy resolution.
How/Where does one start this process?
After the youtube banning debacle, can anybody really trust there are humans working in Google's tech support or that given the amount of traffic they receive that a ticket will be treated in acceptable time?
Of all the things I have heard from Google, this was the least expected.
No, it uses a Media Engagement Index and YouTube is on a list that includes a bunch of other video-focused websites like Netflix.
> The MEI is meant to allow media heavy websites (e.g. YouTube, Netflix) that rely on autoplay for their core experience
https://www.chromium.org/audio-video/autoplay
https://docs.google.com/document/d/1_278v_plodvgtXSgnEJ0yjZJ...
Your interactions influence the default list, minutely, but also only if it’s a popular site. The distortions remain.
Apologies if we have different definitions of hardcoding. I’m not suggesting this list is a compiler argument. Just saying that a Google list decides whether a domain can autopsy video, instead of a more open web approach like “HTTPS can do this”, which is used for many features like WebRTC.
If it was interaction only (so my site with 50 users that all interact with video make the list), that’s fine.
Do you have either of those things to show us?
https://blog.google/products/chrome/improving-autoplay-chrom...
"Chrome does this by learning your preferences. If you don’t have browsing history, Chrome allows autoplay for over 1,000 sites where we see that the highest percentage of visitors play media with sound. As you browse the web, that list changes as Chrome learns and enables autoplay on sites where you play media with sound during most of your visits, and disables it on sites where you don’t. This way, Chrome gives you a personalized, predictable browsing experience."
The inclusion of the "over 1,000 sites" list is what creates a distortion.
Let's take a very popular site indeed. Google.com. Does it get autoplay? Nope. Why not? Because people keep visiting without clicking on any videos.
If I create a new video site tomorrow, let's say I set up a new music label and the whole site is just music videos for our label's artists. That site will start with zero index value, and visitors will need intent to watch the videos. But visitors who come back dozens of times will be eligible for autoplay, because the MEI quickly elevates for them.
But if Google added that same feature to a new tab of the main search site it'd never get MEI high enough to autoplay because most visitors to this popular site of Google.com don't play videos. That's not what they expect this site to do so it doesn't get autoplay.
Most sites I visit which have high MEI don't even use autoplay, but a lot of sites I visit which have very low MEI do use it (and of course Chrome blocks it on those sites). Both causes are the same - the people giving me quality content are also polite enough not to try to shove it down my throat, and those shovelling everything they can down people's throats don't care about quality.
I like Firefox but it wouldn't necessarily exist if it wasn't part of the beast
Using Firefox makes a difference in many ways.
That said I wish Mozilla would allow other forms for income and just be honest about it (unlike Pocket which I think is a good idea destroyed by them being sneaky about it.)
Safari does pretty much exactly this, FWIW.
Even worse if you can use such targeting to figure out who is possibly vulnerable.
Or maybe that’s just paranoid, but I don’t trust an advertising firm (which is all google is at heart) to not do this.
https://haveibeenpwned.com/Passwords
https://www.troyhunt.com/ive-just-launched-pwned-passwords-v...
Mr. Hunt personally paid to operate these services for quite a few years. Then he got sponsorship from 1Password to defray the server costs.
I am happy for Google to use this information for aggregated security purposes (e.g. bots analyse suspicious new pages).
I am not happy with this being associated to the Google account. However, AFAIK, I have no way of knowing this, and in the absence of an explicit statement, I have to assume the worst: that Google will use this to target me ads.
If any Chrome team members are reading this comment, could we get some sort of confirmation that these security features will not be used to collect more data to target me ads?
Any statement they make is open for change at a later date if management changes their mind. You’d be better off thinking about moving to a browser made by a company that whose privacy protections don’t exist in a perpetual conflict of interest with the company’s dominant revenue source
Your concern is fair.
TL answer: I can tell you that this data is used only for improving user security. Legal answer: Please read the Chrome Privacy Notice :)
- Building a fully local system will not provide the same coverage and doesn't allow sharing findings across users (which means that when a new phishing page appears, it will appear as "new" for every single user, instead of being new for N users and then known bad for the rest of users). It also heavily reduces what kind of analysis can be done since you can't just store large datasets on every single device and/or run expensive algorithms on every single web page load on a mobile phone.
- Making it open source would not help. You can't know whether the code you can see is what is deployed remotely, so in the end you just end up trusting a different assertion instead (if you can't trust a privacy policy, why could you trust that the deployment is not backdoored?). It also has some significant cons: malware / phishing / abuse detection is in essence a cat-and-mouse game, and secrecy is unfortunately a key requirement in how everyone is building anti-abuse systems across the industry (not necessarily because they want to, but because nobody knows how it could work otherwise).
- You could even go all the way and have e.g. remote attestations, reproducible builds, etc. that allow proving that indeed the code running remotely is the open source code you want and can audit. This is barely doable with available technology these days, and even if someone was to do it there would maybe be 1K people on this planet able to understand why this is trustworthy. A prime example of this is looking at people in this very thread not understanding the differential privacy scheme for detecting compromised passwords.
Not trusting Google is a personal opinion, and I completely respect that. But implying that there is an alternative to trust for this kind of system is IMO misleading. Using DDG or Protonmail or any other service doesn't change the fact that you have to trust someone, it's just a different someone. You might personally believe their word more than Google's word, but if e.g. DDG started logging your identity and log requests and sell that to ad companies you would have very little way of learning about it either.
Disclaimer: I work for Google, not on Chrome, but I have worked on anti-abuse systems in the past.
There are at least two distinct outfits which pay grey hats to steal data from black hats who in turn obtained it typically through cheesy script kiddie attacks on web sites or phishing.
If you give them not very much money they'll "monitor" their stream of supposedly fresh stolen data (a lot of it isn't fresh because crooks are also often liars) for specific data items you tell them about. This is a bad deal, but apparently some pretty big US companies pay for that service. It's fine though because everybody reading this has unique strong passwords for every account right? Right?
If you give them a LOT of money (or if you were say one of my former employers and you owned the company that does this outright) you just get the raw data. As like UTF-16 XML or CSV files with different types of quoting on each line, or whatever other crazy and inadequately documented nonsense came to mind for each such type of file delivered.
Most of my current life is in unique-per-site and 2FA but the burden of remembering which one(s) are using shared was more than my motivation to fix. This mechanism may take me to fixing faster.
chrome://flags/#username-first-flow
is needed to see the Warn you if passwords are exposed in a data breach
option in Sync and Google services* random number.
It's good that Google is encrypting these hashes, but come on, really? Surely there has to be a better way than just shipping off unsalted password hashes to a centralized location.
Edit: Okay, unsalted scrypt for the password. Thanks for the clarification :)
In the linked paper where they discuss the tradeoffs and take blinding into consideration. They note that the protocol is that a client calls CreateRequest(u, p) which creates a Req that is then sent explicitly to Google. It appears to me that they consider the merits of sending a hash-prefixed password, but do not make it explicitly clear that the final solution sends a hash-prefixed password. I think that they would want to explicitly say that only a partial hash is sent to Google, if that were the case?
I'm not sure if you're actually talking about something else, but the paper says: "Post-canonicalization, the server calculates a computationally expensive hash of both the canonical username and credential password... This 2-byte prefix—while leaking some bits of password material—provides the client with k-anonymity over the universe of all username and password pairs."
IOW, the 3-byte hash prefix sent is of the username and password concatenated. (Note that Google seems to have added another byte to the prefix versus the paper).
They indeed appear to have increased the prefix from 2 to 3 bytes. This makes logistical sense though - with 4 billion items, a 2 byte address yields ~61k items per bucket (and thus sent to the client per request) while a 3 byte address yields only ~240 on average.
static constexpr size_t kHashKeyLength = 32;
static constexpr uint64_t kScryptCost = 1 << 12; // It must be a power of 2.
static constexpr uint64_t kScryptBlockSize = 8;
static constexpr uint64_t kScryptParallelization = 1;
static constexpr size_t kScryptMaxMemory = 1024 * 1024 * 32;