Google wants to get rid of URLs but doesn’t know what to use instead
arstechnica.com
arstechnica.com
- Do not know what or where the "address bar" is.
- To visit their website, they will "google" (the verb) their website address (or biz name) - which leads to newly launched website owners thinking their website isn't accessible/online.
Chrome and other browsers combining the address bar with the search box has only made this situation worse.
I mean, practically, if it doesn't turn up when searched for, it _isn't_ accessible to 99% of the world.
Also, many won't "google" their website address, but instead their business name. For a newly launched website, many national business directory listings (manta, angies list etc.) will outrank the official site for quite some time.
The web is everywhere. It's time to accept that web literacy is necessary for working in an office. This includes things like working with tabs, creating bookmarks and using in-page search.
What do you mean "the site isn't coming up"?
OK, I'll take a look right now... yep, it loads just fine.
Can't find it? Just type the domain name in to the address bar...
No, that's the search bar, I mean, you can use that, but it's only been a couple of days and [search engine] probably hasn't indexed it yet.
OK, if you want to look at it right now, then type the name of your website into the text field at the very top.
What do you mean it still isn't coming up? It loaded fine for me.
...It's not in the list of links... OK, that's still [search engine] results. How about this, type the name of your site, plus the dot com, and nothing else. Got it?
OK, now press enter.
Tada!
Also, this is a story of success. Sometimes it ends with "trust me, your site is up, just give it time".
Back in my tech support days, I found that it is surprising how often you get to the thing you need if you dictate things like Win+R "ncpa.cpl" this way instead of relying a non-techie user to find the next icon to look for.
A part might also be that the other side switches into "this is arcane magic, I have no clue and have to make sure I follow the directions properly" mode...
To boil it down to one example, how do you solve the problem of someone setting up secure.citigroup.accountmanagement.com and have an end-user understand that that's a phishing site? I mean, for goodness' sake, authority in the DNS chain is read backwards. How many users are going to read that and see ".com" as the root and technically most important element, rather than "secure"?
I'm not commenting on any possible solution since this article doesn't even sketch one. I'm just trying to apply the principle-of-charity to the article and come up with the most plausible interpretation, and the one most interesting to have a discussion about.
My own commentary is that I'm not sure if there is an answer that's any better than beating on domain names until they work as well as possible. The only other alternative I see is a centralized authority of some sort, and while that is likely to potentially work better than what we have now for maybe about 5 years, the negative consequences after that as the central authority learns to spread its wings and exert control to extract money from people and abuse its authority to push some agenda outweigh the safety.
URLs, as defined on the web using http and https, specify the use of DNS, which itself brings in a lot of semantics that are not optional. If you fundamentally change how domains are identified on the web, you're going to fundamentally change how HTTP & HTTPS urls are defined, either by rewriting the current definition or putting a new one on the side.
Since we're probably playing telephone here, at least two steps removed from the source, it's difficult to tell if they actually want to rewrite what "https://" means, or if perhaps they would create a new protocol ("ghttp://") which would define domains differently.
Here's a hyper-quick stab at an example of what that could mean. Imagine a world in which
ghttp:Bank of America: {
"server_namespace":"accounts",
"title":"Account Management",
"subhead":"Account Details"
}
is a legitimate "url" that you could use to link to things. The key here is not in the mere embedding of information, since of course you could translate the literal information contents into a URL no sweat, but that the semantics of ghttp would be redefined to A: change the lookup of "Bank of America" to use Google's home-grown domain-like validation system that is run like the highest-end TLS certificates B: contains a specification of how to render such URLs out in the user agent and C: probably contains all sorts of other exciting things that can go in that JSON blob because why not. Certainly the standard will not be a one-paragraph sketch on HN but run into the hundreds of pages, because that's how these things go.One of the benefits, for instance, might be sold as making it easier to parse the URL with a JSON library than bashing regexes together as we do now, which is incredibly error-prone, speaking as one who has made many errors myself.
Technically, that's still a URL, if you ignore the newlines I added for clarity. It turns out that if you read the standard, there's nothing about URLs containing DNS-based domain names; that's a specific detail of HTTP, HTTPS, FTP, and several other schemes, but not technically a characteristic of a URL in general.
So I would consider it a valid interpretation of the article (which, as I said, is rather vague) to say that what is meant by "redefine URL" is "redefine the set of URLs that user agents are used to using" and that Google may want to create a new protocol scheme.
It's also faintly possible they want to actually redefine "https://" itself, but, while there's technically room to do that in some sense, I don't think it will succeed.
Going past that has the tradeoff of requiring more typing which is not going to be well received.
amptp{
culture[en_US];
nice[36];
name[Bank of America];
host[Accounts]
node[/Management/Details];
data{
[view][statement];
[year][2017];
[month][Feb];
[page][2];
[section][debits];
[anchor][`[15`]];
}
}
This roughly translates ashttps: //accounts.bankofamerica.com/Management/Details?view=statement&year=2017&month=Feb&page=2§ion=debits#%5B15%5D
Compacted comparison:
amptp{culture[en_US];nice[36];name[Bank of America];host[Accounts]node[/Management/Details];data{[view][statement];[year][2017];[month][Feb];[page][2];[section][debits];[anchor][`[15`]];}}
GUId per domain. No hierarchy. Peer-to-peer sharing. Cryptographically signed user-comprehensible metadata that's resolved through some quorum protocol or stored in a blockchain.
URLs are fine, but the idea that identity of a website is embedded into its "location" is silly. Google owns gmail.com, but Steam/valve doesn't own steam.com. How am I supposed to know that as a user?
Really, the usecases for domain names names have completely changed since the time DNS was conceived.
That mostly solves technical problems, but technical problems are (probably) not what is driving this effort. Is a user supposed to know that d6998147-a7a2-485a-ac5f-74dddebfe070 is legimately Steam, but d69a2b76-96ca-4d7e-b25c-c5a69848e070 is not? Who verifies the metadata for accuracy? Who says you can't register Bank Of America? Those problems, and the other problems like them, are the hard problems, not cryptographically signing the answers to those hard problems. Crypto is easy by comparison.
1. Aliasing/abstraction of IP addresses. Being able to change physical location or server of a service without reconfiguring everything.
2. Mnemonics for getting to a particular server.
3. Indicator that a server belongs to a particular organization or company.
4. Grouping services together so they can share resources (like cookies).
DNS solves all of those problems, but not particularly well.
#1: You need to deal with a lot of BS when all you need is stable identifier.
#2: There is a lot of contention around domain names. Users prefer to use search instead of typing long domains. There is no particular reason why my mnemonics should be the same as for everyone else.
#3: Similar-looking domain names used for phishing. Users don't understand DNS hierarchy. Expiration. Hijacking. A tree structure isn't enough anymore. And so on. This one is particularly bad. Like you said, it's not just a technical problem.
#4: Just look at HTTP security headers and all the iframe issues.
The problem #3 can be solved separately from the rest. The rest are relatively simple technical problems if you consider them in isolation and we could solve them much better than DNS does.
Funny thing is that Firefox decided to hide the "http://" when Google decided it must be done, and now the "https://" used mostly everywhere is safe, because there is absolutely no excuse to hiding things that change.
But the real problem is that the web is currently dominated by entities that are interested on obfuscating it. Yes, domain names reading backwards is a really bad decision, but there is no reason to keep composing it by hiding the end of URLs, confusing URLs with search terms, and formatting everything the same.
Thanks, but could you please stop trying to control the web Google?
The arrogance coming off of this company is unreal at times.
Also, due to the way DNS works you really are trusting ICANN, then the com registrar, then facebook (CAs add an overlay, but then CAs like Let’s Encrypt ultimately trust … DNS). Making that explicit would probably be a good thing (and might even be a first step towards a rootless, decentralised future).
"People have a really hard time understanding URLs," says Adrienne Porter Felt, Chrome's engineering manager. "They’re hard to read, it’s hard to know which part of them is supposed to be trusted, and in general I don’t think URLs are working as a good way to convey site identity. So we want to move toward a place where web identity is understandable by everyone—they know who they’re talking to when they’re using a website and they can reason about whether they can trust them. But this will mean big changes in how and when Chrome displays URLs. We want to challenge how URLs should be displayed and question it as we’re figuring out the right way to convey identity."
And one of their incentives in doing so is that Google is an alternative to URLs. Already many users Google ‘facebook’ and click on the first link rather than just typing ‘facebook.com’ (or ‘face’ TAB ENTER).
Although I don’t believe I saw it in the article, it’s in Google’s interest to mediate all access to the Internet for consumers — this isn’t fundamentally different from Facebook’s VPN. It’s not in users’ interest, of course.
If they cared about security they could do something like display the domain & the path separately, maybe on different lines (with maybe colour used to distinguish HTTPs vice HTTP):
news.ycombinator.com
reply?id=…They've tried experiments but it's not as easy as it might seem at first thought. Consider what happens with someone setting up a hostname like bankofamerica.secureuserportal.com or www.gmail.personaluserinbox.biz – a fair number of people will look at the left part rather than reading from the right so even displaying it separately won't help, especially if they don't have a huge window with tons of room and all they see is the first part of "www.bankofamerica.comsecurecardholderservices.ru".
Only displaying the domain name will help with that problem but it has usability issues for any company which handles user data under subdomains since it'd be harder to tell user1.example.com from user2.example.com if there's any possibility of spoofing.
In fact, if I was Google (and I was less scrupulous), that's the kind of behavior I would encourage: train users to type destinations into Google, serve up the results using AMP, obfuscate the whole thing by eliminating URLs, then users will just stay in Google's ecosystem and never leave.
I will be surprised if URLs go away completely in the next 50+ years, but a new way to verify the identity of the website will probably be developed, if it's necessary.
They are trying to solve an issue which isn't even part of the responsibility of the URL. The domain is the truthteller here. If you can find a way to display to the user what actual domain he or she is on,you have solved the issue.
I don't understand why we should view computer illiteracy as something we must cater to. You don't have to know everything about how URL works, but being able to change ?page=1 to ?page=2 and knowing what is probably going to happen and similar tasks is knowledge probably needed today and should be teached.
Why does people never learn this? If you teach people to be uneducated this is exactly what they will become.
I strongly disagree. I do tend to think that they're useful and not obviously replaceable, but claiming they have no issues is dubious.
Spoofing, with or without the added level of sophisticated enabled by unicode. Special character handling (e.g. whitespace) is sometimes problematic. URL longevity is a real issue. Server-side security has been compromised in the past using "../" in URLs.
I'm sure I could come up with more, and I imagine there's a "top 10 fallacies programmers believe about URLs" list out there that would provide more examples.
punycode, similar-domain name squatting, SSRF/ generally parsing urls properly, parsing uniformly
probably 100 others but this was just off the top of my head. Some of these they intend to deal with, some they do not.
Unfortunately the easy solution to that seems to be more walled gardens. I'm pretty sure we'll start seeing Google-certified websites. Registering your website with the Google Search Console will start to become mandatory to appear in the top results. Or implement AMP.
So I'm with you but I think it's important to understand that URLs are broken for most people. When you see a perfectly legible piece of text, their brain is automatically blurring it like some piece of flesh in a Japanese Hentai.
This is a truism, but I'm not certain it's true. Is there actual data to back this up? What definition of "normal" is being used here.. the average person using the internet, or non-technically trained people?
The web has been around for decades now, and everyone including my nearly 70 year old mother uses it, and URLs have been pretty ubiquitous for a long time. Not knowing what a URL is seems about as likely as not knowing what a TV channel is, or that they have a bank account number.
It seems more likely to me that normal people do know what they are, but find avoiding them to be more convenient than using them.
Letting a commercial company define a breaking web standard change won't have any repercussions whatsoever.
https://news.ycombinator.com/item?id=17920720 https://www.polemicdigital.com/google-amp-go-to-hell/
There's a great comment about the link between these two over there:
What if there is an alternative system to verify that a website is owned by someone? For example, what if the icon needs to be registered with some kind of trademark-like entity, and the browsers made the icon more prominent?
Using it is more difficult.
Why would making that easy be a good thing, though? Seems like it would just lead to tld squatting and poor performance .
It seems very sensible to me to take the URL for what it is and communicate that back to the user:
- The protocol doesn't matter much to most people, except the implications it has, so maybe show "encrypted" or "not encrypted" (and fallback to protocol name for FTP and the like, most people won't ever touch that without knowing what it is)
- The domain is interesting: there's the TLD, which doesn't _really_ matter, and the second level name, which matters a lot. Subdomains matter less. How about we show it as a much more prominent "[ domain.com ]" with a less obvious subdomain.
- Paths are pretty generic all around: incrementally deeper descriptors separated by forward slashes. Could definitely show that as a breadcrumbs-style thing (might also make people care more about readable paths, which is nice).
- The query string should probably be shown as "tags" with a key and value, since that's what they are.
- The hash is a bit icky, but it might be enough to highlight that you're linking to a place in the page.
Could be displayed as something like this: https://i.imgur.com/RfJoP23.png
But at least Google doesn't have to tolerate something they find aesthetically displeasing.
It's called ShortestSearch and in essence could be looked at as a reverse Google in that you give Google search terms and it gives you a list of websites, but for ShortestSearch you give it a website and it gives you a list of search terms. All with the aim to perfect the art of "It should be the first result if you Google 'x y z' ".
Also found the part of the article where it describes its own URL strangely enjoyable.
This is one of the things that definitely can be improved on. Having something either in the URL or in page metadata to say "this is linkable". This doesn't warrant a replacement, however.
Maybe go back to lists and hierarchies of lists like the old days. With `UUID` behind it. Lists are at least navigatable.
So every URL would be part of at least one list and `URLs` point to lists. Or something like that.
Example: Sites-I-Use List can be shown in the address-bar instead of the url as green while a site I never used before can be shown as red. That’s not even taking organization into account, or all the other myriad of useful usecases.
Anyway how often have you used URLs list-like?! And how often have you used Google(which offers List by Search)?
Breaking/obfuscating/demoting the URL is so fundamental I am shocked it would even be suggested.
Which on one level makes sense. It's good to protect things that have made the web great. However, it's also good to maintain a spirit of innovation and willingness to question those things that people assume has to work a certain way just because it's always worked that way.
When there's a specific proposal, and if that proposal is a bad one, attack it on the merits. But there are real problems of security and user friendliness when it comes to URLs and I don't see anybody else working on it. Does anybody really think that it's impossible to improve and the way we do things now is exactly the way we should be doing it 100 years from now?
It's no different than visiting a shop or a house where you should be able to expect certain regulations to be met (insurance, building codes etc) and I can see more countries starting to enforce country-wide codes for web sites, annual checks, registrations, whatever for people to be allowed to trade in that country.
The route could be displayed less prominently to the right.
Preserve normal omnibar behavior for input.
(ref to Real Names, for those not old enough to remember)
we'll need a protocol. and a port. and a host. and then a resource on that host. and to combine them together.
how about:
foo.html\80:net.host\\:http
totally cool!
Sounds like Google is becoming a bully ;p. They shouldn't reset standards of World Wide Web.
I'd imagine you'd still keep a UI for extracting a sharing the locator unless whatever friendly identifier was adopted also came with a reliable and widely supported way of recovering the corresponding URL.