Some problems of URLs
noncombatant.org
noncombatant.org
First, we can remove parts of the URL we don’t need or which exacerbate our problems.
I'll be the first to comment and say this: No, no, NO!!!
Hiding information from users is never a good thing! All it does is make then even less aware or likely to learn, and that only perpetuates the vicious cycle of ignorance and vulnerability to being tricked.
Incidentally, I was recently forced to use Safari briefly, and absolutely hated the opaqueness --- imagine many tabs all with the same only-hostname in the address bar --- the constant feeling of "where am I, really?" made it very unpleasant and disorienting.
There is plenty of good discussion on this topic here several years ago: https://news.ycombinator.com/item?id=7677898
Users should be educated, not treated like they are stupid.
When false, I see http://example.com/
What is its server headers set to?
What HTTP status code did it report?
What PID is running the current tab?
I agree that we should not hide the full URL from the user, but software hides information all the time. That is the only way computers are remotly usable.
Counterpoint: hiding information from users is often a good thing. Especially when that means reducing the amount of information a user has to process to determine if a URL is legit.
The URL path, query and fragment are the private business of whatever authority controls the URL. If an end user has to consult the path, query or fragment to make a decision about whether the URL is okay, the website has got it wrong.
> imagine many tabs all with the same only-hostname in the address bar --- the constant feeling of "where am I, really?" made it very unpleasant and disorienting
If I tried to navigate tabs by full URL I'd already have that feeling. I can tell from the current URL that I'm replying to a post - but I can tell that from looking at the page itself, and the page title - but 15659459 doesn't tell me very much about what post I'm replying to. Certainly, if I have multiple comments open I don't remember which is why by their id. The same applies on a great many websites.
As far as I remember, the only websites where I do navigate by "hacking" the URL are ones I've made, or it's done for the purpose of deceiving people about what the URL leads to: e.g. changing the product name in an Amazon URL.
For keeping track of where you are, rather than displaying the full URL, I think it'd be better to move breadcrumbs into the browser chrome - https://i.imgur.com/d821tXu.png instead of https://i.imgur.com/UufKDoS.png
(Relatedly, I've no idea which parts of an Amazon URL are "safe" to share. A UI fix to the URL problem should also include the ability for Amazon to tell the browser how to make a safely shareable link.)
The second half clarified most of this but I think the article could do with some revision.
Under any URL-ish scheme, including any simplifications suggested here, nontechnical users will be unlikely to do anything except clicking and copy/pasting. So it really doesn't matter at all to them what the actual text string's structure is. It seems like maybe the article is trying to suggest "advanced" users would benefit from a better system.
Notice that in the screenshots posted, each browser varies somewhat in how it displays the current location. Regardless of the syntax used in a URL-ish string, the browsers can certainly improve the display of this important information.
What about something like this:
https://i.imgur.com/PTs55Tl.png
Where you only show "nonstandard" portions if they are present, and they can be given special coloring to help alert the user.
When the user interacts with that portion of the UI, they can paste in a URL like normal, just when it is displayed, the parsed version can be shown instead to help them see clearly what they are looking at.
1) How people think URLs work, the good parts that people see every day
2) Abstract problems with URLs, followed by concrete examples, maybe links to prominent related bugs
3) Suggested fixes for the problems, including what browers are doing at a UX level
4) Honest look at tradeoffs involved for end users
5) Remaining complications and unanswered questionsThe URL is not https://www.tesla.com it's https://www.tesla.com/.
Instead of treating users like they are stupid, educate them. Show them the entire URL and teach them how the WWW works. People are going to encounter URLs anyway (when pasted), so trying to hide their structure is just going to keep making things worse.
Or go to http://example.com#123 or http://example.com?123 and see where you end up.
It makes things more confusing, by hiding what is really going on.
It leads to things like otherwise technically-savvy people leaving off trailing slashes on longer pathnames, which can cause other problems. For example, http://example.com/some-path can be much different than http://example.com/some-path/
Some browsers were also hiding trailing slashes on longer pathnames, which seems to have stopped (possibly due to breaking sites).
By hiding the real URLs, even programmers get confused, not just the end users.
GET / HTTP/1.1
Host: www.tesla.com
You cannot leave the path after the method empty in HTTP.A better way to look at it is how w3.org defines it: > h t t p : / / hostport [ / path ] [ ? search ] > path : Void | segment [ / path ]
Notice how the slash is in the optional part.
I'd be curious to see if it's possible to serve one at http://example.com?page=home, http://example.com#home, or something like that though.
In any case, http://www.tesla.com is not the URL of their home page -- it's http://www.tesla.com/ even though the browser tries to tell you that you're visiting http://www.tesla.com.
1. how to decouple content from location, i.e. location-addressing vs. content-addressing 2. how to accept more network connection constructions than what URLs permit (proxying, domain fronting, tunnels, etc.)
> We could also imagine a new URL syntax, with fewer and less ambiguous syntactic meta-characters.
That's part of what we're attempting for network addressing with the multiaddr format: https://github.com/multiformats/multiaddr - If the author of the article is reading here, I'd love to hear your thoughts!
We're going with encapsulation and a posix filesystem compatible plan9-ish syntax and approach there. multiaddr describes the semantics of networking namespaces, and then various content systems bring their own namespaces. IPFS brings the /ipfs and /ipns namespaces for its content.
[0] https://www.blackhat.com/docs/us-17/thursday/us-17-Tsai-A-Ne...
Care to elaborate?
https://www.youtube.com/watch?v=D1S-G8rJrEk
He also has a blog post about it:
http://blog.orange.tw/2017/07/how-i-chained-4-vulnerabilitie...
The premise is that URL parsing is complex and libraries get it wrong. This problem is pervasive and leads to server side request forgery vulnerabilities, which Orange was able to escalate to remote code execution on Github.
It does seem a bit silly to have two orthogonal directory structures smooshed together in a single format, especially with different nesting directions. In a way, URIs are a sort of serialization of the application layer packet headers.
I'm curious what a redesigned "address" could look like when optimized for user-friendliness.
Not just browsers but a lot of other things that want a hostname in either DNS or IP format will accept the same thing, because they all use this function to parse it:
http://pubs.opengroup.org/onlinepubs/009695399/functions/ine...
It'd be interesting to see research on how many laypersons understand which way around the "chain of authority" goes for domain names. I can easily imagine somebody not Internet savvy thinking that "facebook.hackable.org" was a legitimate Facebook domain. (And it doesn't help that some organizations spread their site over many domains that - even to a skilled user - are indistinguishable from phishing domains.)
It is unclear how a domain like that could be optimized to ensure a novice user understands that what they're looking at isn't Facebook, even though the domain starts with "facebook" and the page they're looking at looks exactly like Facebook's login page.
One general, slightly off-topic, notion: there should be a protocol whereby a password manager can ask facebook.com what sites are legitimately going to ask for your Facebook credentials - not entirely unlike SPF for passwords. Then even if you search for Facebook in your password manager, it should refuse to automatically provide the credentials to the site.
(More off-topic aside: Password managers should be properly integrated into all browsers. Knowing your passwords should be considered unusual, and actually choosing them yourself downright stupid.).
23.185.0.4 (www.usenix.org) can be abbreviated as 23.185.4
192.0.79.32 (hackaday.com) cannot be abbreviated as 192.79.32