Userdir URLs like https://example.org/~username/ are dangerous
blog.hboeck.de
blog.hboeck.de
The ship has now sailed, but I still think it's worth pointing out that it is Javascript's security model that is broken by design, that XSS vulnerabilities are the result, and the same origin policy is an incomplete workaround as illustrated by the article. This becomes clear when you consider that userdir URLs pre-date Javascript.
The public suffix list is yet another incomplete workaround for the security flaw in the same origin policy (that "same origin" isn't a concept that can be clearly defined).
The modern web stack is a security house of cards, as demonstrated.
How we ended up with billions of people able to interact with an app that we can all build at home seems insane...
X was the solution to the problem MIT had, and as such, the problem X was trying to solve. It worked very, very well for what it was trying to solve.
How we ended up with billions of people able to interact with an app that we can all build at home seems insane...
Java. Java did this. On billions of devices! Everywhere from your smart card to the mainframe. Before Java, Dis! Before Dis, Forth! Good luck running a web browser plus your application on a smart card, or even (decently) on an underpowered netbook.
Sufficiently young people probably suspect these were called netbooks as some joke. They run Emacs just fine and without constantly swapping, but a web browser?
If not, they can’t even run JavaScript ads, because apparently ads like to refresh the screen with a static image at 60fps. Software rasterizers can’t keep up with that (even on 32 core, >200W TDP xeons, from what I’ve seen...)
The web really sucks at everything other than network effects.
Hardware acceleration makes things slower, not faster.
> Software rasterizers can’t keep up with that (even on 32 core, >200W TDP xeons, from what I’ve seen...)
You're nuts. Software rasterizers keep up with DOS games just fine, even on an old 386.
It's the fact that they're trying to use Javascript and other garbage-collected or object-oriented languages to code these rasterizers.
At least JS had the excuse of being designed and implemented by a single developer in three weeks. X Windows was part of Project Athena, a $50 million project from MIT, DEC, and IBM with a five-year schedule.
Here's a nice little article from 1989 telling just how great it was for the time:
The layers of abstraction are there because X is such a godawful piece of shit. Sure, your application will be fast if you only use XCB or Xlib (and XCB is just a modern replacement for Xlib, made because of how deeply flawed Xlib is, and it can’t even fully replace Xlib). But few people are willing to put in the blood, sweat, and code to write something that is pure X.
So if X worked great for what it was aiming to do, it was aiming to do the wrong thing.
Sure, if you're using it to connect to a box in Canada over a satellite connection in Johannesburg during 1995 with countless abstractions, you're going to have a horrible time. Use the right tool for the job.
The X-Windows Disaster
https://medium.com/@donhopkins/the-x-windows-disaster-128d39...
>X: The First Fully Modular Software Disaster
https://news.ycombinator.com/item?id=17056516
John Steinhart wrote XTool, a nice snappy reimplementation of X11 on top of SunView! ;)
https://web.archive.org/web/20171008204348/https://minnie.tu...
X and NeWS History
https://news.ycombinator.com/item?id=15325226
Jon Steinhart wrote:
>XTool was very small and fast compared to the X sample server because I wrote the server from scratch. I think that I'm the only person to write an X server outside of the X Consortium. One of the things that I learned by doing it was that the X Consortium folks were wrong when they said that the documentation was the standard, not the sample server. There were significant differences between the two.
>The only really worthwhile thing about X was the distributed extension registration mechanism. All of the input, graphics and other crap should be moved to extension #1. That way, it won't be mandatory in conforming implementations once that stuff was obsolete. As you probably know, that's where we are today; nobody uses that stuff but it's like the corner of an Intel chip that implements the original instruction set. As an aside, I upset many when working on OpenDoc for Apple and saying the same thing there.
>The atom/property mechanism allows clients to allocate memory in the server that can never be freed. Some way to free memory needs to be added.
>The bit encodings should be part of a separate language binding, not part of the functional description.
>Had he done some real design work and looked at what others were doing he might have realized that at its core, X was a distributed database system in which operations on some of the databases have visual side-effects. I forget the exact number, but X includes around 20 different databases: atoms, properties, contexts, selections, keymaps, etc. each with their own set of API calls. As a result, the X API is wide and shallow like the Mac, and full of interesting race conditions to boot. The whole thing could have been done with less than a dozen API calls.
Recommendation: People who like saying "X-Windows" also enjoy saying "Project Anathema".
Those others I think get it wrong with their "whole desktop" sharing. The experience of having X clients streamed as they are launched and the local WM gets to manage decorations - I like that model a lot more, for say the scenario where I ssh into a machine and launch a GUI program.
RDP is much faster than X. It had the benefit of being designed later I guess.
VNC usually feels laggier than X to me somehow. As if there is a cursor sync issue.
These are just my observations as a user.
I, for one, would imagine something based on a DNS TXT records that either directly contain or point to a URL describing some kind of security domain policy: whether and which subdomains and subdirectories should be considered separate, isolated security domains, with pattern matching, of course. Or maybe put it under /.well-known/ instead of DNS. At the very least, this would solve the unregistration problem of PSL (i.e. when an entry becomes stale, you will not immediately know it, and old versions of your product will still treat it as present).
How would you do it?
I don't think it's that JavaScript and HTML are a bad choice, but there are some things that would have made life a lot easier if they were strongly enforced sooner, including secure cookies by default, SameSite=lax, removal of referer header and CSP - doing them sooner would have stopped bad developer practices while also removing a fair chunk of application security complexities, but at least we're moving towards a better world regarding those now.
I don't know if it'd be technically possible to implement, but additional characters to mark unsafe strings would have a huge impact on webapp security. Reflection of untrusted data at the moment generally relies on one of: HTML encoding, URL encoding or JavaScript escaping and escaping a safe way is highly context-dependent (I've seen an unescaped "\n" cause injection within JavaScript contexts). A way of effectively storing the level of trust a chunk of data has across multiple transports when marking untrusted data including within HTML/JS, SQL statements and interpreted languages like BASH or PHP - this would eliminate a bunch of vulnerabilities and would probably have mitigated a bunch of notable historic vulnerabilities and/or hacks.
I have had a hard time convincing co-workers that if you have php generating sql generating (! yes!) html generating javascript, you need to escape the string for javascript since it's embedded in javascript. Then you need the string escaped for html since it's embedded in html. Then you need the string escaped for sql since it's embedded in sql. Only then can you chuck it into the middle of the string. It is better to not do such craziness; but once you've decided to do such craziness, you must do it properly. The similarities between js and mysql escaping are irrelevant; it must be escaped properly each time it is embedded in another language.
The formats could be so simple: first the length of the data, then raw data of that length
“ The misspelling of referrer originated in the original proposal by computer scientist Phillip Hallam-Baker to incorporate the field into the HTTP specification.[4] The misspelling was set in stone by the time of its incorporation into the Request for Comments standards document RFC 1945; document co-author Roy Fielding has remarked that neither "referrer" nor the misspelling "referer" were recognized by the standard Unix spell checker of the period.[5] "Referer" has since become a widely used spelling in the industry when discussing HTTP referrers; usage of the misspelling is not universal, though, as the correct spelling "referrer" is used in some web specifications such as the Document Object Model.”
(Keep all the legacy crap for when this is not enough)
Say, you have some endpoints which can do a public key crypto signature verification. Part of the cookie data is such an endpoint URL. If a script loaded into a page wants to read a cookie, the endpoint defined by the cookie checks the signature of the script and the browser accordingly allows or not access to the cookie. As you'd sign the scripts during deployment, the private key is not even on the public servers.
For performance reasons probably wouldn't send the entire script just a hash of it so it's like GET https://example.com/check?hash=1234567890&signature=abcdef and there's a simple 200 or 403 response. You could make the API such that you can bundle multiple hashes and signatures together so the entire overhead over the network is a single request for every endpoint. If you want to be fancy, you can have a response body for the 403 response which tells the browser something like "this hash is of an outdated version, plz ignore cache and load a new one".
There's also external resource integrity checks which prevent modification of third party resources without breaking the local site. jQuery CDN code snippets do this by default: https://code.jquery.com/ .
You can't trust one script to access a cookie without trusting all scripts to access the same cookie though - while I can see some merit to the idea when it comes to hiding secrets from XSS/untrusted code, I'd say that in most (99.9%) situations effort would be better spent actually implementing CSP and good data sanitation rather than caring about implementing JavaScript level trust models.
Having the server send a 'this is my site base' header would be helpful even without having path support. At the moment, browsers need to magically know about all multi-level domains, like example.co.uk, so that they don't lump together all *.co.uk domains as the same site.
Unless I'm missing something, it seems that this would eliminate most of the issues raised in the article by using a separate subdomain for each user/site.
Using lighttpd, you can do this with this very simple configuration [0]:
$HTTP["host"] =~ "users\.example\.org" {
evhost.path-pattern = "/home/%4/public_html/"
}
Add a "wildcard" record for "*.users.example.org" pointing to the proper host to your DNS zone and you're all set.---
[0]: https://redmine.lighttpd.net/projects/lighttpd/wiki/Docs_Mod...
Cookies are a terrible security model in general.
If there is a cookie set for the path '/secret' and I can host content at '/attacker', then some of my JavaScript under /attacker could do a fetch request to /secret/something. This fetch request would carry the cookie for /secret, and the response would be readable by my JavaScript (due to Same Origin Policy). I could read the response, extract sensitive content, or even extract CSRF tokens to allow me to do state-changing CSRF-protected things under /secret
Origins are very clearly and rigorously defined[0].
It's just that people have a tendency to group multiple applications/users on the same origin. Now, would it be nice to optionally allow servers to be more restrictive? Sure. There are a few privacy implications there, but I could easily see someone making a case for that. I think I'd possibly support an HTML directive that allowed you to have a stricter same-origin policy (ie, restrict to current URL, or restrict to parent path).
At the same time, setting up subdomains that point to static directories is really stinking easy. And I am skeptical that any web host that refuses to use subdomains is going to care enough about security to opt-in to an even more aggressive scheme. A website that is hosting 3rd-party sites via userdirs is a website that will ignore or circumvent any security model you propose.
> The public suffix list is yet another incomplete workaround for the security flaw in the same origin policy
No, the opposite. The public suffix list is a workaround for the flawed security model of cookies. In your focus on Javascript, you are missing the much bigger problem. See below:
> there is however still a caveat with this: Unfortunately the same origin policy does not apply to all web technologies and particularly it does not apply to Cookies.
The same-origin policy is (overall) fine, the problem is that the old web security model didn't have it. So we are trying to retroactively build a modern security model based on strict, on-by-default isolation on top of a flawed security model that assumed every single domain would be owned by a single entity. Cookie "contexts" are a lot broader than modern definitions of "origins". They can be inherited from parent domains, they can be trivially shared across domains without user consent. And they also aren't by-default isolated across userdirs.
This has all turned out to be, broadly, a bad idea. While I don't think the same-origin policy is perfect, it's pretty obviously better than what we were using before. This is why you're seeing movement from the Chrome team to turn on SameSite isolation by default for cookies. Because we look back at the original web security model and realize that opt-in SameSite and randomized form-submit tokens for security were always a mistake.
All of this is dancing around that you mention that userdir URLs pre-date Javascript. To be clear, userdir URLs were inherently insecure from the moment they were conceived. You should not have been hosting websites this way, even before Javascript existed.
To the extent that people got away with this kind of thing before Javascript, it was only because userdir websites were so ridiculously limited that they often didn't have any serverside capabilities to respond to form submissions, or to run any kind of logic at all. To the extent that people weren't widely phished on userdir-hosted sites by fake forms with hidden fields that sent logic to parent origins or to other sites they were already logged into, it's only because the web was so young and so tiny that we were able to get away with bad security -- the same way that in a rural area I can get away with leaving my garage unlocked.
I'm not exactly happy with every decision modern browsers are making about the direction of web security, but I would challenge anyone to look at the history of CSRF mitigations and then say that the same-origin policy is not a universally better security model for Javascript to adopt.
[0]: https://developer.mozilla.org/en-US/docs/Web/Security/Same-o...
domain.suffix/full/path should have been the definition of an origin because people were literally using it as such. The default config of the most popular web server at the time was doing it. People hosted entirely different sites in subdirectories because it was really easy to just create a directory. TV ads directed people to www.blahblah/sitename.
The real definition of origins is clever since the path before the domain defines an origin path and everything after is owned by that origin. Neat and tidy for sure but broke userspace so to speak.
But that's never how userspace was ever designed to work during the web's history. Hosting sites this way, especially commercial ones, was always insecure because of the way cookies were designed from day one.
Same-site policies broke userspace in the same sense that Wayland is breaking X11 userspace by making it so applications can't just send keypresses to each other. Userspace was already broken, it was an insecure mess -- it's just that we recognize it now. If you were running a userdir site that supported PHP, you probably should not have been doing that.
I'm glad nothing got hacked for the people doing it, but I used to leave my bike unchained at the library growing up, and it didn't get stolen either -- didn't make it secure. I believe you that there was a time when companies advertised on TV using userdirs, because I know there was a time when Facebook only used SSL for its login screen and when Twitter forced users to enable SSM account recovery. Companies do insecure crap sometimes :)
If you don't care about same-site security, then nothing's broken -- the vulnerabilities you're exposed to now in a post-Javascript world for userdir sites are the same ones you were exposed to back then. It's just users today are more likely to actively exploit them.
Like doesn't it seem a little silly that you can't make a.mysite.com and mysite.com/a behave the same way?
The current system is mostly consistent with the way that SSL UX works, it's consistent with the way modern browsers display URLs. From an engineering perspective it's silly, but we're also thinking about ways to easily get across to users 'who owns the thing you're looking at'. It's also nice to have a rigid boundary somewhere, because it allows us to turn on isolation by default; the same-origin policy isn't something programmers have to remember to opt into. So there's UX questions like, "by default should cookies only resolve to a single URL, and programmers have to whitelist subpaths? Would any of the web operators doing this today care enough to actually implement an optional security feature?"
All that being said, I'm broadly sympathetic to an argument that same-origin policies didn't go far enough. I think you can make a very reasonable argument that origins are pretty arbitrary, and that the problems I list above are solvable, and that regardless of the original design we should adapt to what users want to do. That's a fine position to take.
I'm not sympathetic to the top-level argument I was originally responding to; that Javascript has fundamentally made things worse, or that the same-origin policy is a house of cards waiting to collapse, or that that browser policies like same-origin and SameSite aren't basic security improvements over what the web used to be before JS. The truth is, while imperfect, web browser security is still (broadly) very good. I question if there is any user-accessible native platform that comes even close to the web in having decent process/network/data isolation. There is no house of cards collapsing, in reality most native platforms like phones are largely playing catch-up to implement the same sandboxing features that the web has had for nearly a decade.
I get that's not the argument you're making, you're looking at this through the lens of, "should this be supported?". And again, that's reasonable. I do want to re-emphasize though; if you still want to host using userdirs, no browser capabilities have been removed, and nothing is any more insecure than it used to be. You're using the word 'broke', I'm not sure I get what you mean by that. There was an insecure thing people were doing, and they can still do it today with the same consequences. Nothing has fundamentally changed.
That sounds sensible! I always wondered why the origin and cookie scope isn’t limited to the full URL of the resource that set it and only explicitly extendable by something like <link rel=“same-origin”...>
Basically make every URL it’s own security and permissions sandbox but allow the programmer some latitude to extend it (and probably the administrator of parent directories/domains some capability to restrict how far up/across the tree a URL can reach, with the default perhaps being to allow only within same domain and path.)
is this a js problem?
To the extent that it's a browser problem at all, it's a problem with how browsers as a platform handle HTTP requests and site-isolation, which is a security system that predates Javascript.
Once JavaScript execution is obtained* it's possible to inject a JavaScript keylogger and/or rewrite the DOM to request authentication details from the victim (resulting in credential compromise). Alternatively, it's possible use AJAX to perform GET/POST requests to the same domain, routed through the victims browser which includes all cookies etc - effectively this is a time-boxed account compromise (CSRF controls do not apply when requests are executed from the local domain).
It's also possible to coerce a browser into triggering the exploit in a hidden iframe on a completely different page (eg you browse to evil.com, there's a hidden iframe which exploits an XSS vulnerability on facebook.com, compromising your facebook account if your currently logged into facebook on the exploited browser). I'm pretty sure samesite=strict only fix this if the XSS vector on facebook.com requires the user to be authenticated prior to exploitation, similarly, samesite=lax will not prevent attacks which require authenticated POST primitives.
*I'm a pentester, so that's sometimes my job, I don't break laws.
Perhaps, but I feel like major recent changes are fixing most of these vulnerabilities:
1. Content Security Policy locks down most of the vulnerabilities that result from something on a page being able to contact something from a different origin. Of course, it's not set by default and these days requires a lot of trial and error to figure out the "right" way to set it, but it solves a lot of issues.
2. On the cookie front, the recent changes to the default SameSite behavior also pretty much end CSRF bugs.
2. I'm assuming you're talking about Chrome's SameSite value; it's worth to note that this has been rolled back a short while ago because of compatibility issues in larger government organizations having to be accessible especially now with COVID-19. More info here: https://9to5google.com/2020/04/03/chrome-rolls-back-cookie/
2. Note what Chrome is rolling back is the SameSite default change. SameSite has existed for quite some time now, in all browsers, it's just that the default is currently 'None' in Chrome but is changing to 'Lax'. So you can still take advantage of this now, it's just Chrome is delaying changing the default so that it doesn't break sites who aren't prepared for the default change.
So my point is the tools currently available really tighten up the sandbox guarantees of the browser, and make it no more difficult than necessary to build a secure site.
HTML is based on XML and even if you remove JavaScript, you could still create XML tags and attributes that lead to malicious behaviour and information leaks.
Want to guess which vulnerability jumped from not being in the OWASP top 10 to being in the 4th place? XML External Entities (XXE).
This is like complaining that Geocities users shared a security domain. Surprise, there was nothing of value hosted there to start with.
Edit: The "true" marquee was running in the now defunced status bar (window.status) via JS.
After all, if we didn't need to care about user level security, why don't we let everyone be root? Surely we only created accounts for people we can trust, right?
I'm not saying that your point isn't valid that the exposure here isn't very small. I'm just saying that your logic isn't really very sound.
It really has nothing to do with homepaths, or even user-supplied data whatsoever.
Now yes, our uni probably shouldn't be giving out userdirs to all students, much like wifi, you should only let in the people you trust. But that's exactly the point of the article.
Now, if someone knows of a host selling ~username spaces to members of the public, that's a different story.
Except what the article is about; because all sites are on the same domain and origin all sites security affects all other sites security. A page on one site has the right to manipulate all pages on another site.
> Even the web server is centrally managed and the user only provides static HTML
In other news php has suddenly ceased to exist and javascript is no longer a thing. Sorry, handwaving the problem away does not work.
It was vulnerable to the attacks described here. Thankfully no-one on that domain ever exploited it in that way.
https://github.blog/2013-04-05-new-github-pages-domain-githu... and https://github.blog/2013-04-09-yummy-cookies-across-domains/ are relevant
It specifically mentions that and even names github.io as an example.
Also - why would a blog be threatened by XSS? A blog is a collection of static HTML files and I assume the creator would be using a composer or just a text editor + browser to add and edit pages. Why would anyone log into a blog? Nowadays anything like a comment section is generally handled by iframes to some other service (like Disqus) anyway.
Perhaps you've never heard of WordPress [0], the 17-year-old blogging software used by ~60 million web sites as of 2012 [1] and that is, today, "the platform of choice for over 35% of all sites across the web" [2]?
---
[1]: https://web.archive.org/web/20160129215921/http://www.forbes...
It is entirely possible to design your own application which uses unique userdirs without using any of Apache's built in functionality that are completely immune to XSS because it is not possible for any of the users to place malicious scripts on the server.
Likewise it is also entirely possible to build a web app that doesn't use userdirs at all in any form for anything and still be vulnerable to XSS because of underlying input sanitization or code injection problems.
So whatever you're going to do, just do your homework and do it well.
If the path field of a URL contains a tilde directory, do the following: 1) split the path after the tilde directory, 2) add the first part to the host field, 3) use the second part as the new path field.
If there are any concerns regarding compatibility, make it opt-in via a HTTP header field.
But the basic message is: don't expect any protection from cross-directory script attacks, so use domain names for anything serious.
> All of this is primarily an issue if people run non-trivial web applications that have accounts and logins. If the web pages are only used to host static content the issues become much less problematic, though it is still with some limitations possible that one user could show the webpage of another user in a manipulated way.
It is really without limitations possible that one user could show the webpage of another user in a manipulated way, but for a static site that is true whether it can be done dynamically because they share domain or because the phisher can simply read your site content, copy and manipulate it.
The site take the example of apache mod_sites, but alone it's not enough. You need apache mod_sites, and the ability inject a link served by it somewhere else.
Well, yes. If you let user input anything and then inject it anywhere in your website, it's a security whole. Not a surprised.
So if you let your users host entire application on your website, and let other users of your website visit those applications. Well, I mean, sure it's a problem.
The best sandbox is the one you don't need to use.
It was discussed here: https://news.ycombinator.com/item?id=15141594
I'm still of the opinion that there should be an RFC around tilde urls, making it official that they're for hosting static content that isn't vouched for by the owner of the main domain. This has been an important part of our culture for decades, and it should be formally recognized and codified so that people can do it in a secure way.
Having multiple unrelated unvetted outside users with the ability to make your server serve arbitrary JavaScript is dangerous. The route you take to get there is kind of beside the point.
If anything, mod_userdir is helpful in that it's obvious when content is under the control of a different user. Perhaps browsers should treat x.com/~x/ as a different server from x.com/~y/ for XSS purposes.
But in that case this has been true since XHR first came into existence 15-20 years ago: https://en.wikipedia.org/wiki/XMLHttpRequest (sub)domains are a fundamental requirement for web security, they essentially always have been.