Show HN: Web pages stored entirely in the URL
github.com
github.com
How are web pages shared? By distributing links.
How are pages censored? By deplatforming. This usually happens by activists targeting the host right in their cashflow.
If you remove the need to host, you remove a point of failure from publishing. Now the link is the content.
It's an interesting concept whether or not you agree with this implementation.
So instead of sharing the link couldn't you just... Share whatever the message or content is?
Edit: what I'm saying is that "activists" (which is an interesting concern btw) will pressure the platform that the link is being shared on because the link IS the content.
*Though the content is integrated well with modern browsers.
Why not just copy and paste the content directly?
Storing the page in the URL doesn’t remove the need to have a platform to communicate those links to those who want them.
No, the platform is wherever you distribute the URLs to other people. Browsers, for the most part, don't share URLs with eachother without an intermediary server hosting the shared links, either as one or more traditional webpages with links or something like a shared bookmark list.
Prediction: at some point someone will want to censor something and blame the host. The host or hostname will be taken offline, even though it's not storing or distributing the content.
Luckily, the benefit of open source is that anyone can clone the repository and host any number of similar services if they are serious about using them for content distribution.
I don't mind if there are problems with the host since it will mean people are actually making use of the project.
Edit: adding link to https://en.wikipedia.org/wiki/Data_URI_scheme
Telling a normal person to click that link (or, go to decoder.com and paste this text into there, for a more distributed version) is much easier than getting them to view the HTML.
I still think this is really interesting though.
Think about the tools a modern political campaign uses, it would be somewhat similar.
But aren't there any other encoding standards? Perhaps also with encryption?
I'm not knocking this url idea. Simply what-if'ing other possibilities.
I guess it's marginally better since there's no central "place" to take it down. However, the deplatforming concern is still there because you're still probably sharing those links on the major platforms. So they can still take those down. You can argue that having it distributed among random sites is better than having it centralized, but you can achieve the same thing by uploading to multiple sites.
The proper way to do this (that doesn't involve shifting the problem), is to use IPFS/tor. When you're sharing it, you can use a gateway server, but if that gets taken down or blocked, users can still access it using the underlying network if they really wanted.
Mostly, by distributing links on other web pages.
> How are pages censored? By deplatforming. This usually happens by activists targeting the host right in their cashflow.
This does not avoid that problem, since you still need (for effective reach) hosting for the links.
It also adds new problems, like the length limit most browsers impose on URLs limiting the size of content that can be shared by this method rather sharply.
Huh, I just saw a 2018 post on a Chrome support forum indicating that Chrome had the same 2083 char limit as IE...
Checking more deeply I've found more direct Chromium info indicating implying either no limit or a 2^31 character limit, with a 32kB display limit. So, yeah, the URL length limit may not be an issue.
https://www.dreamhost.com/blog/a-win-for-the-web/
One can also host in a different jurisdiction which doesn't have laws or action against your content. Preferrably that can stand up to demands of a major country. Switzerland is popular for this. Example host that I havent vetted but highlights common differentiators:
I originally conceived this as a simple, static CodePen clone, but I felt the "publishing" of pages as URLs was an interesting idea. So I decided to present that aspect of it front and center, even though it wasn't really the point of the project at the beginning.
About a year ago, I had a proof of concept version that I ended up using fairly frequently for sharing quick HTML/CSS/JavaScript experiments (never as a means of seriously publishing and sharing censorship-proof content). I found that if its use is limited to that case, it is actually very handy and robust!
Lots of interesting feedback in the comments, largely more positive than I expected. I was expecting criticism of code quality since I have minimal web development experience, but instead the (valid) criticism is coming from people who are taking the project's applications and use cases seriously, which I appreciate.
I suspect that's why this item is so popular.
I also hope the interest it has generated will inspire others with more knowledge and experience to make better versions or variations that are more practical.
Some hard problems. E.g. URLs provide a consistent place to go to find varying content. If the URL is the content, then it cant vary. How does one solve content discovery without resorting to hosting?
PS: The first person who says blockchain buys the beers.
Distributed search has been around for some time, e.g. eMule uses Kad[1], Tribler[2] has a solution based on Tor, but the most appropriate here is probably YaCy[3]. Of course, all of those are technically "hosted", insofar as p2p involves peers acting as servers. But I bet someone has made a f2f[4] version, meaning you'd only serve to a few, hopefully trusted peers.
I could see this issue becoming problematic with a browser implemented based on this idea, unless I am misunderstanding something.
Example: Alice makes a benign page that uses location data and shares the link with her friends. The friends know what's up and grant location permissions to get the page to work.
What they actually did, however was grant location permissions to "https://jstrieb.github.io/urlpages" and any script served by that origin.
So when some of the friends later open Eve's URL that contains a location harvester, they don't get any prompt at all: Eve's link can just reuse the location permission given to Alice because, as far as the browser is concerned, both scripts belong to the same page.
Or, any time this sort of thing comes up in a thread on a news website like this one, an opportunistic attacker could post a malicious payload in a link and watch as all the excited people blindly click away.
very hard for a normal person to tell the good link from a bad link, and removes the way most people determine if a link should be trusted (by looking at the domain)
The only thing I can think of is, that a service worker has it's own CSP, as opposed to obeying the CSP of the registering script, but this service doesn't use CSP anyways.
That said I couldn't get it to work. You would need to be able to register a service worker file at something like https://jstrieb.github.io/urlpages/sw.js but all the pages you have control over have proper html mime type and are rejected when you try to register them (and have too much actual html junk in them to run as js files anyway).
Just fetch it.
fetch("https://jstrieb.github.io/")
So yea, other sites on his github pages are compromised, that's true. self.addEventListener('fetch', event => {
event.respondWith(
new Response('<h1>I murder kittens for fun</h1>', {
headers: {
'Content-Type': 'text/html; charset=utf-8'
}
})
);
})https://github.com/jsfiddle/jsfiddle-issues/issues/1417
Only issue is this being deployed more widespread, Twitter may need to start scanning explicitly for this software to block it regardless of domain.
Since this has received more attention than I expected, I'd be willing to put the time in to implement this.
But all you did is shift the burden of hosting it to wherever you post the link, which can still take down your link.
But for some cases it can be useful. It makes censorship impossible as there are no server that host the data.
The, "there is no server" argument doesn't hold if you want someone else to ever get your message. It's just either your chat app or some website where the "link" is posted. I don't see any value added.
This allows about 80 unique characters - about 6.5 bits per character. There are dozens of thousands of visible Unicode characters, at least 8 bits each. Using modern compression techniques you can probably get around 4 or maybe even 5 characters of the real text into one visible Unicode character. This would compress the messages so that you can share them over SMS/Twitter (split into a few parts perhaps) or Facebook without posting giant pages of text.
The data could start with a 'normal' not-rare-Unicode character indicating the language, and the same character at the end of the stream to make it easy to copy paste:
a<unintelligible Unicode garbage>a
Multiple translation tables could exist for different languages. Japanese and Indic languages have ~50 alphabet characters but no capitals, so I think you would be able to get similar compression regardless of input language.Users would go to a website that does decompression with JS and just paste in the compressed text to view the plaintext. And vice versa. If the compression format is open then anyone can make and host a [de]compressor.
Version codes may be useful to add, so that more efficient compression techniques don't break backwards compatibility. The first release of the English compression algorithm would start and end with 'aa', a better algorithm released a few years later would use 'ab', etc. Similarly the version code could be used to indicate a more restricted set of characters to allow for better compression - [a-z0-9] and . , space newline is 30 characters which is only 5 bits per character before any further compression.
There are few other implementations of the same idea out there, too. Here's another one: https://github.com/qntm/base65536 ; since it uses a power-of-two sized alphabet, it avoids some of the work mine has to do around partial output bytes. That link also links to some other ideas & implementations
The exact metric I cared about at the time was compressing for screen-space, not for actual byte-for-byte space. I wrote it after spending some time in a situation where doing a proper scp was just a PITA, unfortunately, but copy/pasting into a terminal is almost always a thing that works. But I was also dealing with screen/tmux, which makes scrolling difficult, so I wanted the output to fit in a screen. From that, you can implement a very poor man's scp with tar -cz FILES | base64 ; the base-unicode bit replaces the base64.
Mine doesn't use Huffman coding as it wants to stream data / it leans on gzip for doing that.
This is how we crammed huge text games onto tiny floppy disks in the 80's.
http://inform-fiction.org/zmachine/standards/z1point1/sect03...
Python code to reproduce:
zw = '\u200b\u200c'
enc = lambda text: ''.join(zw[int(bit)] for byte in text.encode('utf-8') for bit in f'{byte:08b}')
dec = lambda code: b''.join(bytes([int(''.join(str(zw.index(bit)) for bit in code[i:i+8]), 2)]) for i in range(0, len(code), 8)).decode('utf-8')
>>> dec(' ')
'Hello world!'Regarding your last point, my plan is to have the main page (index.html on the site) always decode from pure base-64 but to eventually have other "publish" targets or endpoints (e.g. /b65536/index.html#asdf) that decode in different ways. I can then add corresponding "publish" buttons to the editor
data:text/html,%3Ch1%3EHello%2C%20World!%3C%2Fh1%3E
So if you do need to link to a data url you just need to a use an iframe.
Since I released it open-source, my expectation is that anyone who cares about safety from takedowns will clone and host one or more of their own. Or, for that matter, use the files offline with shared base-64 encoded URLs
data url's are truly entirely in the url with no hosting or dependencies.
Another such tool is [1] This gradient generator, which encodes the gradient parameters into the url so you can come back and use the tool to edit it!
Anyway good job on creating this!
[0] https://davidpartson024.github.io/no_host/encryptedPageMaker...
[1] https://www.reddit.com/r/programming/comments/47gjbv/experim...
[2] https://www.reddit.com/r/InternetIsBeautiful/comments/4o76gn...
The worst you could do is exploit a browser zero-day, but you can do that on any static hosting site already!
First, it's not hard to imagine that someone might try to get their account banned for a GitHub terms of service violation keeping in mind that GitHub holds the account owner accountable for content in their repository. This is true even if that content is from other account holders they've given access to their repository. In this case, anonymous access is intentionally being provided which could of course go very, very, very wrong.
"You agree that you will not under any circumstances upload, post, host, or transmit any content that:
is unlawful or promotes unlawful activities; is or contains sexually obscene content; is libelous, defamatory, or fraudulent; is discriminatory or abusive toward any individual or group; gratuitously depicts or glorifies violence, including violent images; contains or installs any active malware or exploits, or uses our platform for exploit delivery (such as part of a command and control system); or infringes on any proprietary right of any party, including patent, trademark, trade secret, copyright, right of publicity, or other rights."
https://help.github.com/en/articles/github-terms-of-service
Understanding what the tool does, GitHub might be forgiving on the ToS violation front. The problem is with the second scenario: law enforcement. It's very likely that in a lot of jurisdictions, law enforcement, prosecutors, etc., wouldn't initially understand what's going on here and even if it can be explained to their satisfaction, I think very few of us would like to spend a night (or more) in jail while attempting to explain.
>a very effective XSS host.
It can only do XSS against jstrieb.github.io which has nothing valuable. So it's not useful for anything. It can't be used in a <script> tag to obfuscate XSS attacks against other websites either, because the response isn't formatted as javascript. I guess it could be used in <iframes> on other websites in order to add obfuscation, but I think the use to attackers would be quite low.
Though I probably should have. Here is an example of a HackerNews login page served with jstrieb.github.com https://tinyurl.com/yypvh3by, you can login to news.ycombinator.com with it, but it easily could have been a phishing site.
My point is, this is a very good idea for offensive operations.
I guess you're right that it's useful for takedown resistance in phishing attacks. It's useless for small, sophisticated, targeted phishing attacks, but for large blunt untargeted phishing attacks it could be useful to have a site that would be difficult to take down and censor.
But I do consider phishing different than XSS.
data:text/html,%3Ch1%3EReinventing%20the%20Wheel%3C%2Fh1%3E%0A%3Cp%3EThis%20page%20is%20in%20the%20url.%3Cp%3E%0A%3Cp%3EHow%20strange.%3C%2Fp%3EEven if this were possible (which, even if tinyurl supported it, would break most browsers) you would be effectively hosting your content on tinyurl.com. Again, this is no different from hosting the file on any of a zillion public data hosting services.
Either they're blocking large-payload qr codes or by url length.
Edit: IE is 2083, Chrome is 1032
PS: IE6-9 don’t count as modern browsers, and haven’t for a long time. To my knowledge, they were the only browsers to limit to exactly 2083 characters.
Might as well just use DNS txt fields.
In theory, an alternative way to serve page content (instead via HTTP 200) should be to return a 3xx redirect and set the location header to a data URI.
This would indeed allow you to use URL shorteners or any kind of open redirect as a de-facto web host.
("In theory", in the sense that HTTP allow redirects to point to any kind of URL, so there is nothing saying it couldn't be a data uri)
I imagine this kind of situation is something that browser developers explicitly want to avoid, because the security implications would be a nightmare. So practically browsers do not follow redirects to data uris.
seriously, though: this has some marginal advantages compared to just sending the sites such as being supported as clickable in almost every program and being allowed in size restricted posts which allow arbitrarily long links
Now, if they were to implement some compression before URL encoding, it might then be feasible to break up large files into a series of links or something...
To use it in api mode, add 2fb.me/ before the url in full.
As you modify or create a language, I generate a deep link that contains all the grammar code and sample code for your language in the url:
Basically, it uses a data URI and Webtorrent for a similar concept.
Doesn't work in Safari (Mojave)
data:text/html;base64,... data:text/html,<p>Hello%20World</p>And furthermore, no internet connection is needed to access data: URIs.
Here is how the link looks like :
https://jstrieb.github.io/urlpages/#JTBBJTNDIURPQ1RZUEUlMjBo...
link to self:
http://url.shortener/some-vanity-url
and then get the vanity url and point it at your site's link.