Hashify.me - store entire website content in the URL
hashify.me
hashify.me
For example, what if a URL were posted to Hacker News, but after the URL was a ?hasifyme=THEHASH, where THEHASH was the Hash of the website linked-to.
This way, if the URL could not be loaded because the server load was to high, you could just forward the URL to Hashify.me and the cache of the plain text from the website would still then be readable.
Boom, instant cache of the website content stored right in the URL!!!!
I guess for this to work well. Dont you need Hashify to provider a service where you pass in hashcode and it returns the HTML to you.
You just wrap the call with good old JSONP and you are good to go.
Not so. hashify.me has no server-side component (beyond nginx serving three static files). _bit.ly_ kindly provides the hash table. ;)
Not having the time to do so, I leave the challenge here: http://bit.ly/ialoWI
The _right_ thing to do would be to build a shortening service designed to handle URLs of arbitrary length. I'm not sure that I'm willing to take on that responsibility, though.
data:text/html;charset=utf-8,However,%20data%20URI%20does%20the%20same%20without%20the%20server.
Here is a self-containing cached document, for simple texts probably with a better efficiency than base64.
EDIT: ah. The shorteners barf if you try to "shorten" the data URIs
With http://preview.tinyurl.com/3maue6t it works, btw.
Though, tinyurl still barfs when I try to shorten the data URI that http://software.hixie.ch/utilities/cgi/data/data gave to me when I let it grab the index of HN.
(This has always been an idea I was thinking about; nice to see it implemented)
That way you only get the text of the page, stored in a cache.
Take entire text of Bram Stoker's Dracula
Chunk into 123 parts
Data URI encode each part
Generate a TinyURL link for each data uri (thanks for having an API guys)
Embed the TinyURL links in the Hashify Markdown editor using object elements
Curses! Even just the objects takes it over the limit. It has to be done in two parts. Create two Hashify pages.
Part 1:
Part 2:
Works in Firefox and Safari. Chrome, Opera and IE9 don't like it.
hashify allows style tags :-)
That's _exactly_ what the site does. Everything happens client-side. nginx serves a single index.html for every request. ;)
(You're uploading the data in order to then convert it client-side. Groxx noticed this too: http://news.ycombinator.com/item?id=2464347)
Seriousy, though - awesome hack.
edit: oooh, another thought: you're essentially uploading the content of the page to view it.
The important thing is that the identified resource is unambiguously identified.
There might be legitimate uses for this, right now I can't think of one. Clever hack though.
Think if it like the difference between the postal service letting you know there is a package that you can go pick up at the post office, and the postal service giving you a package at your home or work that cannot be opened until you go to the store to buy a box cutter, but you have to bring the package with you.
The first example is cheap, since you only receive a pointer or link to where the package is, but you have to do all the work to get it. The second is not cheap, since if the package was a bed from Ikea (for a random large example), the postal service (bit.ly) has to deliver the package to you, and then you have to go somewhere (hashify.me) while carrying that package in order to see what's inside.
[1] pygm.us
Hashify is not really a hash, is it?
The difference is that AFAIK there's no algorithm to take a URL (plus or minus a username) and give you a bit.ly ID, short of looking it up at bit.ly.
Historically, the authors of Perl and Ruby (and WP tells me, Common Lisp?) decided to confuse the interface with the implementation, and use "hash" or "hash table" to refer to the mapping, and not ever Perl hacker has a Computing Science degree, so now we live in a world of people who think that "hash" means the thing that bit.ly does for you.
Once again, Larry Wall Ruins Everything. :P
I love this little hack. Sure, it may have no practical purpose; but it gave me great joy to see this. I'm still smiling.
However, this site disappoints me, it doesn't seem to do anything other than what a data URL can do, except it's vulnerable to downtime because of a centralized website.
Edit: for those of you unfamiliar with what a data URL is. You an store a HTML or image document using a URL like data:text/HTML;base64,hashifystuffhere
Just sayin' ;)
Or just, you know, email people the content if you can give them the same data anyway.
Nonetheless, it's tricky because you have no idea if proxies in the middle will be able to cope, mobile clients, and all sorts of things.. so you're right in the sense that it's pointless (if you want it to be universally acceptable ;-)).
I found that my web server started throwing 413s after only about ~30k characters.
Care to give a peak into how you came up with it?
Anyway, I felt great until 4pm on the Friday, when we presented our creations. Afterwards, I felt flat (as one often does after meeting a deadline or finishing a series of exams). I didn't want the excitement to end.
As I was walking home I had an idea. For some reason I wanted to share thoughts in 72pt Helvetica. I didn't want to broadcast them (I was melancholic after all), but I felt compelled to express them visually.
I began to think about how this might be done. The Web seemed like the obvious platform. I wondered whether it could be done without a database. I remembered something I had heard on a podcast about a site that allowed one play musical notes on a computer keyboard, and would encode these in the URL for easy playback.
This seemed a lot more interesting than sharing my moody thoughts, and now that I had something cool to work on I no longer felt the need to do so anyway.
I think I spent 40 hours working on it that first weekend (yes, I was consumed). I truly believed that I could ship it before showering and leaving for work on Monday! Doing so would have been a mistake – I'm pleased that I spent several weeks ironing out the kinks and integrating with bit.ly and Twitter.
Instead of...
from base64 import b64decode
b64decode(foo)
You can do... foo.decode('base64')
Encoding works too. As well as zip (foo.encode('zip')).Per RFC2045[1]:
(Soft Line Breaks) The Quoted-Printable encoding
REQUIRES that encoded lines be no more than 76
characters long. If longer lines are to be encoded
with the Quoted-Printable encoding, "soft" line breaks
Due to the insertion of these soft-line breaks, encoding is not the same, as you can verify yourself: import os
import base64
import unittest
class Base64Test(unittest.TestCase):
def test_long_string_base64_decoding_and_encoding(self):
byte_seq = os.urandom(500)
mime64_encoded = byte_seq.encode('base64')
self.assertNotEqual(base64.b64encode(byte_seq), mime64_encoded)
self.assertEqual(base64.b64decode(mime64_encoded),
mime64_encoded.decode('base64'))
if __name__ == '__main__':
unittest.main()
Decoding, as you can see above, is fine. This makes a difference when encoding a really long string in an HTTP header.[1] http://www.ietf.org/rfc/rfc2045.txt [2] http://docs.python.org/library/codecs.html#standard-encoding...
ERROR
The requested URL could not be retrieved
While trying to process the request:
GET http://hashify.me/IyBIYXNoaWZ5CgpIYXNoaWZ5IGRvZXMgbm90IHNvbH... HTTP/1.1 Host: hashify.me Proxy-Connection: keep-alive User-Agent: Mozilla/5.0 (X11; U; Linux i686; en-US) AppleWebKit/534.16 (KHTML, like Gecko) Chrome/10.0.648.205 Safari/534.16 Accept: application/xml,application/xhtml+xml,text/html;q=0.9,text/plain;q=0.8,image/png,/;q=0.5 Accept-Encoding: gzip,deflate,sdch Accept-Language: en-US,en;q=0.8 Accept-Charset: ISO-8859-1,utf-8;q=0.7,*;q=0.3
The following error was encountered:
Invalid Request Some aspect of the HTTP Request is invalid. Possible problems:
Missing or unknown request method Missing URL Missing HTTP Identifier (HTTP/1.0) Request is too large Content-Length missing for POST or PUT requests Illegal character in hostname; underscores are not allowed
Check that the address is spelled correctly, or try searching for the site.
On the one hand i am a bit disappointed (that i am too late), but on the other hand hashify.me is made far better I could make it. Great realisation.
Now that it _is_ out in the wild, I can't wait to see what people do with it slash build upon it.
Base64 is an encoding.
What would a trapdoor value be for SHA1 ?
What I can agree is that I would like the cryptographic hash to be a one-way function, yes. But not trapdoor functions, please :-)
(and can you point to the definition of the hash functions that you described ? I'm curious).
As for the parenthetical, I'm not sure I take your meaning properly. If you are asking where I learned that hash functions don't have to be one way it seems to be an odd question, but I just checked Wikipedia and it agrees with me, at least.
As for the trivial hash - now I scrolled down the page on Wikipedia, indeed. I never thought of it this way. (That an identity function on an integer would deserve to be called a hash function :-)
The part that got me was the initial sentence about hash function converting "large, possibly variable amount of data" into a "small datum". As "large" and "small" are implied to be of different sizes, I glazed over a possibility of identity function there.
Thanks!
When I registered the domain I imagined that state would be stored in a Twitter-style hashbang. Thanks to `history.pushState` and `history.replaceState`, this hack is not required in modern browsers. :)
I can't think of a single use case where you would go, "Ah ha! Hashify would work perfectly for this!"
This is clever, in that the entire content of the website is not stored in a database, but in external links. Obviously the biggest problem with this technique is having bots crawl your site, so Google's #! convention is used.
and using bit.libya. i dont trust it.
isn't this also somewhat censorship resistant. since the hashify url without its bitly can be put anywhere on the web that is writable, thus making multiple copies available in a covert way.
It's a very cool idea though.
> For longer documents, Hashify splits the contents into as many as 15 chunks. The chunks are then Base64-encoded and sent to bit.ly in a single request. The bit.ly hashes contained in the response are then "packed" into a URL such as http://hashify.me/unpack:gYi2Ie,g4fpte. Finally, this URL is itself shortened.
copy that to hashify, and let me know the length of the url you get.
Yes, 15 * 2048 = a 30k limit on document size
Actually, now I am wondering if an iframe src could be a data: url in browsers. If so, that could be interesting! Showing content without hitting the server. Probably not though, because of cross-domain security again. Any ideas?
>>> "<h1>Hello, World!</h1>".encode('base64').strip('\n')
'PGgxPkhlbGxvLCBXb3JsZCE8L2gxPg=='
Paste this into your location bar: data:text/html;base64,PGgxPkhlbGxvLCBXb3JsZCE8L2gxPg==Works in Chrome 10.0
Here's what I'm seeing in my browser console:
{ "data": [ ], "status_code": 403, "status_txt": "RATE_LIMIT_EXCEEDED" }
I should have included appropriate error handling for this! Everything except shortening continues to function, though.
One downside came when I tried bookmarking with Delicious (hit the url length limit, truncation would break it). But great for shorter content.
Apache responds with an HTTP code 414-Request URI Too Large once the URI reaches around 8K in length.
Default limits exist in several load balancers as well.
Hashify gets a pretty UI. pen.io removes the need for a DB.
Consider the user experience for the target site on a mobile platform. You have already loaded the site on your mobile device before even taking action, so when you click the link the response is much faster than requesting the site at the click.
As long as they don't have user accounts or database access or such, XSS doesn't let an attacker do anything meaningful. It's not weak security, it's just how the site works.
Edit: To point out the obvious, your iframe trickery is not necessary. Script tags are not escaped, nor are event attributes: http://hashify.me/PGRpdiBvbmNsaWNrPSJhbGVydCgnVGhpcyBpc25cJ3...
Agreed! One can always link someone to a static HTML page with <script>alert('fu')</script> in the body, but no one would tag that "XSS".
Does hashify.me make it easier to send annoying alert messages to your friends? Sure. Annoying, but no more harmful than sending them to the static equivalent.
Note there are risks to hosting arbitrary Javascript beyond stealing cookies. For example, you can steal browser history, discover NAT IP addresses, scan intranet ports, etc. Here's a presentation by Jeremiah Grossman covering some of these attacks: http://www.blackhat.com/presentations/bh-usa-06/BH-US-06-Gro...
Of course, attackers can host malicious content anywhere they control. I could just as easily send someone a bit.ly link to a malicious site I control.