NoSQL in SQL
hakkalabs.co
hakkalabs.co
Essentially I already have a giant precalculation service for all the needed calculations on all the data underneath a parent entity. So, I serialize that using ruby marshall and then use lz4 to compress it before storing to db.
Its actually faster and smaller than using json strings in ruby. The whole tree structure for each entity was 200-400k as raw json strings. It took something like 300ms to serialize to json. I was able to do the ruby marshal and the LZ4 HC compression in something like 20-40ms and it drops the size down to more like 15-30k.
JSON is a pretty cool format, but it's a lot slower in ruby than you might realize and it takes up a lot of space.
I didn't think about the size before compression.
Meanwhile, this serves well for Python: https://groups.google.com/forum/#!topic/google-appengine/WPf...
I shall go ahead and write some tests to see how much space will be taken up :)
I'm flattered it's getting so much attention (I like having my ideas spread!), although there's one thing that confuses me. I've never heard of Hakka Labs, nor did I post the talk, video, or my bio on their site (although their treatment sure looks like I did) -- as far as I can tell, they grabbed it from the site of the SFRails Meetup at which I gave it and posted it online. I'm grateful for the exposure, but some notification or clear notice that I'm not affiliated with Hakka Labs (whatever/whoever they are) would have been nice.
To me, http://www.hakkalabs.co/ looks a great deal like a blog. When I see a blog -- particularly with a name and photo of the author on an article -- I naturally assume that author either created that content specifically for that blog, or authorized/contributed that content specifically to that blog. If that's not the case, I think the blog needs to make it very clear that they are republishing content taken from elsewhere, without the author's knowledge. There's nothing wrong with that (assuming you have permission); it's just about making it clear that that is what's actually happening.
(Underneath, it's about the perception that I am somehow "contributing to" or "endorsing" Hakka Labs by posting content I created there. I'm not saying anything bad about Hakka Labs at all -- I simply don't know enough to judge either way! -- but IMHO it's not cool to create that perception without the author's knowledge. It'd be like a startup GitHub clone suddenly hosting my open-source code under an 'ageweke' account with my name and photo: while they absolutely have every right to do that according to the licenses involved, it gives the perception that I'm a user of their site and uploaded my code there...when, in fact, I've never heard of them before.)
Anyway, don't want to derail this technical discussion any further; feel free to reach out to me directly over email if you want to chat about anything else. You certainly have my email address. ;)
This practice been done by many production sites over the years http://backchannel.org/blog/friendfeed-schemaless-mysql Might not be considered bad practice now.
But every time a developer sees an interesting twist on a piece of technology and goes for it, peers call it a bad practice.
I've been through many cycles like this, and inevitably some time passes, and one day you wake up to see yesterday's bad practices have turned into exciting advancements.
Moral of the story is, ignore the wisdom of the day and go for it, tiger. Stuff that JSON in an SQL table.
https://github.com/perfectsense/dari
Here is the SQL schema: https://github.com/perfectsense/dari/blob/master/db/src/main...
We've used this model for almost five years now with great success. It's simplified rolling out "schema changes" since no tables need to be changed. It's also been optimized to a point where it's extremely fast.
Please?
Its not storage friendly but it is what I believe to be a valid use case.
My own experience is that there's actually more data than you'd expect that can fit into this model. On the other hand, I am absolutely not pushing this as a panacea: if you don't really know what you're doing, tossing JSON in a RDBMS is probably a really, really bad idea. After all, that's part of the talk -- to discuss when it's a good idea and when it isn't.
(I personally think the low-card tables part of the talk -- http://github.com/ageweke/low_card_tables -- is the most interesting idea.)