Storage is Cheap, Don't Mutate - Copy on Write
hashbo.posterous.com
hashbo.posterous.com
For instance, if you're mapping disk space to memory addresses and have large blocks of data, this would mean that every time you needed to update a single byte, you'd either have to duplicate the entire block (possibly several mega- or giga-bytes) or give up linear access. Yuck. There are a lot of use cases that would absolutely be miserable for.
You can get away with such things a lot further if you're implementing these things at the file-system level.
Also related (HN discussion): http://news.ycombinator.com/item?id=1138979
And Cris Okasaki's blog: http://okasaki.blogspot.com/
Yep agreed, I did I believe mention it is most valuable when you have mostly have creates & reads and less updates like in the example of web applications. Of course there are valid in memory level operations, quite famously in the Java world we have ConcurrentHashMap which is a CoW semantic. However naturally one size doesn’t fit all, but then it never really does :-)
Remove the mutation part and now you have a version history,
allowing conflict resolution.
is a little too "hand-wavy" for my tastes. Conflict resolution is not an easier problem than resource contention.My point is more the negative here, you can’t do conflict resolution if you write over the previous state.
Sorry if it was hand wavy, was just a quick Friday afternoon blog :-)
Adding version history to it doesn't change this fact.
Example:
You're iterating over bank records, marking them paid. Another thread adds one to the beginning of the list. Thread 1 finishes marking them paid, then moves them all to the 'paid' section... Including the non-paid one.
The second thread makes a copy of the list and adds a new item.
The next time the first thread goes through to mark everything 'paid', it either gets an empty list, or a list that still contains everything it supposedly moved elsewhere earlier, depending on how you handle it.
paid = (markpaid(x) for x in list_records())The only information available before asking for signup is this;
> Hashbo is a simple innovative, new way to communicate, collaborate and curate.
What the hell is that supposed to mean? Is it a blogging system? Microblogging? A social network? IM? Is it some combination, or a new thing?
People come to a site like this with the question "What is this and why should I care?". At the end of that (quite nice) animation, I have no idea, and that doesn't make me intrigued and curious, it makes me annoyed and frustrated at not getting a straight answer.
Nobody cares about your buzzwords, they care about what your site is, and what it can enable them to do that they couldn't do before.
It’s not attempting to give intrigue, I just really haven’t figured the right copy out yet:) Take a look at some more of my blog articles if you like, they explain at least the general idea of what we’re doing. We hope to get everyone on the beta soon.
The problem for me is that it’s actually difficult to describe what it does and is better experienced then you can make your own valuation as to whether it suits you or not. Much like when the Internet started, it was easier to show than explain. So while we're just in the early stages of beta I didn’t want to prejudice people by giving partial information.
So I’ll see if I can get some feedback from early beta testers to find out what words they use to describe it - that may help me get a more unbiased view.
Glad you liked the animation btw., we’re (mostly) a one man band in reality, and I’m a techy at that, so any complement towards design is gratefully accepted.
All the best and thanks for the feedback, it’s really helpful to know what people think, Neil.
1. I don't want to click another link just to find out the answer to the simplest question I have - 'What is Hashbo?'. You have to remember that people don't care, and won't care unless you give them a reason to care. Every additional click, every extra second, loses people.
2. (minor point) The link opens in a popup, which is blocked by my popup blocker. It really only makes sense to open in a new tab/window if the user is very likely to want to keep the first page open, and there's basically nothing else on that first page, so no-one will want to keep it open.
3. When I get to the blog, there's still no insight into what Hashbo is, just the first blog post, which at time of writing is a post about NoSQL. What percentage of prospective users of your site have ever even heard of SQL? (Non-rhetorical question, I really don't know because I still don't know what the site actually is)
Because of this conversation, I've already spent far, far more time on Hashbo than I normally would give to a random new site, and I still don't know what it is, and I still don't care. And I don't think I'm particularly unusual in my attitudes.
I'm not a 'hater', really I'm not, I'm just pointing out that you have a pretty serious problem here. What's your elevator pitch?
If you product involves education (new model, etc.), you should be pitching to everyone, counting the mean-time-to-understanding, and relentlessly honing it downwards.
http://en.wikipedia.org/wiki/Multiversion_concurrency_contro...
IE the basic point I am making here is that MVCC does a lot of relatively complicated stuff to handle this for you. I don't disagree that you are gaining a lot of power at your remote nodes to do things that would increase performance, but for 99% of applications I think it a huge amount of work for a small amount of benefit.
My experience has been in both high volume financial systems and consumer startups and I personally have found that sometimes ACID behaviour is near essential. Or at least near essential to keep the implementor sane :-) And other times it’s just a hinderance. Nothing in my article implied that I’m out to solve all these, interesting cases I hope - rather that a particular pattern which can be implemented in different ways has a set of benefits. For the record Neo4J supports ACID transactions, in case I have inferred otherwise.
What’s interesting to me is that in various of the comments I have received seem to extrapolate my article into, I believe an attempt to remove or somehow change the working of RDMS systems. Which I can honestly say has me at a loss, goodness knows what I said to trigger that view. No offence to the commenters I just am uncertain where this came from. What I was discussing originally was a design pattern which I have had recent experience of within the context in my case of persistence.
Again forgive me if I misunderstood but I believe you are talking the internals of your database when you’re talking MVCC. This wasn't the problem we’re trying to solve. In fact Neo4J is a very high speed transactional database that we’re more than happy with.
In fact we’re not trying to solve a problem, we needed versioned data so we implemented that, now we’re getting lots of benefits including zero cache invalidations which means we have a safe write through cache, which of course can be caching application style representations of the data (i.e. objects). And the various other benefits I mentioned.
I can only state my practical experience here, which is that we do get a massive benefit from being able to arbitrarily cache data without fear of stale data.
But hey horses for courses, enough from me for tonite - a huge pile of code awaits me and a beta needs to get underway - if you’d like to get back to me, please drop a comment on the original article and I will be pinged. Always happy to chat.
Thanks for the discussion!
Neil
You use caches when it doesn't matter, which is much of the time. If it really really matters, then you use ACID semantics of the DB to ensure that.
Adding a history or audit always seems like a giant hack.
Actually someone else just pointed me in the direction of Event Sourcing - every pattern has a name somewhere - http://martinfowler.com/eaaDev/EventSourcing.html
And having eventual consistency with MVCC implemented with vector clocks really rocks.
Riak takes Brower's CAP theorem seriously and even lets you configure the values per bucket so that you can get more of C, more of A or more of P.
You don't need to store all the history, just the contentious history. Riak pretty much rocks in the sense that you can add nodes to it and it just works, or take them away and it just works. Just like Dynamo. (Not like Cassandra.)