HNHacker News
TopNewBestAskShowJobs

jdf

108 karma · joined November 13, 2010

submissionscomments
jdf··on How can you program if you're blind?
I had a TA at school who was blind and was excellent at helping CS students debug and fix their class projects. He helped me a number of times and was generally able to find bugs very quickly.

While I don't know the name of the program of he used, my vague recollection is that it read aloud the contents of whatever line he had the cursor on, a good deal faster than a human would speak it (I actually found the text to be read too quickly for me to follow by listening - I needed to read it myself).

I actually found this article about him on the Internet:

http://www.dimenet.com/disnews/archive.php?mode=P&id=276

jdf··on HyperDex: A Searchable Distributed Key-Value Store
What happens if a crash happens part way through the acknowledgement chain?

For example, in your insert, if the node containing "x" crashes before it receives the ACK from the node containing "y" - do the dangling "y" and "z" insertions ever need to be cleaned up?

jdf··on HyperDex: A Searchable Distributed Key-Value Store
It looks like they trade the ability to scan ranges of keys for the ability to get single objects via multiple attributes. The value-dependent stuff is also a neat way of solving the consistency issues with multiple node updates. Interesting stuff.

If you were to swap this with your Cassandra cluster, you'd be losing multi-datacenter replication. Although partitions within a data center are pretty rare, you'd also lose some availability there as well. However, Cassandra is usually hash-partitioned so it needs to do broadcast for a scan (AFAICT, it even needs to do a broadcast for a lookup on a single secondary attribute), so you'd probably gain quite a bit of performance with HyperDex.

I can't tell if it's possible to dynamically change the set of secondary attributes being indexed without rebuilding the entire data set. Or how value-chaining works with missing attributes.

Also, apparently consistency has some... gaps... when you search via a secondary attribute:

"The searches are not strongly consistent with concurrently modified objects because there is a small window of time during which a client may observe inconsistency."

jdf··on Show HN: my database engine for GPU
If you look up columnar databases, you can find a whole host of research on how these things store data. Storing values column by column to do vector processing is pretty common, but I can't recall seeing one built for a GPU before. Kudos to the dev for diving in to implementing one, it's pretty neat stuff.

Try finding research papers about MonetDB, Vertica/C-Store, or Vectorwise for some background. Or follow the links here:

http://www.dbms2.com/category/database-theory-practice/colum...

jdf··on Hadoop Reaches 1.0
You can try out MapR's Hadoop distribution, which uses its own filesystem rather than HDFS. The MapR filesystem is not append-only, handles small files well, is easily accessible over NFS, etc.

(*Note: I'm a MapR employee, so obviously I'll be biased towards thinking our stuff is great)

jdf··on Doozer: a consistent, highly-available data store from Heroku labs
There are a number of other systems that allow this same approach of consistency and high availability. For example, Cassandra, which is freely available (as required by the poster), appears to be able to give you this behavior if you set ConsistencyLevel to QUORUM.

Clustrix, the company I work for, offers a full SQL data store with similar quorum semantics. However, it's not free.

Google Megastore allows similar consistency semantics (with its own data model) in a cross data center "cloud" fashion. It's also not free, but it would probably be suitable for some set of Heroku customers, particularly if they're already using Google App Engine.

jdf··on C++ in Coders at Work
Similarly, I've been hoping that Rust (which is somewhat run by Mozilla labs folks) would turn out well:

https://github.com/graydon/rust/wiki/

jdf··on MongoDB vs. Clustrix Benchmark
> Yes, I did not mean that Megastore is eventually consistent. It is not, it's ACID compliant and uses Paxos. Real question: was this not stated clearly in my comment (I'd like to edit to clarify if it is)?

I was interpreting the nod to Megastore in the context of the first sentence, which seemed pretty strongly about eventual consistency. Nothing to worry about, though - your followup comments are very clear!

> Megastore, however, is an example of a "NoSQL" system

Looking at Google's recent papers, they actually seem to be moving more in a SQL-ish direction. Dremel, in particular, appears to be a variant of SQL that allows easier access to nested and/or repeated columns. I thought Megastore was in a similar vein, but at that moment I can't seem to find the link that made me believe that (other than seeing that it uses strongly typed schemas).

Regardless of the benefits of a SQL-like language as an interface, however, I think you have a good point: the underlying architectural decisions about things like consistency or stability are very interesting, but often obscured.

jdf··on MongoDB vs. Clustrix Benchmark
The term "rows" could certainly be replaced with "tuples", "objects", or whatever else you prefer without changing any meaning of the post. If you look at the benchmarks posted on Mongo's site

http://www.mongodb.org/display/DOCS/Benchmarks

you can see that most of them compare Mongo against MySQL. Certainly "rows" are being inserted into the latter.

jdf··on MongoDB vs. Clustrix Benchmark
Eventual consistency often gives up the ability to develop applications against the system, because more reasoning about the state of the system has to be embedded in the application logic. For a better description of this problem, see this post from an AWS employee:

http://perspectives.mvdirona.com/2010/02/24/ILoveEventualCon...

Note that his "right answer" for performance (not availability, so it's a bit of an aside from your comment) is to dynamically repartition, which is what Clustrix does. Note that we can actually do this type of repartitioning on just the "hot" data, using MVCC to avoid any downtime (unlike Mongo, which blocks all writes).

By the way, the Google Megastore project I believe you're referring to is not actually eventually consistent in the same way that, say, Cassandra is. Megastore is fully ACID within an entity group, meaning that lowering eventually consistency down to the per-node level was not the solution that Google went with. BigTable is also not eventually consistent.

*Clustrix employee

jdf··on MongoDB vs. Clustrix Benchmark
Looking at your first SELECT, there's very few RDBMSs on the market that won't evaluate that WHERE clause prior to the JOIN.

(disclaimer: I work for Clustrix)

← PreviousPage 2 of 2