Our app is Java. Even in "short pointer" mode, 30M references to the same string would eat up 4 bytes * 30M = 120M of memory! And, the garbage collector would be displeased because it frequently would have to sift through mature generation regions. So, we are left with disk-based options (a traditional RDBMS included in that mix) or some variation of structs in byte arrays/buffers. [I should mention that we only must provide access to these query results for an hour or so.] While I have found some reasonable flyweight pattern-based schemes for representing structured data in large byte arrays (Javolution has a nice one), I have not found any good data structure libraries which use byte arrays as backing stores. Because the data does not change once queried and inserted into our storage scheme, it seems reasonable to use some look-up schemes (hashes, etc.) which could be persisted in a manner similar to the raw data itself. An additional level of abstraction would be nice, but I'll settle for not writing my own hash implementations.
The route I'm taking now is to use the H2 embeddable, all-Java DB. It has pluggable backing stores, so I can make it pretend that direct byte buffers are disk-like. I don't really need write concurrency, transactionality, and all the other stuff "real" RDBMS systems provide, but this seems like a good trade-off compared to rolling my own in-memory sort-of DB. It may turn out to be fast enough to just run H2 in a disk-based mode.
Does anyone have any thoughts on this?