1) Facebookers: Check
2) Data Scaling experience: Check
3) In-Memory with SQL semantics: Check
4) New-York based software that can be sold to quant funds: Check
I'm excited. It's downloading now.
1) Facebookers: Check
2) Data Scaling experience: Check
3) In-Memory with SQL semantics: Check
4) New-York based software that can be sold to quant funds: Check
I'm excited. It's downloading now.
The bullshit-bingo-lingo on their homepage is mindnumbing.
Meanwhile their actual software seems rather underwhelming, bordering on SnakeOil.
You can buy something like kdb, but that costs $70k a year and requires that your engineers learn some extremely new semantics if they're only used to SQL and (choose any popular language here, Python to C++).
Anyway, why is there not a single benchmark available to support the sales-pitch?
(In your favor I'll just pretend you didn't mention kdb here...)
This ties with my own experience using it.
For example, where can we find your configurations for the MySQL vs. MemSQL benchmark you show in your video? Or how big the dataset was, etc...?
Please explain. What's bad about kdb, besides how nonstandard it is?
At the least they should come up with some seriously impressive benchmarks before dropping names like that.
Did you mean large, fittable-in-memory, and (distributed across multiple computers and/or very reliable)? Because that is much harder.
KEY and INDEX are literally the exact same thing and are just there to support different syntaxes: http://dev.mysql.com/doc/refman/5.1/en/alter-table.html
Perhaps you are referring to primary keys vs secondary indexes, for which there is a difference in how they are physically stored on disk and in memory, particularly with InnoDB. There is a substantial performance gain to be had by having a relevant primary key that is referenced in the where clause as the data is stored in the same page as the primary key, both on disk and in the buffer pool. Secondary indexes reference the primary key, thus require two lookups.
The way I see it, this is very much targeted at financial firms that use kdb+ or products like it that perform really well for things like writing lots of tick data from dozens of exchanges at a time and querying across them quickly in a familiar SQL style.
But that's only a query language. Not a db architecture.
Let's see the source code. That's the only way to verify the approach that is being taken.
Given the source the first thing I'd do is search for "mmap".
They already get one demerit for using a row store.
And another for disk access.
SnakeOilSQL.
It's not a column store. We are thinking about it, maybe some time in the future as we add more analytical feature. It's a row store.
We do write back to dist. We have a transaction log, which we truncate once it grows over a certain threshold, by taking a snapshot.
We see a future where OLTP databases live in memory, and where you have hundreds of terabytes of memory and hundreds to thousands of CPUs at your disposal.
We built MemSQL to make it easy to go directly into memory and get the speed improvements we all need. What we're releasing aims to fulfills that vision.
Also, pricing. Come on, man! Pricing!!
p.s. Isn't a totally reasonable business model for a truly faster DB engine "free, everywhere, bought by Oracle?"