MemSQL is now free to use for databases with up to 128GB of RAM usage
memsql.com
memsql.com
* MemSQL is a distributed, in-memory, SQL database management system.
* It is a relational database management system (RDBMS).
* It compiles Structured Query Language (SQL) into machine code, via termed code generation.
* On April 23, 2013, MemSQL launched its first generally available version of the database to the public.
* MemSQL is wire-compatible with MySQL.
* MemSQL can store database tables either as rowstores or columnstores (The OLAP vs OLTP part I guess).
* A MemSQL database is a distributed database implemented with aggregators and leaf nodes.
* MemSQL durability is slightly different for its in-memory rowstore and an on-disk columnstore.
* A MemSQL cluster can be configured in "High Availability" mode.
* MemSQL gives users the ability to install Apache Spark as part of the MemSQL cluster, and use Spark as an ETL tool.
The main value proposition seems to be the distributed nature, which probably makes it easier to setup out of the box than, say, trying to setup a cluster MySQL or PostgreSQL databases which are not "natively distributed". Also, probably most useful when the data is "big enough" vs resources available on any single server or when reliability is very important.
This announcement, that it's now free to 128GB pretty much saved me from having to do a kickstarter to raise funds for my little SaaS project.
If anyone from MemSQL is reading, really thanks for doing this. I think it's a very shrewd move. Should definitely see an upswing of potential customers adding it to their stack and when they grow, become enterprise customers.
The 128 GB limit applies to the whole cluster. So if you have two nodes in the cluster they would each have to be 64GB or less. If you have four nodes they would all have to be 32 GB or less. To have a highly available system we recommend 4 nodes (a master aggregator, a child aggregator and two leaf nodes). You can read more about the cluster architecture here: https://docs.memsql.com/concepts/v6.7/distributed-architectu...
It's not open-source, but there is plenty of closed-source proprietary software, and plenty of buyers who care about solving their problems and paying money to get that done (and ensure the vendor stays alive).
If you don't want to use a closed-source product then that's your prerogative, but I don't see how you're making a dev/engineering decision by ignoring a product because of that.
That hasn't been my experience, at least not on any suitable time scale.
I strongly suspect that the vast majority of those of us who have worked somewhere "not more capitalized and viable" than the vendor share that experience.
Even when a vendor's support engineer is fully capable of solving the problem, the sense of urgency can't reasonably be expected to match that of a much smaller customer facing potentially catastrophic data loss (or other existential-threat-level consequences).
I do not see how having spare engineering talent capable of reading, editing and running a custom database build is the more realistic or faster option for any business in case of issues.
Those are totally useless during an existential crisis without associated indemnity (which any vendor would be crazy to provide) against loss due to failur to perform.
> I do not see how having spare engineering talent capable of reading, editing and running a custom database build is the more realistic or faster option for any business in case of issues.
I don't see how it isn't, considering that "custom database build" could be so simple as to be trivial. In the GP's case, it was merely using a specific version.
Even the characterization of the required engineering talent as "spare" seems incongruous, as, in small companies, the talent requird to handle unexpected problems with technlogies fundamenta to running the business is essential, not superfluous.
As a counter-example, rethinkdb is open-source but the company failed and nobody cares about using it anymore. What would you do with that? Start building new database features yourself? Or just get your data out and move to a different system?
Databases are so lock-iny and critical that it's only natural for closed source database startups to be considered too risky to touch.
It's better to practice proper vendor management and weigh all the risks and realities instead. If you're not more capitalized and viable then your vendor, then you have more important things to worry about then your vendor disappearing overnight.
A real technical reason might be "the ability to fix the software ourselves" or "easier debugging of the software". Avoiding vendor lock-in isn't a technical decision
There are a variety of ways to try out MemSQL yourself such as installing on Linux, Windows, Mac, AWS, etc, and maybe I am biased since I was an engineer before I became a PM, but we optimize our product currently exactly for technical people such as IT, devops, and of course, engineers. For a list of installation guides, check this link out: https://docs.memsql.com/guides/latest/install-memsql/
Take a look at our docs (docs.memsql.com) and you will see that we all actually are just a bunch of engineers and people with a technical background. Are there certain technical topics you feel are unclear here? I'm also happy to chat privately.
If you still feel this product isn't right for you, that is fine -- MemSQL's focus on query speed may not be for everyone. However, with this release of having a free product for people to try out, we definitely optimized exactly for people that want to try the product out :). I'm actually surprised you mention our product isn't for engineers/technicalPeople, because from our field of view, we actually sometimes see MemSQL as too technical, hence why we focused on usability in this release, ha!
Hope that answers some doubts you may have -- thanks for the comment.
I really don’t see how you could have a problem with that.
I have never seen a productive system that small. I’ve seen thousands.
- MemSQL is transactional and writes transactions on disk
- MemSQL has an excellent implementation of SQL with mature query optimization and query execution. And it get better every release. This is from 6.5 https://www.memsql.com/blog/6.5-performance/
- MemSQL has in-memory and on-disk data storage so you can use MemSQL to store petabytes
- MemSQL has columnstores and vectorized query processing: https://news.ycombinator.com/item?id=16617098
- MemSQL supports geospatial, fulltext search, and json
- MemSQL allows you to stream data from kafka in one command: https://docs.memsql.com/sql-reference/v6.5/create-pipeline/
Genuinely not intended as snark, I'm just curious why memsql is so compelling that I would consider it.
All things being equal, I agree that open source solutions are the best. Things are just not always equal.
Truth be told, the list if features is not very compelling. I mean, JSON support is not a reason to pick a commercial dbms over a FLOSS one.
- In-memory row stores. Super fast for updates and point lookups
- Memory optimized hash joins that minimized cash misses. Great for analytical/reporting use cases
- Vectorization for columnstore query processing. Super fast aggregations that work best when data is cached in memory
Beyond those considerations, why would the same exact same query (executed several times in rapid succession from the console) produce vastly different results? Also, I should clarify, rewriting the query from "select ... from xyz group by ... having ..." to "select ... from (select * from xyz where ...) group by ..." made the inconsistency goes away, without changing the filtering clause. That does not inspire confidence.
select count(*), a from T group by a having b > 0
In this case b is not allowed to be part of having by ANSI standard.We let it run b/c some customers migrate from MySQL and MySQL allows this query. You can set MemSQL to be strict about it by setting this variable:
set session sql_mode = only_full_group_by;MemSQL is a distributed full-featured relational database that has in-memory rowstore and on-disk columnstore tables with rich support for SQL, fulltext search and JSON. It's a fast RDBMS and does really well with analytical queries.
Do you need a fast cache, key/value, messaging system? Or a RDBMS with fast OLTP + OLAP capabilities?
Column-oriented storage itself is many times faster for analytical queries, even if on disk, and combined with the other optimizations of MemSQL will get you far better performance. Along with all the data being able to constantly undergo transactional updates.
This post 7 months ago [0] mentions the same.
It's probably still at this price, but if you are really interested. Yeah, contact sales.
It's not (necessarily) about high throughput, but about low latency.
I guess you could be using MemSQL for post-order or trade analysis but then it would probably overkill since a lot of that can be done considerably slower.
But yes, actual HFT is not an accurate use-case, for any database product.
We're experimenting with MSSQL's memory optimized tabled and native compiled stored procedures. My timings today, I was getting one call to our stored proc in the 300-400us range, that was inserting one record each into 2 tables.
Test setup for all of my scenarios are do all of the same DB ops, I'm alternating which libs I'm using. Best performance so far ive been able to get from linux talking to MSSQL has been using OTL on top of unixodbc with MS Driver 17. Mind these are physical servers sitting a few feet from a shared router.
CMU had guest lecture by memsql founder about the architecture.
Not sure how much of it changed in past 2 yrs tho.
When it comes to commercial products, it is generally good to check if you can afford their paid offering, then only make the tool core part of your infrastructure. For databases, it's better to stick with fully open source popular options from a long-term perspective.
They are entirely different systems. Sqlite is meant for self-contained applications that need some relational data persistence with a single file for storage, not for accessing as a central database with many clients storing TBs and scaling across nodes.
(Quote from their docs)
is MemSQL shared nothing or shared everything, or can be mix of both?
I think most engines guarantee Durability by assuming that once on disk, it won't go anywhere but if it's in RAM, it is susceptible to power outage? If it gets written to disk, it's not as fast as RAM?
There are other reasons why an in memory database can run faster than a disk based database, such as not having a buffer pool manager, but I don’t think that’s what you were worried about.
MemSQL has two storage modes Rowstore and Columnstore. The Rowstore is "in-memory" and the columnstore is "on-disk" but those are oversimplifications. The rowstore data is stored in memory but we keep a snapshot of the data on disk. We also keep the transaction log (a record of all changes since the snapshot was taken) also on disk. So queries can be satisfied fully from memory (because that is where the current data lives) but writes go to memory and to the transaction log on disk. If the machine reboots then the snapshot is loaded from disk back into memory and the transaction log is replayed. When that is complete you are back to where you were when the machine rebooted with no loss of committed data. Columnstore data is always stored on disk although we use a row store in front of the column store that is hidden from the user but acts as a buffer of sorts so that writes in the column store can be pretty fast. More details on how the columnstore works can be found here: https://docs.memsql.com/concepts/v6.7/columnstore/#how-the-m...
> You can do almost anything with MemSQL, using the free tier, that you can do if you have an Enterprise license, including capabilities and production use. The differences are that you can only configure the free tier of MemSQL to use up to 128GB of RAM usage, and support is only community support; for paid MemSQL support, you need an Enterprise license.
We would have recent data in rowstore and move older data into columnstore. You can easily join/union between both for queries. Also some constantly changing data (like budget counters) would always remain in rowstore with many lookups and updates per-second.
Why not share a clear example of exactly what happened, or post on their forum with details, so we can all judge for ourselves?
It is a redis module that embed SQLite, bringing on the table a lot of advantages. Extremely fast and with it you can even upgrade your Redis instance to be your only database.
There is not a huge company behind it, but I really do my best to support it. I don't believe to have disappointed any of our users so far.
Also the tech documentation: http://redbeardlab.tech/rediSQL/references/
For any issues don't hesitate to contact me directly or through GitHub.
Also if you need it for some open source/(do the world a better place) kind of project I am more than happy to provide the PRO version free of any charges or obligations.
Of course it apply to anybody reading!
Note: you have a typo on your home page: "never loose a bit" 'loose' => 'lose'
It works on top of Redis so it inheritance its interface. I guess it would be possible to add an ODBC layer but honestly I have to look into it...