Chartio develops non-blocking MySQL queries using Tornado
chart.io
chart.io
http://jan.kneschke.de/2008/9/9/async-mysql-queries-with-c-a...
with a little more in-depth exploration of the MySQL async API. The problem seems to be that that (undocumented) API does not handle EAGAIN and that there's no way to connect asynchronously (mysql_connnect always blocks).
Compare with PostgreSQL, which has had an async API for a long time:
http://developer.postgresql.org/pgdocs/postgres/libpq-async....
and for which I wrote a Twisted library that does fully asynchronous connection building and query execution: https://github.com/wulczer/txpostgres.
Since the asynchronous API is exposed in psycopg2, the Python connector library, it should be trivial to hook it up to Tornado as well.
I have a much more mature version internally, and my understanding is that you can get gem-ified libraries that do async database stuff for EventMachine today.
I lost interest in this pretty quickly once I added Redis to my stack. Redis is trivial to talk to asynchronously, and by sticking a queue in between your async components and your database layer, it becomes easy to implement the half-async model --- plus you get very smart caching and another layer of indexing "for free".
Also, some Twisted documentation was improved after Tornado was released; there's now "Twisted.Web in 60 seconds": http://twistedmatrix.com/documents/11.0.0/web/howto/web-in-6...
One more reason to be leery of single-threaded eventing systems. You'd never run into this issue with a threaded web app, and it would perform just as well provided you kept your datastructures as independent as they are in your current eventing setup.
Comparing an eventing system and a threading system, the eventing system is inherently less shared-nothing. It shares everything the threading system does and the event thread.
But that's not why I'm commenting. Rather:
You're able to make that last assertion only by shifting the meaning of the word "shared" and denuding it of all its concurrency implications. Yes, event systems "share" the event loop, in all the glory of the word "shared". However, no two contexts in an evented system ever step on each other for access to a shared resource.
I read a little more than that, but I commented on what was interesting to me. If I'm reading the rest of it right, it's basically a tutorial on making a python extension for two specific mysql API functions. That's fine, but it's not that interesting (to me).
> However, no two contexts in an evented system ever step on each other for access to a shared resource.
Isn't that exactly what is happening when other requests are blocked by a blocking mysql call? They are stepping on each other for access to the shared event thread resource, which they need concurrent access to. Is this not the case? Please help me understand if I am misreading you.
While we could've gotten better performance with more processes, we wanted to stick with one process per core. Given this and a non-blocking driver we naturally got a good performance gain.
amysql, as it's called, is up on GitHub: https://github.com/esnme/amysql
The rational hope is that by the time you get to that scale you have people/code that know what they're doing and can fix it. Alas, there's many cases where that's not been true.
To add to the list of truly non-blocking MySQL libraries, check out Erlang's MySQL driver: https://github.com/Eonblast/Emysql
It still doesn't have a native boolean type in SQL, for example. After 30 years.
* It's already in place
* Everyone on your team only knows MySQL
* That's what you know and you are under a deadline
* It's the right tool for the job (as opposed to something very different, from the NoSQL world)
* You are using MySQL Cluster instead of MySQL, but don't have the cash to upgrade to Oracle
If I may toot my own horn, I just upgraded the layout of http://mongodb-is-web-scale.com/ so that you can scroll while the video stays visible.
Clustered and covering indexes. If the data needed by a query is in an index there is no need to retrieve it from the table itself. This leads to better memory usage and cache locality.
MySQL has non transactional tables which use a lot less memory. If you store a large number of small records (say three ints or something) that table will use about a third of the memory a similar postgres (or Oracle) table would.
Mysql:
- Never crashes
- Has more tested/reliable replication than anything else, possibly including Oracle (haven't tried the wacky 3rd party replication things like Goldengate or whatever, and don't believe in spending 6 figures on software licenses in general).
- Supports SELECT, INSERT and UPDATE
Why would I get anything else for a database? For more complicated stuff I have another approach that I employ called "programming".