Tornado Without a GIL on PyPy STM
morepypy.blogspot.com
morepypy.blogspot.com
New users also want a quick explanation as to why
Julia is fast, and whether somehow the same “magic
dust” could also be sprinkled on their traditional
scientific computing language. [...] Julia is
fast because we, the designers, developed it that
way for us, the users. Performance is fragile,
like accuracy, one arithmetic error can ruin an
entire otherwise correct computation. We do not
believe that a language can be designed for the
human, and retrofitted for the computer. Rather a
language must be designed from the start for the
human and the computer.
To me, Python and Ruby are both perfect examples of languages that were designed for the human, and ever since have seen extensive effort to retrofit them for fast execution by computers.I respect the work of the PyPy team, particularly given the raving reviews I've seen lately of how RPython is a boon to language designers who can use it to prototype their languages and get a decently-performing VM in not very much time: http://tratt.net/laurie/blog/entries/fast_enough_vms_in_fast...
But I can't help but think that languages like Python and Ruby will start to fall to languages like Swift, Julia, Go, etc. that were designed with performance in mind. I'm not saying this will happen soon, but these languages are showing that you can have your cake and eat it too.
I'm not sure how JavaScript and Lua fit into this analysis. They weren't specifically designed for performance, but have been very successfully optimized. Lua is a very simple language and Mike Pall is a genius, so LuaJIT has been very successful at speeding up Lua. JavaScript is a little more complicated, but has received an immense amount of resources into optimizing it, and has also been quite successful at getting fast.
Python is the language specification that is probably most impaired by performance on multi-core, due to the fact that threading semantics were so strongly specified. Pypy is now making a compelling argument that even that is not an insurmountable barrier.
If by "retrofitted" they allow backwards-compatible changes to the languages in question (e.g. optional type annotations for dynamic languages) then I would estimate that a language not designed at all with performance in mind would suffer less than a 2x penalty given enough engineering effort.
Where does Lua then fit into your analogy?
http_server = HTTPServer(Application(),
xheaders=True,
)
http_server.bind(port)
http_server.start(0) # Forks one sub-process per core
[1] https://bitbucket.org/kostialopuhin/tornado-stm-bench/src/65...Are most of the interesting problems to solve always limited by Amdahl's Law? Will we never see the gains of single-core speed we saw in the last century again?
Just make the visited node cache public and immutable.
Two, the graph your building itself must be shared, obvs.
If the graph is immutable then there's zero problem with it being shared.
The cache has to mutate and be shared as that's the work completed list. As each thread completes a bit of work (visits a node) it needs to communicate it with the other threads.
var cache = Seq(1,2,3)
cache :+= 4
The cache is immutable and freely shareable. Any other thread and come in and read it and be guaranteed that its current state is valid.Meaning the cache is no longer shared as all threads end up having a thread local cache. You can't update shared changing state and have immutability at the same time.
Immutability is great when one can have it but sometimes its not possible. Shared changing state is something to be avoid as much as possible but sometimes we need it.
But I should say that there is still too much overhead in using STM; you will still be able to very easily (and by a large margin) outperform STM-4 by running 4 instances of Tornado with HAProxy or some other lightweight router ontop. A comparison graph for this should have been the benchmark.
Python's feature set is basically what's easy to do in a naive interpreter. Everything is a dictionary. Anything can be changed from anywhere. With "getattr", you can patch one thread from another. It's elegant, and very difficult to speed up. Google tried, with von Rossum on board. Their "Unladen Swallow" JIT compiler project crashed and burned.
The PyPy group has made a fast Python compiler/interpreter/JIT system. It's really hard. The initial funding from the European Union got them started, but wasn't enough. They really try to handle all the hard cases. This requires two interpreters and a JIT. They have to handle "oh no, someone patched object A from thread R, invalidating code that's running in thread S". There's a "backup interpreter" that kicks in for hard cases, and once control is out of the area in trouble, the JIT can recompile it. (This is an old and oversimplified description.)
This transactional memory thing is very clever. It has to separate things at run time that probably should have been separated at compile time, of course. It's impressive that they can get it to work. It's a lot like how a superscalar CPU works, including transaction commit and backup at the retirement unit.
Python gets into this mess because, like C and C++, the language doesn't really know about concurrency. (Threads came late to UNIX, and C predates threads. So C has an excuse for backing into concurrency.) Python has the C model of concurrency; at the user level, it's treated as a library issue. Internally, though, it needs a lot of locking, because there's so much mutable state in the interpreter.
It would be a lot easier if the language were restricted a little. But then It Wouldn't Be Python(tm). The price of this is huge complexity layered on a simple model, and probably years of obscure bugs in PyPy.
Other than Erlang, which modern language doesn't?
Java has e.g. "synchronized" keyword, but I regard it as merely syntactic sugar over a standard library implementing thread - semantically, it still basically has the C model.
Go has channels and stuff - but it's still library level (in fact, it's basically syntactic sugar over the Plan9/Inferno/Aleph/DontRememberName standard library primitives)
There were other OS with threads and co-routines being explored as design while UNIX was being developed at AT&T.
As for the rest I agree with you.
Personally I don't have any use for Python besides the occasional shell script, but as a user of applications written in Python I would like them to perform fast.
Unladen swallow was attempted at a time when llvm jit was buggy and slow.
Maybe if there were competing implementations to ruby, it would still be popular outside the rails community. But it seems as if the vast majority of ruby core developers work on the mainline implementation, or forks thereof.
People don't like to say it to your face, but Python is often derided and considered to be pretty silly. Oh, and its syntax is like Lego. Neat at first until you want something like a multiline lambda.
People don't like to say it to your face, but C is often derided and considered to be pretty silly. Oh, and its syntax is like Lego. Neat at first until you want to understand what something is a pointer to.
It used to be the assumption that there was a direct tradeoff between performance and convenience/productivity of a programming language. I think newer languages are showing that this doesn't have to be true, at least not to nearly the same degree.
PyPy is a project to bring some optimization to Python. Basically make Python run faster. It is also effectively another implementation of Python. CPython is the default one (the one you download at python.org). There is also Jython (running Python on the JVM), and PyPy, and a few others.
NodeJS is the marriage of the V8 Javascript interpreter and JIT with an asynchronous IO library (libuv) + a large ecosystem of modules.
You might want to compare nodejs with PyPY+Tornado or with PyPy+eventlet. Read about STM and the idea behind it. STM lets you take advantage of multiple cores. nodejs is single threaded. In practice nodejs might be faster currently just because V8 is very good and depending on workload if most of the stuff it does is just proxying data from one stream to another, it might do pretty well. But if you start doing a large number of concurrent requests where each request has do to some logic, PyPy might come out on top.
In general anywhere with complicated business logic or a large number of steps needed in the backend to handle requests, I wouldn't use nodejs. I never liked the callback/errback paradigm for large concurrent applications. That works for demos and short web tutorials, in practice, I don't like how it looks. I like green threads, lightweight processes better.
Of course not, Maciej ;-) PyPy is awesome and thank you and the whole team you have been doing a great job so far. Just looking at performance graphs and speedups gained over the years. It looks very impressive.
Yeah I don't have a citation but from experience, I have noticed where there are a few steps in each request processing (think proxies) solutions based event loops (epoll, kqueue and friends) can outperform those that spawn a thread/process/context. For example haproxy is certainly a very well done fast proxy, it is single threaded and it seems to work for it.
Again sorry for misunderstanding, I was just speculating without any benchmarks or even particular applications in mind.
Cheeky comment aside, PyPy is an alternative interpreter for Python and, more broadly, a meta-frame work to make writing tracing JIT for the language of your choice easy.
It's much faster than regular Python. The only downside (which, admittedly, is very, very big) is that the PyPy developers released their own FFI library for interfacing with C-code (which works brilliantly), but using c-types and the c-python c-api is very kludgy. This means that a lot of very important python libraries which are basically thin wrappers around c/fortran code (SciPy, NumPy, etc) are not really working.
There is some effort around reimplementing NumPy in PyPy, but it's going very slowly.
Edit: PyPy is also a heroic effort by a small team of developers who receive some donations. JS is probably the language that received the most amount of money and attention into making it run fast from Google/Mozilla, etc - in terms of results/money, PyPy is incredible.
It's easy enough to use CPython extension modules via cpyext, but it's really slow.
Python also varies in speed over time and implementation. Implementations have speed. There are reasons that people use "slow languages". It's not just developer happiness. You really do have to look at performance holistically.
If you want to speak to microbenchmarks, I can. For example, CPython's std lib JSON processing is done in C. It's very fast. As of Go 1.2, PyPy blew the doors off Go in my tests processing 10,000 JSON records. CPython 2 & 3 also beat Go in my tests as well. What's this mean? Does it mean Go is a slow language? Well, it's one microbenchmark vs another. It really doesn't mean much to your application.
"I don't know why slow-language people bother throwing up the argument that benchmarks aren't everything."
I'd go even further than 'aren't everything'. Benchmarks that aren't your application do not mean anything.
Many are ignorant of how CPython is even built and are shocked when I show them how fast many standard library modules are. It's just not so simple. Don't believe the hype.