What Blocks Ruby and Python from Getting to JavaScript V8 Speed?
stackoverflow.com
stackoverflow.com
I talked with the Unladen Swallow guys sometime after the project was canceled. They said that the main issue was not getting Python to run fast on LLVM, where they said they had some really encouraging results. Instead, it was:
1. Bugs in LLVM which they had to spend a lot of engineering effort working around. This is probably much less of an issue now, with Clang, XCode, Julia, and Rust all being used in production with a good deal of success.
2. Compatibility with the C API. This was considered a "non-negotiable" for Unladen Swallow's design, and one major reason they started by forking CPython rather than building from scratch on LLVM. The problem is, the C API makes a lot of assumptions about how objects are laid out in memory, how they are memory-managed, how they are accessed, etc. If you must support defining & calling methods from C, there's a lot less freedom to, say, stick a trampoline in the method preamble or a PIC at the call site or switch to using a tagged representation of pointers and ints. That means that the VM has to fall back on the interpreter (and possibly waste cycles converting & marshalling data) for a lot of things that are supposed to be fast, which kills many of the real-world performance gains.
V8 had the luxury of defining its own C interface. I've often wondered whether a Python that gives up C extension compatibility could get similar speed gains - I even ran the idea by a former Python core team member who worked on my team at Google, and he said there was no theoretical reason why not. But then, Python without C extension compatibility would be Python without Numpy, Scipy, scikit-learn, PIL/Pillow, JSON, StringIO, and many other critical packages. At that point you might as well start over with a new language.
Would we really miss them if Python was then fast enough to do these things natively?
For the language I work on we were able to write macros and functions to allow our C libraries to continue to work when we changed our VM, but we've spent considerable effort optimising things so the C has to reach back into objects as little as possible, and we could only do this because recompilation and refactoring was okay.
I assume you're unaware of the numpypy project then, which is doing exactly that, with reasonable results.
I also think you overestimate the performance of numpy/scipy when compared to, for instance, C++ template-based libraries which are able to make large performance gains from compiling specializations for the specific array shapes being used. JITs can potentially go further, being able to compile specializations in places where array shapes are not known at compile-time.
You may not be aware that NumPy does it's heavy lifting by offloading the work to C and FORTRAN, and that's where the speed comes from. You set up your data in Python, use the nice high level language to parse files, etc, then you pass chunks of memory containing matrices off to the FORTRAN to crunch, then use Python again to say write it back to a file.
Nah. Note that these are optional even in the original numpy.
I'm not sure why Fortran wouldn't be considered a high level language, if that's the implication (and it hasn't been SHOUTED since the 70s).
Anyway, this thought experiment has basically been done. "Ditch the backwards compatibility, start with a new language, make it fast by design, and rewrite all the packages that make it super popular." It's called Julia. And it's popular among some people, but for a lot of others, the maturity of the Python ecosystem still makes it a lot more attractive despite the performance & multi-language issues.
Arguably, lack of commercial support is a bigger problem for mature projects than young projects. Whether or not Julia pulls away users from R probably has more to do with how much R adapts to new directions in statistics than with anything Julia does (assuming that julia passes a minimum threshold of credibility, which it probably doesn't yet for most statisticians.) A lot of my interest in julia comes from a specific simulation that I need to run to prototype a new estimator that would be infeasible in straight R. If R were faster, I wouldn't bother.... (To preempt the obvious reply, I could always write the slow part in C. But if I need to introduce a different language anyway, why not look at other options?)
PIL a little more, still quite fast to port though (a month or so).
No reason why Numpy and Scipy couldn't be written for a new C extension interface too, though that would take much more time (1-2 years).
Thing is, with 10-20 million dollars, Google could have sponsored all of the above.
Writing the code is rarely the hard part. There's a whole ecosystem around libraries and languages that takes time to build.
For a JSON parser? I think you overestimate the difficulty in implementing one. If anything, a parser for such a format is one of the easiest things to know you got correct -- there are extensive test suites.
As for switching to it, Python users already use several JSON libs anyway, so it's not like adopting another would be any great deal.
For the other libs, like numpy etc, yes, it would be more effort, but nothing unsurmountable. And speed is a far greater hook to get users to adopt that version of Python than "we now got unicode all over and fixed a couple more annoyances".
Isn't there a pure-python json parser in the standard lib already?
import json
Does in fact give you a working JSON parser. I wasn't aware there was a different one to use.* the standard library's json, which stems from a merge of simplejson way back
* simplejson
* ujson which can be much faster but is not exactly sound[0]
* cjson
* jsonlib
* demjson
* pyyaml (since yaml is a json superset)
Maybe money and promises of better performance would solve that problem, too.
> Our new solution in JRuby+Truffle is pretty radical - we're going to interpret the C source code of your extension. We use the same high performance language implementation framework, Truffle, and dynamic compiler Graal, to implement C in the same way as we have implemented Ruby.
That is not quite correct. While you can't add your own C-extensions, some popular extensions that aren't pure Python are included in the provided platform (including, particularly, numpy.)
A number of projects (e.g. this bcrypt package -- https://pypi.python.org/pypi/bcrypt) have switched from using CPython's API for C extensions, and switched to the cffi package (https://pypi.python.org/pypi/cffi), giving them what seems to be a much cleaner (and less CPython-specific) break between the C-level and the Python-level code.
I say "seems to be" because I'm not too familiar with the internals to know for sure :) What cffi did give them is that bcrypt (with it's C component) can be used by both CPython and PyPy -- even though PyPy doesn't support any of CPython's native C API.
Meaning that PyPy just has to support cffi, and is otherwise freed from any fixed C interface itself. I think the reason cffi gives it this freedom is that cffi provides Python-side metadata about the C datastructures being worked with through cffi, allowing the PyPy JIT to efficiently integrate that information into it's compilation process.
Though, it seems much easier to work around for the implementation than the existing Python C API.
Google has historically been good at high performance computing, and chose to spend enough money to get and fund a team of experts on building high performance VMs.
Microsoft has historically been good at adding backwards compatible feature sets, and chose to spend enough money to get and fund a team of experts on tasteful design and evolution of programming languages.
Put on hold as too broad by rene, davidism, vaultah, Sam, iCodez 2 hours ago
There are either too many possible answers, or good answers
would be too long for this format. Please add details to
narrow the answer set or to isolate an issue that can be
answered in a few paragraphs.
If this question can be reworded to fit the rules in the
help center, please edit the question or leave a comment.
Meta: It's interesting that 5 users thought that it would benefit the SO community to put this question on hold, while 100+ thought it was an interesting question. In one way, SO does keep the noise down pretty well, but this doesn't feel to me like noise. Does avoiding questions like these actually make SO a better place? Is there a reworded version of this question that would actually elicit better responses? Or is this just the joy of playing the enforcer?You should see how excited the moderators get when they decide there's a new rule against some new kind of question
I don't agree with closing questions like this either, but there are rules, those rules work most of the time and the question is broad. SO is not a Quora, it's not a discussion site, even though it would be nice to use the brain trust of SO in that way.
Why is python not as fast as JavaScript is a completely different question than "An internal error occurred during updating maven project". The later seeks a solution to a problem while the former only encourages discussion but is impossible to solve (because it's not a solvable question).
So, I feel this question was appropriately tagged, or perhaps even not aggressively enough.
Honestly people should read the manual then. That will give you the specific answer to your question. Stackoverflow wanted to replace other message/discussion boards in regards to programming. It's veered away from that purpose and IMHO has become cumbersome to use. If you have to worry more about how you word your question than the question itself is it worth asking there?
well, you are supposed to ask a question after you have researched the solution, so the extra worry wouldn't actually be relevant.
Sure, the answers won't help to solve my _specific_ problem. However in general I like to known how things work. It is not enough for me to read ordinary info about Ruby symbols. I have to read Ruby source code to understand them better (1) Thanks to this I understand Ruby symbols gc issue, a new feature of Ruby 2.2
Even more important is _intuition_ one can build reading the answers. Intuition about languages in general. I really like duck typing languages. Sometimes I need raw performance. So should I plan migration to statically typed languages (Java, Scala)? Or is it better to invest my time into learning Dart? Does Dart have _potential_ to solve my specific requirements?
As a software architect I have to predict a future a bit. And I must say that nostrademons answer helped me to learn something new today.
Regarding Dart: It looks nice. At work I'm forced to work in JS sometimes (I work at a game company that does contracting, and lately clients have started needing/wanting HTML5), and my opinion is that Dart looks better than it in every meaningful way. I've also begun to get the impression that Google has moved most of the V8 engineers over to Dart, but I don't have any hard evidence for this.
Honestly without knowing your problem, it's impossible to say whether or not either of these would solve it. My guess is that either of them could, as they're both languages that are orders of magnitude faster than Ruby and both seem to have large, fully featured standard libraries.
SO also keeps jQuery signals loudest.
It speaks more to the popularity of jQuery and the use of JS by those inexperienced with it than a specific signal/noise issue.
Edit: PL's aren't physics.
> The vast numbers of CL and Scheme compilers are relatively easy to write, exactly because the languages leave (somewhat) less to decide / change at runtime than Python/Ruby/JS.
I would say it's a mistake to conflate Scheme and Common Lisp. Scheme isn't even defined well enough to say something implementation independently about whether the environments should decide things at runtime or not. But Common Lisp as a language has special/dynamically scoped variables and runtime definable macros which make it much more difficult to write a compiler for it. I would say that Common Lisp has way more things to decide at runtime than python or ruby. Maybe they're about the same in terms of difficulty of writing a compiler for, if only because lisp has no parsing/lexing pass. Again, in the end, it's mostly not a language issue. It's how much money and how many compiler experts work on your implementation.
It does seem like it would be hard to add the 5.3 features to LuaJIT without compromising on efficiency, but given how clever Mike Pall is, I'd suspect it can be done.
1) Python & Ruby have lots of dependencies on the C FFI — PyPy is great, but not replacing CPython because of this.
2) A huge amount of effort is going into v8. Perhaps that level of effort will catch up to the effort put into the JVM.
3) Using esoteric, hard to optimize features in Python and Ruby is apparently a sign of cleverness. And those features are used in the major frameworks like RoR.
4) Incompatibility between other engines. Right now though you have to be backwards compatible with the web, it is acceptable to implement a small subset of the new features. People wouldn't be happy with several versions of Java / Python / Ruby all implementing some of the new features but not all of them.
But really, besides a highly optimized regex engine, v8 is still kind of slow. Just happens to be the fastest of the slow.
Also, JavaScript has a much smaller API surface area. In addition to optimizing core code execution, a decent amount of work goes into optimizing the core libraries and built-in types: RegExps, string operations, collections, hashes, etc.
In JavaScript, that set of stuff is pretty small. In Ruby and Python, it's enormous.
Python only needs to be fast enough because the running time of a typical program is largely affected by i/o time.
Secondly, if we assume the question is saying that V8 is faster, is it faster than vectorized numpy functions? Without specifics in mind, this question is just to start a language war.
While python and ruby may be getting faster they're not seeing those order of magnitude speedups that we've seen in JavaScript.
Don't believe me? look here: https://jakevdp.github.io/blog/2014/05/09/why-python-is-slow...
perl is not much older, but in a majority of tasks its faster than python, even c-API python. (it even threads!)
Yes it could be engineered around, I assume thats what pypy is doing (although STM is is going to limit performance in a threaded environment)
Javascript was designed to be an embedded language, so it was had to be quick to initialize, and have a simple object model.
Python was designed to be an easy object oriented language.
Being harder to optimize is not being slow. This is an important distinction. If you want to know what that means and the real reason why Python is slow, go read this instead: https://speakerdeck.com/alex/why-python-ruby-and-javascript-...
What an amazing "obviously"!
Ruby has what, a handful (if that?) of people at Heroku who work on Ruby some of the time? Not sure if Matz does full-time or not.
There are a bunch of core contributors that are very active. Several of their companies allow them to spend time contributing (most of them are from Japan). Another example is Aaron, he is on Ruby core and his company allows him some time to contribute to open source, so this is sponsorship in a way. I while a company stands to benefit from sponsoring a full time developer, they benefit just as much if someone else sponsors a full time developer. Right now there's not enough companies with either a business incentive, altruism, or interest. Perhaps there are companies out there interested in sponsoring full time devs who just don't know how (and to your point cannot find them) but I think that would be the minority case.
> Oracle Labs pays two people full time (including me) [to work on JRuby] and a PhD student, and more soon.
There's actually been a bunch of great performance increases in the past few years in addition to the GC. Optimized method cache invalidation by the late James Golick, frozen string pool for hash keys by tmm1, using vfork instead of fork, etc, copy on write GC. There's also new features, like Ruby's ability to GC symbols that let developers use symbols in more places and spend less time converting back and forth between strings.
So yes, money is a factor. Having employees work on it full time helps. Despite only having 3 full time employees Ruby has made some pretty impressive improvements recently.
Holy shit, how did I miss this?
edit: more details here: https://news.ycombinator.com/item?id=8804624
That's really sad...good guy
And embrace types, either optional or inferenced. Optimize on type information.
v8 doesn't really do multithreading, for example, so it's able to take a lot of shortcuts that save a lot of time. v8 doesn't have built-in Bignum support, and doesn't support non-UTF-16 character encodings, so there is a lot of work that it doesn't have to do, where more full-featured languages have to make sanity checks and do conversions under the hood frequently. Ruby's metaprogramming constructs are extremely powerful, but their rulesets are fairly complex, so the VM has a lot of bookkeeping to do to ensure that everything works well. I don't know Python's internals that well, but I do know that Ruby spends a lot of time making sure that all those things play nicely together without the programmer having to exert much effort, and that does incur a performance penalty.
v8 is an absolutely excellent piece of engineering, but it's a tool that occupies a slightly different problem space than the Ruby and Python VMs do. That's not to say that Ruby and Python don't have gains they can and should make, but that it's not as simple as pouring money in one end and watching performance come out.
If you need multiple, separate VMs (ie, to drive multiple isolated scripting engines in a game), then v8 isolates or Lua contexts will do you just fine, but that's a different use case than most places where you'd want to use concurrent Python.
https://docs.python.org/2/extending/embedding.html
It's just that at the time Python was created, multithreading was not something you generally did in a C program. The POSIX thread standard didn't come out until 1996. Python was started in 1989, first released in 1991, and reached 1.0 in 1994. Common rules of thumb for dealing with multithreaded programs (eg. "Avoid global or static data") didn't really become popularized until the 2000s, and many programmers in less well-informed circles still don't know them.
Not really. Basically on par.
Maybe my information is out of date, but I just now picked a random benchmark (the fasta one from the language shootout) and ran it against v8-3.14.5.10 and luajit-2.0.3 (Fedora 21, latest available via yum), and luajit came out quite a bit ahead. I grant that it's a really naive benchmark setup and shouldn't be taken seriously, but my investment here is pretty minimal :)
$ time luajit-2.0.3 fasta.lua 25000000 > /dev/null
luajit-2.0.3 fasta.lua 25000000 > /dev/null 8.60s user 0.01s system 100% cpu 8.613 total
$ time d8 --nodebugger fasta.js > /dev/null -- 25000000
d8 --nodebugger fasta.js -- 25000000 > /dev/null 12.80s user 0.51s system 100% cpu 13.209 total
Edit: https://gist.github.com/spion/3049314 is another microbenchmark where Luajit quite handily outperforms not just v8, but equivalent C code!Those are in the same order of magnitude even! Not even 2x the speed. Hardly makes a difference in picking one or the other, all other things being the same.
Now, compare V8 or LuaJit with Python or Ruby -- that would be "wrecking it".
I'm not attacking v8 here. I'm simply pointing out that there are VMs for interpreted languages which run faster than v8, which isn't to say "v8 is bad", but rather "different languages inherently have different performance profiles".
Though I must admit V8 could handle dictionaries a bit better - but at the moment it does not.
In my opinion, V8 would have been much more powerful if it was fully parameterizable in all the types that it supports, etc. In that case Python could simply be translated into V8 intermediate code (or javascript).
Personally I've been working with JavaScript since the mid 90's, and node.js since very early on. I like JS more than most as it's my favorite language, for all its' warts. I do feel that for most work loads a modular dynamic language implementation for a given task is a better idea than building a compiled version. If you need more performance from something, or it's something relatively obvious (A/V processing comes to mind) it's similar to a premature optimization to go compiled first these days.
Kind of more than you are/were looking for.. but you may want to look into IronRuby/IronPython ... it's a pretty nice way of working with the languages.
If you just write in Fortran, you don't have to have to debug two versions.
Or Decaf http://trydecaf.org/
Or http://rubyfiddle.com/ if you want to be literal about Ruby in the browser :-)
As a programming language, Ruby really pushes the limits of what we are used to.
What I want to say is: In many cases starting with a scripting language makes sense, just to get a working prototype quickly, but when you know beforehand that performance is critical, don't overestimate the ability to just patch it up later. The whole 'premature optimization' viewpoint should (if at all) only be applied up to a certain experience level, after which it's rather damaging than doing any good.
Ruby pushes what?! I bet you never used a commercial Smalltalk system.
Ruby is popular because it's powerful and expressive in a way that many other environments are not. 'Power' isn't everything - ability to easily use it is equally important. Hell, it's a direct spiritual descendent of Smalltalk to a greater degree than any other popular language!
The performance of Ruby caused by the extensiveness of its dynamism had definitely been one of the downsides of the language, and the fact that newer languages place more emphasis on that while learning from Ruby is great - it's how the field evolves!
Check also atom vs sublime text as desktop applications...
Javascript is fast enough,the problem is that Atom is basically using browser rendering for the ui,which isn't as fast as using c++ widgets.