Announcing Topaz: A New Ruby
docs.topazruby.com
docs.topazruby.com
$ time ruby -e "puts 'hello world'"
hello world
real 0m0.184s
user 0m0.079s
sys 0m0.092s
$ time ~/Downloads/topaz/bin/topaz -e "puts 'hello world'"
hello world
real 0m0.007s
user 0m0.002s
sys 0m0.004s $ ruby -v
ruby 1.9.3p194 (2012-04-20 revision 35410) [x86_64-darwin11.3.0]
$ ruby bench_neural_net.rb
ruby bench_neural_net.rb 17,74s user 0,02s system 99% cpu 17,771 total
$ bin/topaz bench_neural_net.rb
bin/topaz bench_neural_net.rb 3,43s user 0,03s system 99% cpu 3,466 totalhttp://anholt.net/compare-perf/
Example output:
+------------------------------------------------------------------------------+
| + x |
| + x |
| + + x x x |
| + ++ +++ + x xxx xx x |
|++ ++++++++++++++++++ x x x xxx xxx xxxxx |
|++ ++++++++++++++++++ +++ + ++ xxxxxxxxxx xxxxxxxxxxxxxxxx xx x xx|
| |______MA______| |________A________| |
+------------------------------------------------------------------------------+
N Min Max Median Avg Stddev
x 57 45.62364 46.437353 45.93506 45.951554 0.19060973
+ 57 44.785579 45.534727 45.042576 45.056702 0.16634531
Difference at 95.0% confidence
-0.894852 +/- 0.0656777
-1.94738% +/- 0.142928%
(Student's t, pooled s = 0.178889) $ time ruby -e "10000.times { puts 'hello world' }" > /dev/null
real 0m0.102s
user 0m0.096s
sys 0m0.005s
and $ time ./topaz -e "10000.times { puts 'hello world' }" > /dev/null
real 0m0.098s
user 0m0.071s
sys 0m0.026s
Any idea why I don't see such big difference? $ ruby --version
ruby 1.8.7 (2010-08-16 patchlevel 302) [i486-linux]So downloaded & built Cardinal (which went seamlessly however I did have Parrot already installed) then I did same benchmarks alongside ruby1.8 here:
$ time ruby -e "puts 'hello world'"
hello world
real 0m0.130s
user 0m0.049s
sys 0m0.071s
$ time parrot-cardinal -e "puts 'hello world'"
hello world
real 0m0.057s
user 0m0.037s
sys 0m0.019s
Very interesting because I thought Cardinal was supposed to be slow!I think more diverse benchmarks are required. And when time permitting I might add Topaz & ruby1.9 into the mix.
I wish more frontpages for these kinds of projects would do that.
I disagree. Zero benchmarks is definitely better than specious benchmarks.
As to your point, obviously no one should be making any decisions off of flawed benchmarks, but flawed benchmarks (not so far as outright lies, just flawed) at least give me an objective justification to investigate further.
Even some flawed benchmarks could help turn the initial tide of responses like "this is X written in Blub, it's bound to be better!" or "this is a faster X! Now everything will be twice as fast!" They're silly examples, but it seems like every time a new technology comes out, these are the kinds of knee-jerk, overly-optimistic reactions people tend to have.
"Specious" does mean, more or less, flawed.
Zero benchmarks are better than flawed benchmarks.
So you are saying you would prefer wrong information?
1 obsolete : showy 2: having deceptive attraction or allure 3: having a false look of truth or genuineness : sophistic <specious reasoning>
So instead of taking the common meaning of specious, a deceptively attractive benchmark we have to go to a less common use of specious in order to construct an specious argument about the proper use of specious? Nevermind that it is distracting from the main point of the discussion about benchmarks in the context of Topaz.
So I'd argue that amalog's definition IS the common one.
I would also direct you to the usage examples in your link. All of which use specious in a context that implies deception of outright falsity. I have never ever seen specious used synonymously with obsolescence.
But in the case of advertising a new library/project, which is arguably one of the main functions of the frontpage, flawed benchmarks (though not so far as outright lies) at least give me an objective reason to investigate further.
With no benchmarks, generally I'll open the page, mutter "that's nice" and move on with my business. I'd imagine I'm far from the only person who does that. Young projects don't help themselves when they don't effectively advertise themselves.
Looks promising for Topaz! https://gist.github.com/havenwood/4724778
Deleted comment
I don't understand why the lack of features would affect the time of "hello world."
time echo "hello world" hello world
real 0m0.000s
user 0m0.000s
sys 0m0.000s $ type /bin/echo
/bin/echo is /bin/echo
$ time /bin/echo "hello world"
hello world
real 0m0.009s
user 0m0.002s
sys 0m0.004s
$ type echo
echo is a shell builtin
$ time echo "hello world"
hello world
real 0m0.000s
user 0m0.000s
sys 0m0.000sIn general a benchmark is probably the worst metric you could ever use for deciding on an implementation. Unless the profit margin of your business is razor thin and dependant eeking out every last drop of performance, and even then most of those gains will be from extremely small sections of code that are probably best written in assembler by a programming God, and you should investigate FPGAs, ASICs, and other high performance solutions.
If your benchmark (infrastructure) involves a database (or anything that uses disks) that's probably going to be the problem long before the speed of your language / language implementation.
In all comparisons, you should remove confounding variables. Yes, you should benchmark something you actually care about, otherwise what's the point? That doesn't mean all other variables are immediately null and void. That's why I said said if your goal is to measure ruby execution time, you should remove startup time.
As for the practice of benchmarking in general, you're partially right. Micro-benchmarks are usually useless because they don't map to real work load. But profiling and speeding up small portions that are used heavily can have drastic improvements that in isolation seem small -- the death by a thousand cuts problem. Not all improvements come from isolated instances with very slow performance profiles.
This fallacy about DB access and not needing to optimize really needs to go away though. Even if 50% of your app is spent hitting DB, you have opportunity to speed up the other 50% and it's likely far easier. Ruby in particular is ripe for improvements on the CPU side. I managed to reduce my entire test suite time by 30% by speeding up psych. I managed to cut the number of servers I need in EC2 in half by switching from MRI, Pasenger, and resque to JRuby, TorqueBox, and Sidekiq. And I've managed to speed up my page rendering time anywhere from 8 - 40x by switching from haml to slim. None of these changes required modifications to my DB, none required me to write assembly, none required me to switch to custom-built hardware, and each helped reduce the expenses for my bootstrapped startup, while improving the overall experience for my customers.
What is the percentage increase in profitability yielded from these optimizations?
What was the percentage increase in profitability yielded from the last A/B test of your homepage CTA copy?
Better than that, this savings isn't one-time. It's recurring, as EC2 is recurring. But we've also reduced the expense growth curve (the savings wasn't linear), so we can continue to add customers for cheaper.
The A/B testing thing is a complete non sequitur. a) there's no reason you can't do both. b) most A/B testing yields modest improvements.
If it's the latter, I'd seriously consider colo as you can probably reduce costs by another 80%.
I was illustrating that there is real world gain to be had by doing something as simple as switching to a new Ruby or spending some time with a profiler. These weren't drastic code rewrites. They didn't require layers of caching or sharding of my database. I fail to see what's even contentious about this.
Additionally, the X time exceeds Y cost argument really only works when people are optimally efficient. Clearly those of us posting HN comments have holes in our schedules that might be able to be filled with something else.
It's overly simplistic to say the only option is to cache everything. Or that your DB is going to be your ultimate bottleneck, so the other N - 1 items are worth investigating.
And even in the link you supplied, the illustrative example is getting a 20 hour process down to 1 hour without speeding up the single task that takes 1 hour. It suggests there's an upper limit, not that because there is an upper limit you can't possibly do better than the status quo.
Amdahl's law advocates starting from the part that takes the most time. In a database application, it can be interpreted as either A) improving the connector or B) reducing the application's demand for database resources.
"Why did Ruby 1.9 bother with a new VM? Why try to improve GC? Why bother with invokedynamic? Why speed up JSON parsing? Why bother with speeding up YAML? Yet there's obviously value in improving all these areas and they speed up almost every Ruby app."
JSON parsing improves those applications that use JSON parsing, and in many applications JSON parsing is the main operation. There are many other applications for which garbage collection is the limiting factor. You are taking my comment, which was addressing the parent comment's remark that "This fallacy about DB access and not needing to optimize really needs to go away though.", way out of context. It's not a fallacy -- you need to know what is dominating execution time and how to improve that aspect.
Take it to the logical extreme -- you could just write in x86 assembly directly. The program would be faster than ruby, but the development time would not make assembly a worthwhile target.
"And even in the link you supplied"
What link did I supply? I recommend the Hennessy and Patterson "Computer Architecture" book :)
In any event, we probably agree on more than we disagree. I never disagreed with working on the DB if that's truly the bulk of your cost. But, you do actually need to measure that. It seems quite common nowadays to say "if you use a DB, that's where your cost is". And I routinely see this as an argument to justify practices that are almost certainly going to cause performance issues.
Put another way, I routinely see the argument put forth that the DB access is going to be the slowest part, so there's little need to reduce the other hotspots because you're just going to hit that wall anyway. And then the next logical argument is all you need is caching. The number of Rubyists I've encountered that know how to profile an app or have ever done so is alarmingly small. Which is fine, but you can't really argue about performance otherwise.
You can run into the same problem with MRI and its GC settings. If too low for your test, you're going to hit GC hard. It's best to normalize that out so you have an even comparison. Confounding variables and all that.
There's a lot you can do outside the Ruby container to influence startup time. When comparing two implementations, the defaults are certainly something to consider, but not when trying to see which actually executes Ruby faster. They are two different metrics of performance and should be compared in isolation.
Deleted comment
$ time ruby -e "puts 'hello world'"
hello world
real 0m0.011s
user 0m0.008s
sys 0m0.003s time ruby -e "puts 'hello world'"
hello world
real 0m0.221s
user 0m0.005s
sys 0m0.006s
subsequent times: time ruby -e "puts 'hello world'"
hello world
real 0m0.008s
user 0m0.005s
sys 0m0.003s
So, I guess he ran ruby first followed by topaz and ended up with those results $ time ruby -e "puts 'hello world'"
The program 'ruby' can be found in the following packages:
* ruby1.8
* ruby1.9.1
Ask your administrator to install one of them
real 0m0.060s
user 0m0.040s
sys 0m0.016s1. Pick one (just for this session): $ rbenv shell ruby1.9.1
2. And then run the example.
By the way 1.9.1 is really old already, 1.9.3 has a lot more bug fixes.
The Ruby language changed between 1.9 and 1.9.1, so a new package name had to be created.
If 1.9.1 was just called "1.9" it would break all of the packages in Debian that depend on whatever language features were different between 1.9 and 1.9.1.
"ruby1.9.1" in Debian 7.0 provides version 1.9.3.194.
It's "trivial" to make a fast language that looks quite a bit like Ruby. It is a lot harder that make a language that remains fast in the face of handling all the quirks of the full Ruby semantics, though, such as selectively handling the risk of someone going bananas with monkey-patching core classes that could happen at any "eval()" point.
2) Python is a very similar language to Ruby, and Pypy already runs python very rapidly. This doesn't guarantee that the Ruby interpreter will be anywhere near as fast, of course, but it does give evidence that it's possible.
3) As kingkilr notes below, they've taken into account your argument and they believe they've implemented enough of the language to be confident that they can run it rapidly. No reason you need to believe him, but it's worth listening to.
refs:
- http://topaz.sourceforge.net
- Topaz: Perl for the 22nd Century http://www.perl.com/pub/1999/09/topaz.html
- Historical Implementations/Perl6 http://www.perlfoundation.org/perl6/index.cgi?historical_imp...
Not criticizing at all by the way, just curious about the motivation / background / context for the project, which is missing from the docs.
a) Because it's fun
b) To prove RPython is a great platform
c) To mess with people's heads, it's crazy!
Congratulations for the release :-)
[0]:
$ ls -l playground/pypy
total 29160
drwxr-xr-x 12 lloeki staff 408 18 Feb 2011 ply-3.4
drwxr-xr-x 15 lloeki staff 510 3 Apr 2012 pyby
drwxr-xr-x 23 lloeki staff 782 27 Mar 2012 pypy-pypy-2346207d9946
drwxr-xr-x 18 lloeki staff 612 27 Mar 2012 pypy-tutorial
-rw-r--r-- 1 lloeki staff 14927806 27 Mar 2012 release-1.8.tar.bz2
drwxr-xr-x 10 lloeki staff 340 3 Apr 2012 unholy[0]: https://github.com/topazproject/topaz/blob/master/topaz/lexe...
1. (optional) install PyPy "JIT compiler" binaries [0]
2. get and extract the PyPy source[1] from bitbucket release packages, which includes the RPython translator, written itself in Python and, although working on CPython, best run under PyPy as installed on step 1.
3. write a (R)Python module including a def target returning the entry point function[2], and call the translator upon your module, like so:
python ./pypy/pypy/translator/goal/translate.py example2.py
I wish getting the translator was more 'packaged' and did not involve getting the whole PyPy source but I hear this is in the works, and it is reasonably easy already. The hardest part is actually figuring the translator is not part of PyPy binary release, and that it's available straight from bitbucket.[0]: http://pypy.org/download.html
[1]: https://bitbucket.org/pypy/pypy/get/release-1.9.tar.bz2
[2]: http://morepypy.blogspot.fr/2011/04/tutorial-writing-interpr...
Although the most well-known language implemented in the pypy "vm" is python, the toolchain is completely language agnostic.
So, this is more akin to Apple writing a C interpreter on top of llvm than it is to building a ruby interpreter on top of python.
* RPython comes with a good garbage collector
* the language where you specify what's going on is RPython in which you write an interpreter. Then you get a JIT using a few hints. This is difference than "interpreter in C + compiler to LLVM" scenario by quite a bit.
* RPython comes with a set of data structures that are higher level (lists, dicts, etc.) and is a GCed language. JIT is also well aware of those.
In much the same way, comparing "Apple writing a C interpreter on top of llvm" to "Alex Gaynor writing a Ruby interpreter on top of RPython" makes a very good analogy, even though Apple is not much like Alex Gaynor, Ruby is not much like C, and PyPy is not much like LLVM.
You're right that the analogy doesn't quite capture everything here. On the other hand the analogy gets across most of the point.
I would have said "That's mostly right, but the interesting thing about sharing code on a VM is access to VM features and code shared on the platform, not just that it's possible to decouple a front end from a VM."
or something to that effect.
"C'mon dude, don't let flaming beget flaming."
To be clear, this isn't "Ruby running on top of Python." This is "Ruby specified in RPython."
If Topaz was able to give me the ability to access NLTK, but still write in Ruby? I'd be overjoyed.
I doubt that's going to happen. Topaz is not a Ruby running on the CPython interpreter, the Topaz VM is developed in RPython. As far as I know, RPython doesn't provide that capability either (between VMs coded in RPython), which fijal seems to confirm.
Zach Raines did 99% of the heavy lifting (I've got a very small part to the whole project, and no time to work on it right now), and it works pretty nicely overall.
Especially when writing a code generator. Targeting a significant whitespace language is potentially more tricky than one with explicit delimiters. I love the Haskell approach to this.
Yes, I hate it that much, and I know quite a few Ruby devs with similar views...
Of course, I also use Coffeescript and Haml...
I'm curious: why not use Python then as well?
For e.g.
- Javascript | https://bitbucket.org/pypy/lang-js
- Smalltalk | https://bitbucket.org/pypy/lang-smalltalk
- Schema | https://bitbucket.org/pypy/lang-scheme
> To run Topaz directly on top of Python you can do:
> $ python -m topaz /path/to/file.rb
At least once a year there's a maelstrom of posts about a new Ruby implementation with stellar numbers. These numbers are usually based on very early experimental code, and they are rarely accompanied by information on compatibility. And of course we love to see crazy performance numbers, so many of us eat this stuff up.
No. It's Ruby in RPython, where PyPy is Python in RPython. The pypy project is currently in the process of splitting "RPython" (the VM-development framework) from PyPy (the Python VM) to make that clearer.
RPython is a general purpose, language-agnostic (ish?) core for implementing JIT-ed garbage-collected languages, it's kind-of similar to LLVM being a framework for implementing compilers.
(now because rpython is a proper subset of python you can run your rpython VM on a python interpreter without translation, but that doesn't give you any specific python bridge, in the same way coding your VM in C doesn't give you a c FFI for free)
The important distinction really is between RPython (a language and framework for building virtual machines) and PyPy (a virtual machine implemented in/on RPython)
The fact that it used the name "Topaz" should have absolutely no bearing on any current project.
I also dislike the "high performance" claims without showing a single performance comparison. Not to mention the state of implementation completeness of the language.
Additionally, erlang has avoided scaling issues by being functional. This might be good or bad, depending on your viewpoint, but just implementing a ruby interpreter in erlang won't cut, because you're not answering any of the hard questions about the state. You might answer "meh, ruby is just a broken language", but well, this has nothing to do with that article.
"Topaz - An implementation of the Ruby programming language, in Python, using the RPython VM toolchain." [2]
[1] http://docs.topazruby.com/en/latest/blog/announcing-topaz/ [2] https://github.com/topazproject/topaz/blob/master/README.rst
I don't believe introducing a lot more dependencies and complexities will cure any problem. Sure, it's a nice project and the work shows the great skills of the developer. But it's not practical to use this besides some non-crictical fun projects.
(Also, LLVM is not a language...)
As far as I understood is that this allows for low-level concepts to be changed more easily, while low-level implementations always suffer from basic decisions (e.g. you will never get reference-counting out of CPython, while you might be free in to change that in an RPython-implemented language).
So, the entire point of PyPy is being able to implement other high-level languages in RPython, so an attempt to implement Ruby is definitely very interesting! Also, the resulting binary is not depending on Python whatsoever.