Ruby's Date/DateTime classes rewritten in C.. 20-200x perf improvement
github.com
github.com
+1 for duck typing, a custom non-Date object can look like a date for operations that only use a subset of the methods.
If you know your format (iso8601, etc), you're better off just using the right parser instead of Date.parse.
Also the internal representation uses Rational, which can be pretty slow itself.
Date.parse: http://github.com/ruby/ruby/blob/trunk/lib/date/format.rb#L1...
edit: correction to how parse works.
I don't know why their version is so slow, I can only imagine that the conversion to Astronomical Julian Day (the Ruby Date internal representation) is expensive, inefficient or both.
I used year/month/day ints as internal representation, which was fast for my use cases (parse date string, format date string, compare two dates).
Dropping to C at the slightest hint of a performance issue is a band-aid, papering over deficiencies in the VM.
And one of the major purposes of using C is to write fast libraries that can then be called from high level languages.
[1] http://www.haskell.org/haskellwiki/Regular_expressions#regex...
Internally, things like Data.Vector work just like the C equivalent. The internal representation of the data is the same (i.e., Vector.Unboxed Double == a big block of memory with CDoubles in it), and the machine instructions that run on the data are the same.
High-level languages need not be slow. It's just that Perl, Python, Ruby, and PHP are.
[Edit] And just to be completely clear and also reap all the downvotes that I can get from disgruntled web devs: Python and Ruby are indeed totally unsuitable for writing CPU intensive, complex algorithms. If you write Ruby or Python code, avoid loops! Try to be declarative so all the loops stay in the C code.
Languages in which CPU-intensive work is fairly expensive are the ones which benefits the most from a C-rewrite. Ruby is notoriously slow in that regard, which makes it a prime candidate for such a rewrite. And before people scold me: Synthetic benchmarks, like the alioth shootout:
http://shootout.alioth.debian.org/u64/which-programming-lang...
is a good indicator of exactly the CPU-intensive power of a given language. Note however, that CPU-usage is not all constraints of a modern program. Which is, as an aside, why Ruby and PHP are viable options for writing code in.
Most Ruby developers will never use C unless they really care about extracting every last drip of performance from something. In most production deployments I've seen, the only gem that is compiled from C is the mysql one, for example.
Except having lots of legacy Ruby libraries written in C ignores the real problem and prevents the Ruby VM from evolving.
For this reason alone Ruby will never have the capabilities of the JVM ... which has it all ... kernel threads (no GIL), async IO and userspace threads (through libraries like Kilim) ... making all kinds of parallelism / concurrency paradigms feasible (in some tests Kilim scales better than Erlang).
For most web apps this doesn't matter because the heavy processing is done by the DB, but have you ever seen a usable DB written in Ruby?
Here's a freakishly scalable one written in Java with no C libs dependencies ... http://neo4j.org/
I wish people would start supporting projects like Rubinius more ... who's main bottleneck is the limited support it provides for C extensions, because for speed improvements it pays better long-term to invest in the VM rather then improving bottlenecks by dropping to C (which IMHO actively hurts the platform).
Dropping to C should only be done for code reuse.
Ruby 1.9 has 90% of the features you mention above and RBX will soon remove the GIL on top of it. Async IO is arguably further along in Ruby then in Java.
Ruby will never be as fast as Java, but micro optimizations like the OP will not hold back the inevitable progress of ruby vms.
You compare a many thousand man years VM (HotSpot JVM)
to one that has been largely written by one person in
isolation ( Ruby 1.8 series ).
LuaJIT was also written by "one person in isolation" and yet it beats Java in some benchmarks.There are techniques first researched in the Smalltalk/Self implementations that have been known for at least 20 years (roughly the same time Ruby was born) that could've been used in Ruby 1.9.
Ruby 1.9 has 90% of the features you mention above and
RBX will soon remove the GIL on top of it.
You're talking as if there's a list of checkboxes that just needs to be checked.It doesn't work that way.
One reason the GIL is here to stay is because of the many libraries written in C that depend on it.
Another reason would be performance degradation on single threaded programs ... removing the GIL on the whole needs major architectural changes ... just putting fine-grained locks on all mutable data structures won't cut it.
And yet another would be the garbage collector which also needs to be optimized for true multi-threading, otherwise it becomes a bottleneck. So add this on your list ... Ruby also needs a top-notch generational garbage-collector.
Async IO is arguably further along in Ruby then in Java
You are probably kidding.The solution seems obvious. Wrap them in a mutex by default and introducing an optional API call to remove the lock.
This way the important libraries will be fixed up gradually by the community and after a bit of time you won't have any C ext's left that will make use of the mutex and therefore the GIL will be gone.
That's the approach that rubinius is set to take.
Useful libraries written in other languages (C in this case) can benefit the platform (Ruby). The most common case for this is for speed benefits, but there are other possible goals as well.
To paraphrase the JRuby team "We write Java so you don't have to".
Finally, Rubinius is architected totally differently than YARV. Rubinius uses a JIT, YARV is just an interpreter.
It's funny, for some language communities (i.e. Java) being "self-hosted" is a big deal. I can't speak for Ruby but in the Python world, we simply don't consider it a major issue.