Why Bundler 1.1 will be much faster
patshaughnessy.net
patshaughnessy.net
Also, why don't you do the recursive lookup internally instead of having bundler do several calls? It seems like performance would be strictly better because it saves network round trips.
I could definitely see future iterations going down more levels and caching the result as well. Right now each dependency lookup is not cached, so that might cause issues once 1.1 is actually released :)
awesome!
They did the simplest thing that could possibly work, and ended up delivering 65% solution; now they're iterating. The hard work was getting people to adopt Bundler in the first place.
I'd rather have waited an extra 25 seconds on every bundle install over the past year than being forced to manage gem dependencies manually for another year.
For a solution to the rubygems index download they needed the changes to the rubygems api, which they could probably not have affected before being a successful project.
Having said that, I cannot at this moment tell you how to take over a Ruby runtime with a malicious Marshal byte string.
I'd much rather be using JSON but I was told Marshal or plaintext...I'll go with Marshal. :/
Also, Marshal doesn't allow any kind of code to be included into the stream, so there is no ability for stream to perform remote code injection.
Marshal call back into Ruby for non-builtin types, but it does so by simply calling a method on the constant and passing either the raw Marshal data or a previous created object tree. This provides enough protection that there haven't been any reported cases of it being exploited and no know issues exist with it.
tl;dr: Why, oh why, Marshal, and not, say, JSON?
As for why not JSON, because there is no JSON parser as part of the standard library and rubygems needs to be extremely careful about what dependencies it has.
Basically, I'm not convinced Marshal is necessarily any more risky than something like YAML would be, even though it feels scarier. But I haven't done an extensive audit or anything — I just looked over the Marshal code a while back because I was curious what it was doing.
I would love to see a more apt-get like system where you cache things locally, but that has its own implications/difficulties as well.
Basically, it's hard to keep the time down from "gem push" to "gem install" if you go towards more distributed/delayed indexing systems...but I'm willing to compromise if it's way better.
Another option would be to query for all dependencies of a specific version of a gem. That would reduce the amount of data, but it would still mean more work for the server and it would produce a lot more data to cache (much of which would never be queried more than once)
With the currently applied method, only the data that's really required must be queried and it can easily be cached on the server side too.
See: http://stackoverflow.com/questions/3642085/make-bundler-use-...