I am very emphatically _not_ saying that you're wrong, but without enumerating what your problems are, you can't get them solved.
I am very emphatically _not_ saying that you're wrong, but without enumerating what your problems are, you can't get them solved.
* You can't override a dependency:
This may very well be an issue with rubygems, but I'd have hoped bundler adoption would have fixed it if that's the case. Since dependencies and their versions cannot be overridden, transitive dependency conflicts are a constant minefield. The naive solution is to modify your gemspec to be overly permissive and I get pressure to do this frequently with gems I work on. But, now I'm asserting that my gem works with some hypothetical dependency that hasn't been release yet. And I've watched that blow up many times. Semantic versioning does not fix this.
* Dependencies can disappear on you:
Yanked gems irk me, but I appreciate their value in preventing the installation of an unwanted version when one isn't specified. However, it also explicitly blocks installation of a particular version, effectively ignoring the version specification in Gemfile. I may be romanticizing things here, but I don't recall ever seeing a non-SNAPSHOT published artifact being removed when I was working with maven or ivy.
What's frustrating is the gem still exists, but not in the index. I guess I would expect Bundler to realize I do want to install the version I've specified and fetch the file and do the installation. Ultimate control in my dependency graph should rest with me, not the whim of an upstream developer. I realize this is a rubygems.org thing, but it also strikes me as a solvable problem that many don't see as a problem. And it means that checking out historical copies of my app almost certainly will not run without changing the Gemfile(.lock), which seems like a core value in a dependency management tool.
In any event, you will only discover this when deploying to a machine that has come up since the gem was yanked. Common staging server and CI strategies don't catch this since it's essentially a race condition in the gem ecosystem. You simply won't find out until your deploy fails. It undermines any trust in the ecosystem and the only viable solutions are: 1) run your own gem server; 2) modify gemcutter to dismiss all yanks, or 3) vendor every gem your app uses.
* It has weird rules:
gem 'x', platform: :mri
gem 'y', platform: :jruby
works the way you'd expect, but gem 'z', git: 'whatever', branch: 'mri_compat'
gem 'z', git: 'whatever', branch: 'jruby_compat'
will not. Apparently the dependency graph can't be resolved if the gem names aren't unique. The platform part isn't taken into consideration from what I gather. This made it really hard when I was trying to port some C ext. gems to a JRuby equivalent.* It promotes bad practices:
This one is admittedly contentious, but given bundler was created basically for Rails 3 and Rails is the worst abuser, it's also hard to divorce the two. But requiring your entire dependency graph up front is just not a good idea. It's bad for performance and it's bad for memory. It's why apps take 30s to boot. Most non-trivial Rails apps I've come across are basically multi-tiered applications in a monolithic codebase. That means fog is getting loaded in controllers and haml is getting loaded in Sidekiq jobs. It's just a very odd situation. I've tried to defer loading, but this is a battle that's hardly worth fighting because of railties. Auditing every gem to see if it has a railtie is tiresome and invalidated as soon as a new version comes out. If a gem has a railtie, it needs to be required at a very specific point in the Rails boot cycle, otherwise your app just won't work the same. Figuring out that difference will likely drive you insane.
Likewise, every new gem created from Bundler has a gemspec that shells out to git at least once. So, if you have a dependency on a gem sourced from git in your Gemfile, you get to pay that cost on every app load. And if you end up somehow getting different results on different machines, then you're not really running the same thing, which seems odd to me for a dependency management tool.
* It's slow:
Gem installation has gotten a lot faster since the early days. And dependency resolution has gotten better as well. So, I'm not saying no work is being done here, but it is still slow. This is a complaint that is always levied at maven, too, so Bundler certainly isn't unique in this regard and I'd argue it got faster in a much shorter timeframe than maven did, so that's promising.
* It was designed for MRI:
This one may be unfair since I don't use rbx, but bundler wasn't really built with JRuby in mind. "bundle exec" is used everywhere now and it forks a process. Forking the JVM is anything but light. The solution is to use binstubs, which avoid the forking. But also means you can't really use both JRuby and MRI in the same codebase.
I really hope that Cargo learns from some of the pitfalls of Bundler. It seemed like Bundler didn't really learn from the pitfalls of other package managers. I wasn't involved with any of the design decisions, so I certainly don't want to say it was pulled together haphazardly. But it also seemed to overlook problems that tools like maven had solved over the past decade. I'm sure a fair bit of that had to do with the underlying rubygems system, but the distinction is also a bit moot from the user perspective.
Bundler in general is a nice tool, and Bundler + shared pile of gems is a nicer workflow than gemsets (though I am not a fan of "remembered options"; I'd rather make my own config file than forget that Bundler "remembered" my experimenting).
The more I think about Bundler, the more I wish that general purpose tools like Nix Package Manager [2] got more love. I've been playing around with using Bundler for the dependency resolution step, then Nix for the packaging and deployment steps (each app might have a separate Nix profile where it could install its particular set of gem versions), but there are missing pieces.
[1] http://www.youtube.com/watch?v=3soqhbnh0jY#t=35m22s [2] http://nixos.org/nix/
Yup. This has gotten way faster, recent versions of Bundler include a -j. Also, the bundler-api project has been steadily working on making this faster. Without getting into the details, let's say the gems format and API does not make doing this efficiently easy.
> But this seems to trigger the (NP-complete, last I checked [1]) dependency resolver yet again,
It does not. It _does_ check to ensure that what's in the Gemfile matches what's in the Gemfile.lock, though.
I much prefer how npm handles things.
> I much prefer how npm handles things.
Can you elaborate on how npm handles things in a better way?
NPM also has the ability to have isolated dependencies from each other dependency. For example, module A can pin module B to version 2, whereas module C can pin module B to version 3. So you have independent copies.
Now, that worked for a dynamic language where you don't have to deal with static/dynamic linking. Would it be appropriate to statically link two modules that are the same, but at different versions? Maybe not. That would lead to massive binary sizes.
> I love local by default.
Roger. It's hard to articulate my own thoughts about this, but I feel like Bundler enables the best of both worlds here.[1] That said, it's obviously a preference.
> NPM also has the ability to have isolated dependencies from each other dependency.
Ahh yes, I've heard about this. In Ruby, it's not really feasible due to the language, and I'm a bit skeptical that it doesn't lead to super extra complexity, but then again, I haven't used it myself...
Totally agree regarding static/dynamic linking. Though isn't it the same thing with NPM: you still have two copies of the library.
1: http://words.steveklabnik.com/how-to-not-rely-on-rubygemsorg...
It's braindead simple. It's an emergent property of the way node loads modules, which is also pretty darn simple.
When you give Node a package to require, if it doesn't refer to a relative or absolute path, it walks up the directory structure looking for `node_modules`. When it finds one, it looks inside for the module you asked for. If it doesn't find it, it keeps walking up the directory tree.
So you can have multiple module repositories in a directory tree. A given file will always find the one that's "closest."
Npm takes advantage of this by installing every module's dependencies in a `node_modules` directory in the root directory of that module. Subdependencies' dependencies go in their node_modules directory. Voila! Isolated dependencies, braindead simple, albeit with a crazy deep directory structure. (node_modules/foo/node_modules/bar/node_modules/baz...)
Node/npm has the best dependency handling I've yet seen, and I think it's because it doesn't try to outsmart me or go for anachronistic disk space optimizations.
Consider this: A frequently used module in one of my trees is "glob". It's repeated four times in two different versions. It takes 207K each time, including subdependencies. The wasted space is 414K... more than a whole floppy! ;-)
Or to put it another way, I've wasted 0.00017% of my rather small 250GB SSD.
Optimizing that is anachronistic.
> that's nothing that can't be solved via hardlink deduplication
Yep, and `npm dedupe` [1] does something similar. This does have the potential to become massively complex, though. You have to re-dupe when one dependency upgrades but another doesn't, and you also need to deal with the fact that two modules' dependencies may be sharing memory space that a module author was expecting to have to herself. (Modules are cached, so changes to "module global" variables are shared, but modules loaded from different locations are cached independently.)
On the other hand, my Java .m2 directory counts 2469 jar files for a total of 2.4G. My considerably smaller .cabal is still 253MB. I won't complain if hardlinking prevents package repo sizes from going out of hands.
> Yep, and `npm dedupe` [1] does something similar.
Not quite. From what I understand, it attempts to manipulate the package hierarchy. On the other hand, hard link deduplication doesn't need to do any such thing, it simply needs to hardlink files with the same checksum. No intelligence required. Some time ago, I used rsync to do this kind of thing when deploying large binary dependencies, and it is very practical (IMHO, hard links get too little love).
If you were trying to namespace each of those symbols as you went because of a global symbol table (>_> ... c), then this would be a much more complex task.
It's important to realize that rust libraries are c libraries, complete with symbol information; Rust does mangle symbol names per crate by default, but I'm not 100% sure would stop you from having some issues where the public symbols in a crate caused global symbol table conflicts.
Don't know though, just speculating.
That seems like a bad idea even in a dynamic language, perhaps especially in a dynamic language. What if you return a type from module A that is defined in module BV2 to module C suddenly it can't find the method(or the method does something different then expected). At least in a static language you could prevent this to some extent. It seems like it would be better to have version ranges and fetch the newest released version or fail if no version is acceptable with interactive/cmdline overrides.
One way this manifests itself in a nice way is with local installs - let's say you have a package that depends on modules A, B, and C. Each of those modules has its own dependencies. The default behavior when you install the package is that A, B, and C have copies of all their dependencies (so if all 3 depend on D, you get 3 copies of D), and so forth all the way down. This does waste disk space, but for the most part node modules are small and disk space is cheap. The advantage is that you don't have to worry about conflicting dependencies, or subtle changes breaking things in hard-to-find ways.
Like a sibling commenter said, I have no idea if this model will transfer well to Rust.
For one of my larger Ruby apps I resorted to saving a copy of $LOAD_PATH before initializing Bundler, and then using some hacks to reset $LOAD_PATH to the basic load path + the results of grep'ing the "post Bundler" load path for each of the relevant gems and its dependencies, so that each require only worked on the minimal required $LOAD_PATH. The result cut down the number of stat calls from 100k+ on startup to <20k without any other changes, and cut tens of seconds of the startup time....
The stuff I did is a giant hack, but frankly Bundler needs serious work there - for almost all Ruby code I've worked on, poor load path handling in Rubygems and Bundler accounts for 90%+ of startup time.
I don't think Bundler and Rubygems' path handling would ever have been written and released in the state they are now if someone had seriously looked at the actual system calls it generates for any app with dependencies on more than 2-3 gems....
Especially because there's a simple-ish solution: create manifests of files for each gem, load them, and check a hash of the combined set of files first, and fall back to a much trimmed down load path if the file isn't from any of the specified gems. Bundler could even easily create an aggregate manifest of all the gems the app depends on to make the check very cheap.
I keep meaning to try to hack something together, but for now I have something half-assed that cuts enough seconds of my app start times that I haven't been able to prioritise it.