1) We couldn't be sure that we hadn't broken something because we knew we didn't have 100% test coverage.
2) Some of the gems we depended on conflicted so severely that we had to rip them out and implement the solution ourselves or pick a different gem.
3) Our tests themselves of course contained code with breaking API changes. That means we had to maintain the tests as well as the production code through the upgrade, and had to make changes to many of those tests.
All of this uncertainty means that this was not your typical test-driven confident refactor. QA still had to do massive regression testing, and we're pretty sure we introduced at least a new bug or two. It took a pair of devs 4 months of non-stop work to get these services up to Rails 4.1. The upgrade was absolutely necessary as the Rails core team had already stopped fixing major security holes in 3.1 long ago. The company incurred a tremendous cost during this upgrade process. If they would have kept things up to date all along they could've saved money, but of course that would have eaten into the supposed time-savings of Rails.