Making ActiveRecord 2x faster
tenderlovemaking.com
tenderlovemaking.com
It's been very discouraging to me to see responses to performance issues that range from lukewarm to plain "I don't care". The Rails community could benefit a lot from a visible push by the core team to make performance a priority.
That said, I will still probably end up migrating toward Clojure/cljs and js-focused apps...
EDIT: It's a waste to bleed a client if all your blood-money is going to Heroku! XD
It's treated like a hobgoblin.
I've tried to get into the Java web world, but it's always felt like the tools weren't what I was looking for. Partly, it always feels like the Java world builds so many ways to do things that it's hard to get a feeling on what is canonical. Even with something like Play, should I use Java or Scala? If I use Java, should I use Ebean or Hibernate? If I use Hibernate, should I use XML or. . .
In some ways, it feels like Java advocacy is focused on non-coding managers. Or maybe I just don't know where to look for developer-focused advocacy.
For example, in the C# world, Microsoft has done a decent job telling developers how to get started and why they should care about .NET MVC. In the Java world, it's hard to know where to get started: Spring, Play, Ninja, Wicket, Dropwizard, Tapestry, Wicket, etc.
I guess I'd love to hear your recommendations and experience in moving to Java.
So yeah I'd say I'm pretty early in my endeavors (I mean I code using a Java tool for a corporation but there is a pretty strict workflow for that) but as far as open source & my actual education is concerned, ok I'll give you tech advice --
short answer: https://spring.io/guides
1. Java has a number of MVC solutions. Pick one & study it, probably with a tutorial/guide. Don't focus on the amount of configuration available (the cruft), just follow a tutorial to get a few hello world kindof pages. You'll see that it's not nearly as difficult as one might imagine. I first studied JSF, then Spring MVC (in retrospect, Spring MVC is probably a better starting point. I think it is also really hot in the job market.)
2. Pay attention to the terminology (research every unknown acronym) and see how it is an analogue to something you're familiar with.
3. Make note of the features/points that are not available in the language/framework you currently use. For me, data-binding to a form was the first thing where I thought "woah thats so much simpler...". Look up PrimeFaces or something like that if you want to see an explicit/obvious example of this, though personally I'm not sure I'd use it in practice (Spring itself offers data-binding stuff).
4. Research some JVM framework that is similar to one you are already familiar with (Play is a good start here or again Spring MVC if you're not afraid of Java syntax) and figure out how to make some basic pages, then add some logic.
5. Study the relationships between servlets/classes/beans/etc, figure out how they are scoped. A lot of it is pretty much free-form code and frameworks like Spring will let you do pretty transparent things before you get deeper into more nuanced app-flows.
6. Protocol is usually pretty easily determined by app needs. I like to default to json since I do a lot of javascript and it seems to fit better with REST, but internally lots of java apps & web services use XML. This hasn't really hurt my brain at any point yet.
7. And finally, when in doubt, try a pure Java solution until you realize that you need a tool. I think this has really been key for me. I could build apps on a similar level to what I was doing in RoR without many fancy libraries, and now I'm learning enough of each that I can make a more educated decision when the application actually demands it. I think all the choices can be daunting but the power of the minimal tools is enough that the entire chain isn't necessary for every application.
8. Oh, plus one more! If you like some other language (for me it's Clojure) study the language in it's pure form plus throw in a little study of Java interoperability in case you'd like to use it sparsely within one of the tried & true Java frameworks. This may not be possible without some meditation on it...
There are lots of templating languages available so as always you'll often just be injecting stuff into html. Many frameworks are just turning into a home base to shove fancy javascript frameworks into so depending on your app, you may not even need much server-side code (in fact I think this is why some of the simplistic monolith frameworks are surviving so long).
And of course, I don't know everything.... so maybe the JVM isn't the answer. I'm just having a lot of fun learning it, and I was once a Rails guy.
I still really like Ruby for its expressiveness and malleability, and don't plan to move away from it any time soon, but a lot of that means digging into these libraries, finding problems, and fixing them. I really enjoy that process, but it's rather inefficient for me to be responsible for the performance of every library in my stack.
Ruby has a reputation for being slow, and while it's true that you're not likely to ever reach statically-typed language speeds, it can be pretty decently fast. The danger with tons of abstraction is that a performance hotspot buried three libraries down can turn your otherwise-fast application into a tarpit. If Ruby library developers made performance analysis and optimization a goal alongside "beautiful code" and "DRY", Rails applications across the board would get substantially faster without changes to the underlying VM.
For example, Sprockets uses Pathname heavily. Ruby's stdlib Pathname is extremely robust and extremely slow. This slows down path resolution for every single asset quite a bit. My custom-patched Sprockets does asset resolution and compilation around 43% faster than stock Sprockets. As you can imagine, this has tangible benefits in the speed of my development cycle, and substantially increases my happiness. I opened an issue on the Sprockets project, but it was closed as requiring too-substantial a refactor. Imagine how many developer-hours would be saved if the Sprockets maintainers were to fix that hotspot?
but it's trueeee I prefer scripting languages to Java but there are a few of them available on the JVM. To be honest, I think the right balance is in the dynamic-OPTIONALLY-static typed languages.
Java interoperability solves this in a way... I've been taking a look at Dart for similar reasons. I just don't think the answer for me is RoR. Once I learned a Lisp (Clojure) and a lot of javascript I realized a lot of the syntax features I like about Ruby exist in many other places
What RoR really has going for it is the gem community, but if they're hard to trust and come with a slew of other problems...? Tough choices
Edit: if you just want to install a gem and skip the why, the parent has this blog post: https://www.coffeepowered.net/2014/02/09/faster-pathname/
Web services that DO use Ruby today are almost always IO bound. For web services that don't require data crunching CPU bound processes, Ruby as a language is hardly the bottleneck.
While faster view renders might not improve your theoretical maximum throughput, it will certainly have a positive impact on your users' experience.
Not an afterthought, but I'm more interested in providing features for my customers. The site is fast enough, and the time to develop new features is more important than making the thing scale for zillions of users when I only have 100's hitting the web site.
I don't ignore performance, but if you're charging enough, it's certainly not the primary consideration.
Ideally, in the Rails world, the people building client apps would only have to worry about the algorithmic complexity of their code, and the libraries should be operating efficiently under the covers. To that end, you're completely right in that you shouldn't have to worry about performance - the people writing the software you build on should have already have done that. But when they don't, that costs you time and money, and I think that's a shame.
The same concept applies to gems. You just construct a scenario that stresses your gem internals, run it with ruby-prof, pop open the resulting dump in kcachegrind or similar, and explore to find the hotspots. I did something like this for MongoMapper, and documented the process here: https://www.coffeepowered.net/2013/07/29/mongomapper-perform...
Cf. the comical performance of the entire rubygems ecosystem (no least bundler).
A programming language community can hardly make a stronger statement about their general mindset wrt performance than tolerating such an unmitigated trainwreck at the very core of their ecosystem and _every single developers_ daily workflow.
If you could magically drop support for Ruby 1.8, and make every library use require_relative to load its internals rather than require, I'm pretty sure most of the problems with Bundler would disappear.
Gem developers: If you aren't supporting 1.8, switch to exclusively using require_relative to load your app internals. Consider separate 1.8/post-1.8 code blocks for requiring code. If you'd like to understand why, strace a bundler-backed app load sometime.
Sorry but come on, these apologisms are part of the reason why these things never get fixed. Ruby 1.8 went end-of-life two years ago. And which precious ambiguousity are we enshrining here that could possibly justify this magnitude of wasted developer lifetime?
There is no excuse. If anyone in charge gave a flying f*ck and was not a complete code illiterate we'd have a --fast flag by now. In fact we'd probably not even have it anymore because it would have become the de facto default over night...
The issue is the require means "search the load paths for files matching this filename, filename.rb, filename.so, then search relative to the requiring file for the same if nothing was found". Bundler manages loads by finding all the gems in the bundle (which have been pre-resolved and locked), then creating a load path which satisfies the dependency order. So, when you `require "foo"`, the (extensively large) load path is iterated looking for a match. Given the number of gems (and thus the number of items in the load path) in a typical project, it's not difficult to recognize why this might be slow. Over the course of a load, you will perform hundreds or thousands of require invocations, each of which has to perform this iterative search.
This seems unnecessary, except you have to know how to resolve ambiguities. Consider something like this file: https://github.com/rails/docrails/blob/master/activesupport/...
What does that require statement mean? It's a require statement that asks Ruby to require something called 'benchmark', from within a directory that contains a 'benchmark' file. There are, in fact, multiple matches for "benchmark" in the load path. In order to ensure that the right one is loaded, you have to perform that iterative search, in order, to make sure that you always arrive at the stdlib "benchmark" before the ActiveSupport core_ext "benchmark". Yet, in the majority of cases, by convention, when you require some file in a library, you're actually wanting to require some file relative to your requiring file. It's obviously inefficient to iterate the entire load path looking for some file that you know lives in a subdirectory relative to the requiring file, but since Ruby doesn't have any way of knowing whether you mean "load the first hit of this file from the load path" or "load the file relative to the requiring file", it has to perform the slower, more conservative search.
This is what require_relative is for. That lets you tell Ruby "Hey, skip all the load path shenanigans and just load this file relative to the requiring file". Since you don't have to do the extensive load path iteration, the search time is constant, rather than linear.
I've experimented with a few "--fast" systems, including pre-emptive load path iteration and caching (doesn't really make a difference; you aren't going to beat stat by any significant margin), and load path "baking" (for each file loaded in the manifest, keep the resolved absolute path and replace relative paths with absolute paths"), but I've yet to come up with a solution that properly replicates Ruby's search-and-load behavior, and invariably ends up breaking on some edge case. Fuzzy semantics are kind of a bitch like that.
Yes. I haven't interacted with the bundler team but I did try to contribute to rubygems during their security debacle. I hope these are not the same people.
but I've yet to come up with a solution that properly replicates Ruby's search-and-load behavior, and invariably ends up breaking on some edge case.
I have a hard time believing that'd be your only bottleneck?
If that's really all then do exactly what you said. Break the edge case, replace your iterative search with the O(1) lookup. There you have your --fast flag. Document that require_relative is to be used for relative imports for the flag to work.
Then count the days until someone patches up a 'fast-rails'.
And to the gem that auto-converts projects into --fast compatibility.
Hadn't personally heard of this before.
I'm glad to see the gradual progression since 3.2 stable!
For those interested in good caching without bypassing parts of ActiveRecord, Memcache + Dalli (gem) is a great solution. Caching is not always the right solution, though. Sometimes the best choice is reversing how you get to that data or optimizing your query directly using SQL.
They are only caching calculations internal to AR. So you don't need to worry about using cached data when the actual external data has changed, is one thing.
The caching here is only of things that can't possibly change, essentially the output of a function for a given input.
Your memcache suggestions are a different sort of thing entirely, which certainly may be helpful for some apps in some circumstances.
It's like
cloudflare/varnish/nginx -> rack -> rails -> app
-> memcache
Anything you can use to take load off the DB, which in turn may hit some sort of blocking I/O call (SSD or HDD) is a very Good Thing (TM). Or a sleeping heroku dyno. ;)Unless I'm misunderstanding.
Not asking for the same thing twice (or N-1 or N^2-1 extra times) if has to go over the network is a good thing.
Databases often aren't properly configured OOTB, especially when it comes to temporary tablespace and query caching.
It might. But I definitely wouldn't assume it does. Micro-optimization.
Of course, it doesn't hurt either way, unless it does in code complexity, or opportunity cost.
e.g. `Article.where(xxx: value)` or `Article.where(xxx: value).last(3)` will not benefit from this improvement yet.
statement = undocumented_method_call { |params|
Article.where(xxx: params[:xxx])
}
statement.execute xxx: "some value"
https://github.com/rails/rails/blob/e5e440f477a0b5e06b008ee7...The trick is that the cache only supports caching ActiveRecord::Relation objects, and not everyone knows which calls will return a Relation object. For now, I want to hide this API and make it as "automatic" as possible so that people don't have to change app code, but still get performance boosts.
1. http://edgeguides.rubyonrails.org/upgrading_ruby_on_rails.ht...
That one is staying, most of the others are going (e.g. find_all_by_xxx, find_last_by_xxx)