How to write fast code in Ruby on Rails
engineering.shopify.com
engineering.shopify.com
You can get 100x speedups, but the downside is that you wind up with big nasty SQL queries that duplicate Ruby logic and are hard to maintain. There was a nice gem that would automatically produce JSON-generating SQL [0], but it is abandoned now. It only supports Rails 4 and ActiveModelSerializers 0.8, which are both quite old. I just published a similar gem [1] that works for Rails 5 and AMS 0.10. Unlike the old gem, mine outputs JSON:API (common in Ember projects). I hope to add the traditional JSON format too. AMS 0.10 makes this easier, since it introduced the concept of "adapters". My gem is super new and doesn't yet cover a lot of serializer functionality, but I'm hopeful about supporting more & more. Feedback is welcome!
On the writing side, I've found activerecord-import [2] to be very useful. It batches up INSERTs and UPDATEs and for Postgres it even knows how to use INSERT ON CONFLICT DO UPDATE.
I also have a handful of Postgres extensions that use arrays to do fast math/stats calculations. [3, 4, 5] If you are thinking about moving something into C, it's natural to consider a gem with native extensions, but perhaps a Postgres C extension will be even better.
[0] https://github.com/DavyJonesLocker/postgres_ext-serializers
[1] https://github.com/pjungwir/active_model_serializers_pg
[2] https://github.com/zdennis/activerecord-import
[3] https://github.com/pjungwir/aggs_for_arrays
I also recommend using Oj [1] which is a very fast JSON parser/generator for Ruby.
JSON generating through DB seems like a good idea, but imo it's a bit too complicated solution.
[0] https://github.com/procore/blueprinter [1] https://github.com/ohler55/oj
Instead of doing
Rails.cache.fetch("blog_#{blog.id}_posts_#{posts.max(:updated_at)}") do
blog.posts.as_json
end
You can (and should) actually use something like Rails.cache.fetch(Post.all) do
blog.posts.as_json
end
Or even specific scopes, e.g. Rails.cache.fetch(blog.posts.published) do
blog.posts.published.as_json
end
The cache key will include the query hash, item count in the results and the max updated_at automatically for you.In some rare cases, you might want more control and then probably best to use something like `ActiveSupport::Cache.expand_cache_key`[0]
e.g.
key = ActiveSupport::Cache.expand_cache_key(["blog posts", blog.posts])
Rails.cache.fetch(key) do ... end
[0] https://api.rubyonrails.org/classes/ActiveSupport/Cache.html...That said, I wouldn’t use a query as a cache key since there may be more than one thing I want to cache about that query (GP’s example is their as_json representation, but what if I wanted something different to be computed? Like maybe html snippets? I would expect the cache key to mention everything about the thing it’s caching.)
This is exactly where you'd use something like ActiveSupport::Cache.expand_cache_key(["blog posts", blog.posts]). You can add bits of info to your cache keys to distinguish them from one another.
It's probably also worth mentioning that the `cache` view helper does parts of it for you already, and takes into account the digest of the view code itself among other things. Rails provides a really neat abstraction here that does make caching easier.
It has some rough edges, and some gotchas, but overall it's pretty smart.
Don't forget that even fetching the cache data itself takes time, even with fast cache storage like redis, especially over the network (where some cache layers live when they are shared).
I wrote a post how why one should still learn Ruby - https://bilalbudhani.com/why-you-should-learn-ruby-regardles...
There is so much code that is not part of the codebase driving the behavior of the system — and that third party code is very tightly coupled to the actual behavior that will be observed by running the code. Figuring how anything actually works is made massively harder than required and developers are encouraged towards designs that will require ridiculously inefficient database interactions.
Rails is terrible I think for experienced developers because there are no mechanisms in the codebase — instead there are layers and layers of “conventions” which often only exist to try to avoid bugs that would’ve been much better to catch or prevent with a combination of a type system, high level tests, and less mutation.
A big point of using a framework is to leverage an ecosystem other people's libraries (and stackoverflow questions!), and they will typically only work well if your code is relatively "normal".
If your code looks like "beginner" code in the framework, you're doing it right.
As a point of comparison, I switched to Django for a few years after a few years of Rails. Django was a lot less magical and was more debuggable. But Rails' magic actually encourages a better project structure. Django doesn't care about your structure.
I also found the thoroughness and completeness of the ecosystem surrounding Rails to be stronger. That is the real selling point to me. For most things you want to do in Rails, you will find a library that integrates with the project and DB structure you already have.
> avoid doing unusual things
If by "unusual" you mean "not the Rails way", then this makes sense. But the problem is, the Rails way doesn't always line up with the way an experienced programmer might want to do something.
In the specific example of fetching data from a DB - an experienced programmer might intuitively avoid multiple round-trips to the DB, and want to find the appropriate hook point(s) in a request handler to optimise this. Making 20 requests for 20 models is "unusual".
Rails doesn't make things like this easy. It's quite opinionated on how you should work. The core assumption in play here is that databases will either be fast, or can be made faster, essentially externalising the problem. Or if you want to address this - find a Gem to do it, or dive into the layered guts of it to figure it out.
Related - if I recall correctly, foreign keys were not deemed an important enough feature for inclusion in Active Record, and the Rails way is to embed the logic for joins into your models rather than the database. I'm sure many of us here could name several scenarios where this is just plain wrong.
Eg. You have a dot.net service that needs to write some data to the database, and doing it via REST endpoints on your rails app is really slow compared to accessing the database directly. But now you don't have any constraints on your relationships, and it's easy to shoot yourself in the foot. Extra sad, since you do have some constraints (belongs_to etc) - they're just not communicated to your database.
I have some legacy rails projects I maintain, and my job would've been easier if previous developers followed the rails way more closely - and whenever I manage to refractor my way back, code becomes shorter, clearer, more efficient and less error prone.
Things like only half of a relationship being defined, and joins being hand-rolled in one direction is a "favorite"...
Rails, when learned properly, enables complete and total dominion over the entire sphere of web development. Nothing else compares, nothing else even comes close.
Anything you'd ever want to do with the web, save the kind of scaling that led Twitter to replatform 5 years in, can be done 10x faster and 10x more reliably with Rails. The main bottleneck is understanding. There's a zen to Rails that you have to appreciate before you can unlock its potential.
It greatly saddens me that Ruby and Rails has been falling out of favor.
I didn't know if there were other architectures to consider for rails -- like having a few read replicas for web traffic, and then a single large writable master for sidekiq/background jobs, for instance.
https://www.citusdata.com/blog/2019/05/23/rails-6-multiple-d...
Rails 6 introduced multiple databases natively in activerecord. You can more easily have read only databases or segment off high write workloads.
Foreign tables are a way to allow bringing in data from multiple databases into one so it “looks” like everything is in a single db. This got better/faster, particularly with joins in postgres 12.
Also consider things like materialized views or even plain views which can help bring together just the right data you need ahead of time.
Just a couple options off the top of my head.
And yeah I can probably lean on views a bit more to save some search time for point/line-in-polygon operations.
I think a good first step would be bringing in a PostGIS consultant who can give you some expert advice on tuning for your current workload.
I'm assuming Django also faces similar issues
I have developed in Django and avoided RoR due to this possibility, but am always interested in learning more, provided the tools have a decent shelf life...
[0] https://arstechnica.com/information-technology/2012/03/hacke...
I loove ruby, it's my first choice for scripting wherever performance doesn't matter. No way around the fact that it very slow though.
It's also very easy to scale horizontally the execution of the web code. That's why, for most web software, programmer productivity is more important than raw language performance
You're right about price though; if the cost of compute is relevant in your business model - which, again, usually isn't the case - then doing that computation with Ruby on AWS isn't the best choice.
Surprise bills from AWS means it is time for refactoring ( and thus optimization time ;)
Plus if it is performant enough for Shopify and Github it seems like it would be performant for 90%+ of use cases? The speed of development and the flexibility of development is much greater than most stacks once you add rails.
Even more of the article has to do with Rails, and not ruby in general.
But the larger point, are you suggesting that in a language that is "performant enough", developers need no performance advice, they can write however they want and get adequate performance, for any performance needs?
Maybe, although I'm dubious. The existence of advice for programming does not mean a language/platform isn't "something enough" until we have AI writing code. At any rate, if such a language exists it is not ruby. But I disagree with that suggestion. The existence of optimization advice does not mean that a language "is not performant enough". If Shoppify is still happily using ruby/rails, I would say it indeed demonstrates they are performant enough for them.
Also describing common performance pitfalls does not imply that something is or isn’t “performant enough”. Only that you can make mistakes that will affect performance.
That, I’m afraid, is true for every language or framework.
It's very misleading for large companies to say "Ruby on Rails is performant enough", as they rely on other languages for their performance critical workflows.
I remember sometime ago on HN a GitLab developer repeating the same sentence. Actually, GitLab does use Golang for some microservices.
It's wonderful that a combination of Ruby and <performant_language> is easy to develop/maintain and fast at the same time, however, the distinction with a pure Ruby system is not a trivial semantic matter, as it perpetuates the wrong and misleading idea that pure Ruby is fast enough to develop a complex system.
https://www.slideshare.net/RedisLabs/redisday-london-2018-ho...
That said, the difference isn’t enough to matter. Both languages are in the same performance magnitude and thus suited for the same kind of tasks.
But then you have JIT based ones as well, TruffleRuby, RubyMotion, Rubinius.
Ruby
IO.foreach('logs1.txt') {|x| puts x if /\b\w{15}\b/.match? x }
Python with open('logs1.txt', 'r') as fh:
for line in fh:
if re.search(r'\b\w{15}\b', line): print(line, end='')
Ruby's best was 3.3 seconds whilst Python's was 3.9 seconds. If you factor-out Ruby taking 100ms longer than Python to startup Ruby is 18% faster than Python for string parsing. Results probably differ for Python's numeric performance but that depends whether you're including C-based third party libs such as numpy.For everything else, it's great, and it's almost always a good default to choose speed of development and engineer productivity over runtime performance.
Doesn't Elixir on Phoenix[1] provide the elegance of Ruby on Rails without its drawbacks?
Have you seen anything different? Most enthusiastic vendors I’ve talked to have moved away from elixir due to low adoption and demand.
Tiobe, google trends, etc tell the same story.
I’d love to find evidence to the contrary but the low numbers of elixir devs seems to be both a supply AND demand problem... ie lack of traction and interest.
In the latter case, I feel confident enough about Elixir that I'd train someone to use it rather than opting for a more 'mainstream' stack.
I suppose if getting new developers on-board asap is a concern, Elixir might not be the best choice. But personally if I ever were to build a product that would eventually need more programmers, I'd pick Elixir even just for the 'filter' effect. It's easy enough to pick up for a good developer, but would filter out many of the shittier devs I'd encounter if I went for PHP/Python/Ruby. and the benefits are worth it to me.
All that aside, while I am not confident that Elixir will gain much traction, I think there's a small chance LiveView and Telemetry, among other projects, might divert Elixir from the path that, say, Clojure, seems to have taken when it comes to traction.
I'm not sure I would say elegance is the main feature of Rails, but I think productivity is. IMO nothing out there beats Rails in terms of productivity on a new project. Phoenix is clearly inspired by Rails, but if you'll permit me to pull a semi-random number out of my ass I don't think Phoenix will ever be able to do better than about 80% of Rails's productivity. Some things are just not as simple (for example, just due to the nature of things Ecto is not as simple to write with as ActiveRecord is).
That said, I've totally moved away from Ruby on Rails and over to Elixir/Phoenix for my stuff. I absolutely love Elixir. So far in my experience I'm finding it more performant, easier to maintain, but a little less productive to write in. But still plenty productive enough.
I haven't used Elixir/Phoenix myself, but my main issue with Ruby on Rails was how much of long-term productivity was sacrificed for short-term productivity gains.
In other words, what made it extremely easy to create the first version of the application with Rails, also made it very hard to maintain and scale it later on.
Therefore, I would consider it a win, if 20% decrease in short-term productivity provided an equal increase in long-term productivity, which is much harder to measure.
I not saying your wrong, just that a hand wavy “rails is bad in the long term” is not an argument.
It took me 4 hours to learn LiveView, 3 to get it up and running with Phoenix PubSub (which I didn't know before either), and another 2 to get Presence in there. So I now have a live dashboard for the backend of our service (and it will probably be similar to the frontend). The rest of the week was developing a system so that I could run acceptance tests concurrently (you have to shard your Registries and your Channels and use a secret erlang feature to track who's what). Everything "just worked".
I think this 10 hour enterprise would have taken me at least a month to do with a standard RESTful interface and a backend off the shelf PubSub like RabbitMQ, and probably two months to do with an non-featureful backend websockets system (I find websockets confusing).
It's hard for me to believe that any other system is as productive.
You can run Sinatra on TruffleRuby today and it's very competitive with e.g. Go web frameworks. It's not quite production ready but has commercial support from Shopify (Chris Seaton) and Oracle Labs. It will get there soon.
There's also no reason I know of that you couldn't build a simple optimizing JIT for CRuby like LuaJIT. I'm trying a proof of concept at the moment.
But yeah, we keep mixing implementations with languages.