Something that I found a bit ironic about the video linked in the article is at 5 mins 30 secs in; he states, "Why not give the memory address a name that makes sense to a person?" Here is referring to assembly and the abstraction it makes and how this relates to the abstraction that is variables. Symbols are really a cross system immutable memory address abstraction; I don't see how this is a bad thing.
I wish I had been clearer in my talk but I only had 30 minutes and wanted to cover other topics. Here is a more comprehensive argument against symbols in Ruby:
In every instance where you use a literal symbol in your Ruby source code, you could you could replace it with the equivalent string (i.e. calling Symbol#to_s on it) without changing the semantics of your program. Symbols exist purely as a performance optimization. Specifically, the optimization is: instead of allocating new memory every time a literal string is used, lookup that symbol in a hash table, which can be done in constant time. There is also a memory savings from not having to re-allocate memory for existing symbols. As of Ruby 2.1.0, both of these benefits are redundant. You can get the same performance benefits by using frozen strings instead of symbols.
"string".freeze.object_id == "string".freeze.object_id
Since this is now true, symbols have become a vestigial type. Their main function is maintaining backward compatibility with existing code. Here is a short benchmark: def measure
t0 = Time.now
yield
t1 = Time.now
return t1 - t0
end
N = 1_000_000
puts measure { N.times { "string" } }
puts measure { N.times { "string".freeze } }
puts measure { N.times { :symbol } }
There are a few things to take away from this benchmark:1. Symbols and frozen strings offer identical performance, as I claim above.
2. Allocating a million strings takes about twice as long as allocating one string, putting it in into a hash table, and looking it up a million times.
3. You can allocate a million strings on your 2015 computer in about a tenth of a second.
If you’ve optimized your code to the point where string allocation is your bottleneck and you still need it to run faster, you probably shouldn’t be using Ruby.
With respect to memory consumption, at the time when Matz began working on Ruby, most laptops had 8 megabytes of memory. Today, I am typing this on a laptop with 8 gigabytes. Servers have terabytes. I’m not arguing that we shouldn’t be worried about memory consumption. I’m just pointing out that it is literally 1,000 times less important that it was when Ruby was designed.
Ruby was designed to be a high-level language, meaning that the programmer should be able to think about the program in human terms and not have to think about low-level computer concerns, like managing memory. This is why Ruby has a garbage collector. It trades off some memory efficiency and performance to make it easier for the programmer. New programmers don’t need to understand or perform memory management. They don’t need to know what memory is. They don’t even need to know that the garbage collector exists (let alone what it does or how it does it). This makes the language much easier to learn and allows programmers to be more productive, faster.
Symbols require the programmer to understand and think about memory all the time. This adds conceptual overhead, making the language harder to learn, and forcing programmers to make the following decision over and over again: Should I use a symbol or a string? The answer to this question is almost certainly inconsequential but, in the aggregate, it has consumed hours upon hours of my (and your) valuable time.
This has culminated in objects like Hashie, ActiveSupport’s HashWithIndifferentAccess, and extlib’s Mash, which exist to abstract away the difference between symbols and strings. If you search GitHub for "def stringify_keys" or "def symbolize_keys", you will find over 15,000 Ruby implementations (or copies) of these methods to convert back and forth between symbols and strings. Why? Because the vast majority of the time it doesn’t matter. Programmers just want to consistently use one or the other.
Beyond questions of language design, symbols aren’t merely a harmless, vestigial appendage to Ruby. They have been a denial of service attack vector (e.g. CVE-2014-0082), since they weren’t garbage collected until Ruby 2.2. Now that they are garbage collected, their behavior is even closer to a frozen string. So, tell me: Why do we need symbols, again?
I should mention, I’d be okay with :foo being syntactic sugar for a frozen string, as long as :foo == "foo" is true. This would go a long way toward making existing code backward compatible (of course, this would cause some other code to break, so—like everything—it’s a tradeoff).
Apparently Rails 5 will take advantage of symbols GC to use symbols as keys for the params hash instead of strings. It uses strings now only to prevent DoS because of uncollectable symbols.
I find it a bit surprising though that an article that is generally in favor of functional programming is bashing Ruby symbols. Ruby symbols come from Lisp and are just a form of atoms which are common place in functional languages. Erlang, Lisp, Elixir, Clojure, Scale to name a few all have atom/symbol support in some form. They are immutable even across systems. It just seems to run counter to the article's argument. Also there was no supporting reason why symbols should be killed. The article is released after GC was introduced for symbols so I'm assuming that's not the reason to "kill symbols".
I'm not sure there isn't more about symbols than that but I'll think about it the next time I'll write some Ruby. I'll pretend I have only strings and see what happens. One thing for sure: we'd need a syntactical shortcut because having to type "string".freeze everytime is unbearable. If the right shortcut is :string or the lispy 'string, I don't know but if we need it symbols are fine even if symbol.class could end up being String.
As far as the difference between "string".freeze and :string; the symbol will resolve to the same object_id every time across systems and processes. If you spin up irb and type :foo.object_id you will see 1092508. Now one way this is commonly used in Ruby is when we use Object#send. We pass a symbol in place of the function name as the first argument and internally this is used as a system optimization to lookup the function that is being called dynamically.
Now I haven't been knee deep in the code of any of the Rack web servers, but I would imagine that something like Puma which is built for concurrency and multi-threading could take advantage of this behavior as well for some shared process to enable spinning up lighter weight additional processes or threads thus leading to lower memory consumption on your web server and the ability to handle more traffic on the same hardware.
Erik is right in saying that strings and symbols are "converging"; they have not yet converged however. At the point at which strings converge to symbols in the sense that both resolve to the same object across systems, then what we have left is symbols and so it is not symbols that have been killed, but rather strings supplanted by symbols. If instead we go the direction of saying memory and performance don't matter so let's punt on that whole symbols are immutable thing I think we are giving up too much.
I use Ruby because the high level abstraction is great and saves me developer cycles, but I also want it to eek out as much performance as possible. Honestly in general I'm tired of hearing the argument that X should die now because it is a performance hack. Some performance hacks are good.
(For reference: In the video he discusses symbols vs freezing strings from 22:00 to 25:07)
object_id for symbols is not consistent across processes.
% ruby -e "p :foo.object_id"
396968
% ruby -e "p :bar.object_id"
396968
% ruby -e ":bar; p :foo.object_id"
397128
You're getting 1092508 consistently not because it's :foo but because it's the next object_id available for a symbol after all the ones created during IRB startup.