Why Ruby Has Symbols
dmitrytsepelev.dev
dmitrytsepelev.dev
Ruby has symbols in all probability because Lisp and Smalltalk have symbols.
It could get most of the same practical upside of symbols from interned strings - the important thing is being able to compare using pointer equality and look up hash tables without needing to walk a string. What symbols at the type level do is ensure that these string-like things have already been interned, that is, de-duplicated, when they hit lookup points like member access.
But the implementation could do something very similar behind the scenes by setting a bit on interned string values. Besides, symbols aren't enough for the more advanced dynamic language optimization techniques like you see in V8.
I don't think so - unless you mean always interning all strings. The point of symbols is you can do a single address comparison. How can you do that if you could have two strings that are the same but have different addresses?
There is also a down-side of symbols - they by definition always escape the compilation unit since they're interned!
You intern all the literals (which includes lexical symbols) and are 99% if the way there.
I don't understand - if you're not 100% of the way there then you can't rely on address comparison. 1% of your string comparisons would fail!
I get a lot of use out of both of those decisions (immutable strings and string/symbol identity), they work well together, and I'd (much) rather have the problem of string-builders than the problem of tracking references to strings and copying them if I need both the original and revision.
That'd be catastrophic for performance in Ruby - every string allocation would always have to be reified, and would always need to access a shared data structure.
Yes but if you're regularly creating strings, which is what Ruby web servers do all the time, then your intern table is going to become a white-hot hotspot, contended by all threads all the time.
It's more accurate to say that Ruby has mutable strings as part of its semantics, and that can't be changed, and would interact disastrously with immutability.
Languages with immutable strings simply architect a web server around that semantic, the problem you're describing doesn't happen. A mutable string is replaced with (in Lua) a mutable array containing immutable string fragments, which are made into a string using the builtin table.concat.
That's what TruffleRuby does internally! We give you 'mutable' Ruby strings, but really they're made up of fragments of immutable strings.
Edit: I'm guessing you're saying they're immutable but not necessarily interned. Ok.
We already have the best solution... symbols. An alternative set of semantics that work well with our hardware.
In older Sun/Oracle JVMs, String looked something like
class String {
private char[] contents;
private int length; // substring() used to share contents with parent string
private int offset;
}
Since they switched to memoizing hashCode calculation, it looks something like class String {
private char[] contents;
private int length; // Allows Substrings to share contents with parents if they are a prefix of the parent
private int hashCode;
}
The fast path for equals() should first check for address equality between the two Strings, and if that fails, check if length and contents are equal, and if that fails, have a fast path for returning false if they both have memoized non-equal hash codes. If both strings compare as equal in the slow path, then one of them should be patched up to save memory and also hit the fast path next time equals() is called on them.When two Strings compare as equal, you probably want to use some heuristic to determine which one is likely to live longer into the future, and switch the other String to use the former's contents. Maybe in the Java case you'd only enable this optimization if length == contents.length to avoid some trickiness around prefix sharing optimizations. For a mark-sweep-compact garbage collector where lower addressed objects are more likely than not to have been allocated earlier, using the lower addressed contents as a tie breaker is a decent heuristic. Even in cases where address comparison is only as good as a coin flip in guessing age, it's a decent arbitrary tie breaker that in the long run leads to equal strings all converging on sharing the same contents char[].
Note that maximum String length in Java is 2*31-1 chars, so the sign bit of String.length is free for use as a flag for either keeping track if hashCode is memoized (if you don't want to use a sentinel value to indicate non-memoized hashCode) or if this String's contents should be the preferred version (for instance, if its backing store is in a shared library's read-only data section).
In languages with mutable strings, if maximum string lengths are also 2*31-1 characters (or if String hash codes can be safely truncated to 31 bits, etc.), then a single bit could be found to keep track of which contents arrays are copy-on-write.
If your strings are immutable and interned, they are as good as symbols; this is why Python does not have symbols.
ECMASript introduced symbols because JavaScript strings, while immutable, are not necessarily interned. Symbols are much cheaper to compare for equality: you only need to compare the pointers / ids, not actual string bytes.
Lisp has symbols for the same reason: Lisp strings are vectors, which are also mutable.
Since Lisp symbols serve a central role as identifiers and structured objects, they are not like what Ruby uses. Lisp uses symbols also for named interned things, but that is only one purpose.
In Common Lisp symbols have a name, a value, a function, a package and a property list (a list of keys and their values). By default in a call like (mult 1 2 3), the global function will be retrieved from the symbol and the function will be called with the arguments. The property list sometimes will be used by an IDE to store information about the symbol: like where it was defined, what its definition is and similar.
If you have five instances of `:foo` in Ruby, you can guarantee they will be IDENTICAL.
If you have five instances of `Symbol('foo')` in Javascript, you are guaranteed they will be completely DIFFERENT.
JS symbols are like Common Lisp's `gensym` which it uses to guarantee macro variables won't collide with existing variable names.
Aren't symbols and interned strings the same thing? Of course you can get all the upside of symbols by having symbols...?
I've never heard "lexing" used this way, and I believe it's simply incorrect. Lexing (tokenizing) precedes parsing (parse tree and then syntax tree construction). It isn't syntax tree validation.
Or so I thought. Are there other examples (besides this article) of "lexing" also being used to mean something else?
Thanks for the submission!
Indeed, I always thought that it's called "syntax checking" :)
I don't know what it was about this article but I feel slightly more confused now. It jumps to bytecode before explaining how they're useful at the higher level of abstraction that is the developer.
How do symbols help you think about and find solutions to problems and then implement those solutions? I.e., what is the facility they provide that is not present in some more traditional OOP language (say PHP?)
I think thr syntax is slightly nicer than quotes, but it's also more syntax, and there is a limit to how much you can have if you don't want code to look like Perl(Which Ruby is approaching).
People are saying it can can be used for some kind of better compile time checking, would be interesting to see that as the main focus?
Aren't Ruby's strings mutable? I'm not sure, but I seem to recall they were/are, in which case you can't really intern them. Python and Java have immutable strings, with the ability to optimize allocations being (probably?) one of the reasons.
On the other hand, Symbols in Ruby seem to be immutable, which allows for their interning.
e.g. This pseudo code must stand for mutation to be present.
a = ...
b = a
mutate(a)
/* b now mirrors a */Most places where you can use symbols instead of strings you lose nothing and gain speed.
I am not sure how it is in ruby though.
OTOH, symbols return true for both: they are the exact same object.
It's not just about speed though: symbols are somewhat limited in what they can be made of. They follow the same limitations as methods and variables. So often symbols are used when dynamically calling methods or assigning variables.
"Foo".public_send(:strip!) ¹. Which is slightly different from "Foo".public_send('strip!'). Not in outcome, but in calling. Because this is invalid syntax: "Foo".public_send(:one-two three) whereas this isn't: "Foo".public_send('one-two three'). Technically, I guess Ruby can have a method that is named "one-two three" but that would be really nasty to call. Symbols protect a lot against this.
And therefore are used in this context a lot.
¹ The exclamation mark can be a part of a method and symbol in ruby. As can the question-mark and some other sugar-ish stuff like [].
Prefix notation has none of those pesky limitations if you can live with it :)
Edit: oh. Scheme is painfully monomorphic. Equality for.strings is string=?. Equality for chars is char=?.
Then there is object equality (eq? ...), eqv? ("Normally eq?") and equal? which is a generic equality predicate that works for all objects (including circular data structures).
Eq? is the one you would use for symbols. Symbols are always (almost, at least) eq?. One string is only eq? To itself, but not a string containing the same content.
Only in the unquoted literal syntax. The :symbol form follows Ruby's usual identifier rules but there's also a :"quoted symbol" syntax. You can also send :to_sym or :intern to any string and it will be converted to a symbol.
https://ruby-doc.org/core/String.html#method-i-to_sym
> This can also be used to create symbols that cannot be represented using the :xxx notation.
'cat and dog'.to_sym #=> :"cat and dog"Those are more focused and simpler than strings in say PHP. And they are first class, unlike members in Java.
They are primarily first class names in your program. Think of them as distinct elements in a set, as opposed to arbitrary text to be transformed and parsed.
If I give you a symbol (or keyword) then you know it is a name. If I give you a string, it could be anything really.
Some of these use cases are little tokens like single words that are used in as values in function arguments, or a switch statement. On the other hand, storing the user's inputted name as a symbol whilst copying that data to your model object is probably not a good idea.
Yes this explanation leaves out a lot of detail.
Of course, the only thing prohibiting :err from being part of the data is convention. But if you're just looking over the code, the atom stands out almost like a syntactic feature, so it's easier to hold to. Plus, since they're interned strings underneath, you can use them in macros to make things like schemas unambiguously, and convert them into migrations with very little magic. So that's it, nothing you couldn't do with interned strings, variable names, and enums, but all in one handy little first class datatype.
--andrew
Python notably doesn't, and as such you get functions that take arguments that are strings with special meaning, which I always found a bit clunky even before I discovered Ruby.
Both cases would require some synchronization of data structures (across time or over the network), and with user-defined types this can get complicated, and atoms make the lack of user-defined types much more pleasant (and more performant than strings).
I don't know anything about Ruby or Erlang so I don't know if that's really relevant, just your comment seems to imply it doesn't.
In the context of computer algebra systems, which are much about manipulating abstract syntax trees, mathematical variables are usually represented as "symbols".
Beyond that, this page gives databases as an example, which is in fact very nice. Beyond being fast and efficient, using symbols allows certain errors to be compile-time instead of runtime, where typos are only detected on an application level and not on a code level. This is where symbols can play out their advantage. Think a bit of ENUMs in other languages.
i think the idiom you were going for was perhaps "you guessed it" or "you called it", as if poking fun at how, obviously, native symbols are helpful for symbolic programming, because it's the same word
Separating the two is useful semantically because it lets you differentiate between the two - and because these two kinds of string are better off being implemented and optimised in different ways.
So it's both UX and practical mechanical sympathy.
After a while, it was one of the things I liked the most. Maybe my favorite part is I didn't have to read and write so many damned " characters!
"Converting" "text" "like" "this" :to :this :format, really helps me read it.
I might be the weird guy on this. Wouldn't be the first time.
The technical benefits are nice, but this type of ergonomic feature is why ruby has remained my favorite language for over a decade.
I more and more dislike how Ruby (arbitrarily) allows omitting brackets. but not always. Often making the code harder to read. What is the call-chain in this rspec magic: `expect(something).to be >= 1` (quick: where and how do you add a custom failure message).
And while `attr_accessor :time, :date, :state` are really neat, I more and more dislike constructs like `validates :name, :login, :email, presence: true`. And prefer to write them explicit and unambiguous: `validates_presence_of(:name) etc`. Which is only a very slightly improvement over `validates_presence_of('name')`.
And don't get me started on "saving time" by typing less characters or shorter lines of code: if this is what makes you Go To Market faster, there's something very wrong with your IDE, editor or typing skills. If anything, those short things have cost me time in Rails codebases living years and years.
And fewer characters often is proportionally easier to read.
Intent sometimes becomes clearer with less characters. E.g. "attr_accessor" is, IMO, vastly superior to a large list of getter and setter method definitions. Easier to read, clearer in intent. Especially when there is that one getter or setter: you can be confident its doing more than just setting/getting (which probably is a smell, but I digress).
Details, however, hardly ever become clearer. A single `has_many :tags, :through => :taggings, delete: :cascade` may seem easier to read than explicit method declarations and callback registrations, but its a faux abstraction. It also rapidly falls apart when you continue developing on this for years and end with things like `has_many :posts, :through => :taggings, :source => :taggable, :source_type => 'Post'`
The abstraction remains in tact with a `define_relation(:taggings, DatabaseJoinTable.new(:taggins))` an `delegate :tags, to: :taggins` and a `register_callback(:delete, InlineTagginsRemover.new(self.taggings)`. I just made this up. But I tried to design an interface that is explicit rather than implicit. One that uses dependency injection and common Ruby-isms over a framework DSL.
Point is: behind those seemingly "easy to read" lines, there's a large world of black magick, lurking. I've dug through these forbidden forests on numerous occasions when our Rails app started misbehaving, race-conditions popped up, performance degraded, or even random dataloss. It's only easy to read on the surface. And while that is where we spend a lot of time reading, readability of the underlying stuff is even more important, because that is where the details matter.
Rails sacrifices the readability and understandability of what happens below the hood for readability and understandability on the surface. This seems a deliberate choice. But I dislike it. Severely.
That's Rails, not Ruby. Although Ruby allows it because of how flexible it is + metaprogramming.
But it is enabled by Ruby, as you state, by how flexible Ruby is. It may seem a nice touch that Ruby hands you the freedom to choose to e.g. omit brackets. But I think this is a bad freedom. As Rails shows, its a freedom that leads to, IMO, harder to read, and harder to reason about code.
With any language design, the limitations as well as its features, is what make the language. Limitations are an important feature of a language, IMO.
Now you can use frozen string literals and there's no benefit to symbols. Throwing "# frozen_string_literal: true" in the top of the memory benchmark script I get:
Calculating -------------------------------------
strings 0.000 memsize ( 0.000 retained)
0.000 objects ( 0.000 retained)
0.000 strings ( 0.000 retained)
symbols 0.000 memsize ( 0.000 retained)
0.000 objects ( 0.000 retained)
0.000 strings ( 0.000 retained)
Comparison:
strings: 0 allocated
symbols: 0 allocated - same
At this point with no practical difference between them STRINGS AND SYMBOLS BEING DIFFERENT ARE A MISTAKE. When you serialize to something like JSON you lose the distinction (the operation is singular and does not have an inverse transform) and you have to pick either symbols or strings to get back. On a long enough timescale this causes enormous confusion, and leads to the creation of hashes with indifferent access (which helps with the problem, but doesn't fix it).Ideally it would be good at this point to make symbols and frozen strings completely equivalent ("foo".freeze == :foo being true) but that would likely break too much existing code. The differentiation between strings and symbols though only causes code bugs (mostly biting the new and intermediate level programmers). It is just syntactical sugar with a footgun.
Designing a language from scratch these days, it should have immutable strings by default from the start and should not introduce symbols, unless they are purely syntactic sugar around creating an immutable string.
I somewhat miss this when I go back to Ruby, but then realize that symbols often can be used for `&str`. Often. Not always.
Gah no! This isn't true! Even if you turn on frozen string literals comparing two strings is slower because they have to test for a non-frozen and non-interned string also happening to be the same.
https://twitter.com/ChrisGSeaton/status/1514603665801109508
There's a pointer comparison, but behind it is on the failure side is a full-byte-comparison. Atrocious for cache even if the strings are tiny. If they aren't you're checking every byte!
Really both features (symbols and mutable strings) aren't worth the literally endless bugs that they cause.
Newbies try to use strings as enums, because they don't understand enums. Symbols provide tradeoffs compared with both.
This is what you get with a dynamic language where people try to overload functions with "I accept a scalar OR an array!!!"
https://twitter.com/chrisgseaton/status/1514603665801109508?...
Ruby took some things the designer liked from Lisp, from Smalltalk, from Perl, from other places. He liked symbols because they're good for performance (and compile-time correctness) at a low cognitive cost, and that's why Ruby has symbols.
IMHO it’s a style choice
const String ACCOUNT_FUNDS_EXCEEDED = "ACCOUNT_FUNDS_EXCEEDED"
is plain retarded :D