The evolution of Ruby's Range class
zverok.space
zverok.space
I guess this isn't super constructive, but to me the whole thing smells of not just a lack of discipline, but a lack of _interest_ in correctness that seems to be endemic in the Ruby community.
We ended up writing our own `TimeRange` class to paper over the base Range bugs that show up if you use it with times.
One awkward takeaway from the experience is that I've come to believe that sometimes it _is_ worth unit testing other people's code, counter to the popular advice.
My pet frustration in Ruby is `.slice` behavior.
> "abc".slice(0, 10)
"abc"
> "abc".slice(2, 10)
"c"
> "abc".slice(3, 10)
""
> "abc".slice(4, 10)
null
In Python, the last would also return "".edit: updated the code to reflect a basic repl, without using puts or the 'or' statement, which handles the falsy behavior that "" and null both have in Python.
I would argue that any input that doesn't have exactly one obviously correct output should result in an exception being raised. But Ruby often shies away from doing that. My interpretation is that they prefer to return _something_ rather than blowing up, so the show can go on.
I guess I sit at the opposite end of the philosophical spectrum from Ruby, so it's no wonder choices like this frustrate me. My personal philosophy is "explode early, explode often", because propagating invalid states only leads to future pain and suffering.
> [...] so much code now relies on its behavior, predictably or otherwise, that they couldn't change it.
I'm not denying this is a reason for keeping the questionably behaviour, but it's kind of funny because Ruby also introduces subtly backward-incompatible changes semi-regularly, including in "minor" releases. (Maybe that last bit is unfair, because I don't think that Ruby has ever claimed to follow semver.)
I think I saw this horror movie, but it was titled "PHP." :p
Coming from a C++ background my time in Ruby was weird and enlightening, it was less specified than C++, but pretty much everything is. I don't think it was underspecified, but a lot of the behavior was meant to be intuited and follows some kind of underlying pattern. If you miss the pattern you miss a lot of the functionality of that part of the language.
I agree that failing early and often is generally preferable. I don't see that Ruby doesn't do that, it just has more definitions for generic behaviors. Consider the amount of things in the range from the original post that produce type errors. Last time I checked both the C++ language and the C++ standard Library both had about twice as many pages in The Standard than all of Ruby in its standard Library have, yet there is still undefined and implementation to find behavior. So there are places where we get weird an unexpected Behavior that can result in bad State being passed around even in very formally specified places.
But, 3 is not a valid index into a string of 3 characters. By your criteria, "abc"[3] correctly returns nil. Whereas "abc"[3,1] returns ""
The language designers really just did not think this through.
There are two (intertwined) explanations:
- In Ruby exceptions are costly to generate.
- Similar to Go, Ruby's stdlib design favours using return values for signalling, keeping exceptions for exceptional cases.
So in plain† Ruby there are a lot of places that use `nil` to signal that something's off, e.g `[1,2,3][4]` returns `nil` because it's an out of bounds access, or `"123".downcase!` returns `nil` because nothing has been downcased, or `__FILE__` returns `nil` when a file is `eval`'d.
So yes, you have to check for things like `nil` (or explicitly swallow with `&.` which is kind of a Maybe monad pattern) before doing stuff with values, just like you check for `if err == nil` in Go.
Running a type checker like `steep` helps a lot linting for these cases.
† Rails likes to generate exceptions, but Rails != the whole of Ruby.
And your answer for Python is not quite correct: "" is falsy in Python, and both of the last two when translated to Python give "null".
And yes, I also understand the logic of the API; but if you're used to using slice to protect against random NPEs and out-of-bounds exceptions - which is something I do and am used to being able to trust in as a general pattern.
What gives you that impression? "abc".slice(4, 10) is perfectly valid and accepted, assuming the code above is accurate.
That's why if you put five or more in for the first index it fails to produce a result entirely. I think I might I preferred an exception or a failure code being returned, but I can't say the current design is truly awful.
Where does this come from? Are these discrepancies stemming from different Ruby implementations/versions behaving differently? "abc".slice(5, 10) returns the same value as "abc".slice(4, 10) [which, curiously, does not return the same value as the original comment] under MRI 2.6.1 that I had handy.
The inconsistency here is that when you call "abc".slice(2, 10) and get "c", Ruby has implicitly truncated the range to return whatever characters are available, even though it can't go all the way to 10 because the string isn't long enough. But then when you call "abc".slice(4, 10), it doesn't just give you all available characters from index 4 (which would be an empty string), it gives you null instead.
I don't see the inconsistency. slice on Array works the same way. Where is the inconsistency?
> (which would be an empty string)
What other aspect of Ruby would suggest that it is an empty string?
If what you are struggling to say is that different languages are different, then okay. "Japanese is unlike the English I know and therefore is inconsistent" would be a rather bizarre take, though.
Between what happens when the start index is greater than the length of the input, and what happens when the end index is greater than the length of the input. If the end index is greater than the length of the input, it returns a string (as long as the start index is not greater than the length of the input). But if the start index is greater than the length of the input, it does not return a string: it returns null, which is not a string.
My suggestion is that the behavior would have made more sense if it either returned a string in both cases (i.e., if it returned a string even if the start index is greater than the length of the input), or returned null in both cases (i.e., if it returned null whenever the end index is greater than the length of the input).
Again, what makes that an inconsistency and not just a different language?
> My suggestion is that the behavior would have made more sense
On the basis of the start and end indices being equivalent. But are they? What attributes of the language should see us consider them to be?
> What attributes of the language should see us consider them to be?
None.
I'm not sure where "care" enters into the picture. It's a computer language. For what reason would emotions be assigned to it?
I ask for a range whose start is in bounds but whose end is out of bounds.
Why should those return two entirely different types?
"abc".slice(4, 10) || ""
_and_ check those inputs in one go if needed. substr = "abc".slice(start, len)
raise "wrong inputs" if substr.nil?
just that one feature (only nil and false are falsy) of ruby makes it positively stand out among all the languages that have "null" as conceptRelying on a library you don't have any tests for puts a lot of faith in it working. Also in it continuing to behave as you expected when you change the version.
People have weird ideas about other people's code though. Something being found on GitHub may mean it doesn't need it be reviewed or even glanced through before putting it into production while code written in house is obsessed over.
To be fair Range is part of the language implementation. If you have no faith in the language implementation, why are you using it?
The popular advice is that you shouldn't test implementation, only interface behaviour. Which means that you shouldn't explicitly test the usage of someone else's code within your code. Their code is just an implementation detail. If the implementation is faulty, testing of the interface should still reveal that faulty behaviour.
That doesn't necessarily mean you shouldn't ever test someone else's code (with their interfaces). But if someone else's code is not tested it is undefined, and relying on undefined behaviour is fraught with problems. It is unlikely you would want to use it in the first place, which, practically speaking, leaves little justification to invest in improving the situation.
I've never heard that advice
I basically always unit test the code I depend on -- to learn how it works!
---
libc in particular is worth unit testing. C has lots of "creative" APIs that are easy to misuse.
Like functions that return the same static buffer over and over again. I think the only way you find that is by running tests with ASAN enabled ...
libc functions also have so many error conditions, and they may differ between platforms.
Can you predict the output of the following?
(1..20).each do |i|
puts i if i.odd?..i.prime?
endI thought I remembered Matz once saying it would be removed, But I could be wrong. Maybe someone uses it.
I've personally never used for anything else than strings, but when I do, it's very useful.
(1..20).each do |i|
puts i if i.odd? || i.prime?
end
edit - upon testing, I just realized the parent's code prints 10 & 16, my code does not, so not the same.“The form of the flip-flop is an expression that indicates when the flip-flop turns on, .. (or ...), then an expression that indicates when the flip-flop will turn off. While the flip-flop is on it will continue to evaluate to true, and false when off.”