What are your Ruby Regex Idioms?
blog.samstokes.co.uk
blog.samstokes.co.uk
The latter supports several languages' regex dialects, and I like the visual way it displays group matches, particularly when groups are nested. It can be a bit clunky, though, whereas Rubular is pretty clean and usable.
"foo@example.com".slice(/@(.*)/, 1) _, username, domain = */([^@]+)@(.+$)/.match("foo@example.com") username, domain = [*/([^@]+)@(.+$)/.match("foo@example.com")][1..-1]
And it is super readable. username,domain = */([^@]+)@(.+$)/.match("foo@example.com").capturesFWIW, I personally prefer the #match method. Some of the examples in this article just make me want to hurt people. Especially usage of $1, $2, etc.
username, domain = */([^@]+)@(.+$)/.match("foo@example.com").captures
Of course, your string better match, otherwise you'll get a NoMethodError.caller[0][/`([^']*)'/, 1]
Note that you only need to do this for Ruby 1.8. In Ruby 1.9, there will be a __method__() and __callee__() method that does the same thing.
caller[0][/`([^']*)'/, 1]
isn't really so much shorter than: re.search("`([^']*)'", caller[0]).group(1)
and the latter is quite a bit more explicit. >>> re.search('a(x)a', 'axa').group(1)
'x'
>>> re.search('a(x)a', 'aya').group(1)
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
AttributeError: 'NoneType' object has no attribute 'group'
The idiomatic python for the job is: match = re.search("`([^']*)'", caller[0])
if match:
match.group(1)
else:
None
(or some nonsense involving try/except, which would be even worse). try: re.search("`([^']*)'", caller[0]).group(1)
except: NoneWhat's interesting is that it's actually not special syntax - at least, it's not special syntax for regex usage. The syntax obj[arg, arg, arg] looks weird because in most languages things that look like subscripts can't have multiple arguments, but it's just sugar for the method call obj.[](arg, arg, arg), which is pretty mundane. The semantics are just the semantics of that method on the class of 'obj', and can therefore be defined in library or user code.
Note I'm not defending the specific use of it here to write odd-looking regex code - I just think it's interesting that the language is flexible enough to allow that kind of usage without special-casing it.
Raganwald wrote a blog post [1] along similar lines, arguing that it's a core part of Ruby's language philosophy, about this snippet:
(1..100).inject(&:+)
[1]: http://weblog.raganwald.com/2008/02/1100inject.htmlI think it would be more accurate to say that Ruby's String#[] aka String#slice method accepts arguments other than just indices (including regular expressions and other strings) for returning substrings/subpatterns. It's just a different approach, intuitive and rather uncontroversial, IMO.
In a dynamically-typed language there's no reason not to pass a Regexp, or indeed a Fruitbat, to a slice expression, so long as it behaves in some expected way. What I don't like about it is that there's no obvious convention for what that should be. String#[] has two fairly different behaviours depending on whether the argument behaves in one way or another. Because they have the same syntax, they look like they should mean similar things, but whether they're similar is debatable at best; and even if that case is debatable, the semantics of this are completely different:
Proc.new {|x, y| x + y}[2, 3] # => 5
(That is, Proc#[] is aliased to Proc#call, to work around the fact that Ruby's method call syntax is a special case whose semantics are not programmable, and can't be used on Proc objects.)More generally, I think it's a net win that the language detaches syntax from semantics in this way, and allows the semantics to be programmed; but for readability, syntax should still suggest semantics - i.e. a convention that should be followed when defining the semantics. The authors of the Ruby standard library didn't follow such a convention for [].
(foo){2,5}