SRL – Simple Regex Language
simple-regex.com
simple-regex.com
- You don't need to remember which regex syntax the library is using. Is it using the emacs one ? The perl one ? The javascript one ? that new "real language" one ?
- It's "self documenting". Your combinators are just functions, so you just expose them and give them type signatures, and the usual documentation/autocompletion/whatevertooling works.
- It composes better. You don't have to mash string together to compose your regex, you can name intermediary regexs with normal variables, etc.
- Related to the point above: No string quoting hell.
- You stay in your home language. No sublanguage involved, just function calls.
- Capturing is much cleaner. You don't need to conflate "parenthesis for capture" and "parenthesis for grouping" (since you can use the host's languages parens).
Only for the time being, while only one regex combinator API/implementation exists for your language.
Even if the actual combinators are different, it's still a much better situation than the basilions regex syntaxes with completely different quoting and escaping mechanisms.
This fails at compilation with Opening paren has no matching closing paren. at position 0 in string "(ab"
(defun scan-ab (s)
(ppcre:scan "(ab" s))
By the way, here is the alternative syntax: (defun scan-ab (s)
(ppcre:scan '(:register "ab") s))
The alternative syntax allows to embed string regexes: (defun wrap-regex (regex)
(typecase regex
(string `(:regex ,regex))
(t regex)))
This is useful for combining regexes: (defun exactly-some (&rest choices)
(let ((choices (mapcar #'wrap-regex choices)))
`(:sequence :start-anchor
,@(if (rest choices)
`((:alternation ,@choices))
choices)
:end-anchor)))
(exactly-some "t.*")
(:SEQUENCE :START-ANCHOR (:REGEX "t.*") :END-ANCHOR)
(exactly-some "t.*" "a.a" ':whitespace-char-class)
(:SEQUENCE :START-ANCHOR
(:ALTERNATION (:REGEX "t.*")
(:REGEX "a.a")
:WHITESPACE-CHAR-CLASS)
:END-ANCHOR)Suppose all regex operator characters have to be backslashed (and \\ stands for a single \). Then it's clear. No backslash means it's literal; otherwise it's an operator (and a backslash on a nonexistent operator is a parse error).
The ambiguities exist because regex aficionados want common operators to be just one character long.
How readable is this:
`\^\[-+\]\?\[0\-9\]\*\.\?\[0\-9\]\+\$`
Some examples are a superfluous closing parenthesis or a [ bracket in a character class.
A sane regex syntax treats these characters as a separate lexical category from literal characters, in all contexts, and only produces literals out of them when they are escaped.
https://groups.csail.mit.edu/mac/users/gjs/6.945/psets/ps01/...
For programming it could be nice. It's an interesting idea. But I've found that good highlighting solves most of the problem. When I type "\\(" into emacs, it immediately highlights that \\( in a different color, along with \\| and \\). That way you can differentiate instantly between capture groups vs matching literal parens. Here's an example: http://i.imgur.com/b417O2o.png That small snippet would become way, way longer if you use a regex combinator. Maybe good, maybe bad, but it's hard to read a lot of code.
Yeah, it's ugly. And there are a bazillion small variations between regex engines. But for raw productivity, it seems hard to beat.
So SRL sounds like a great idea: us mortals can use it to painfully stitch together our our poor regexes, and when we get good enough we can skip SRL and become fully fledged regex gods.
Off topic, if SRL is a regex compiler, I'd love a regex decompiler. To be able to turn an impossible jumble of weird characters into a structured description of what it does would be hugely useful.
I can also recommend the webbased regex101.com, as it also supports non-JS regex flavors and explains things quite well.
[1]: https://groups.csail.mit.edu/mac/users/gjs/6.945/psets/ps01/....
About your example ... on the contrary, it would become way cleaner with regex combinators. The string manipulation would be replaced with proper composition and the escaping madness would disappear. For example
(concat foo "\\|" bar "\||" baz)
becomes (alt foo bar baz)
You seem to be familiar with elisp. The lisp family is particularly adapted to combinators approach, please read the link above. :)Combinators are very convenient and precise, but the tradeoff is that the code is longer. And I have to look up what to write every time I want to write one. But that's a personal bias.
Thanks for the reference. I'll study it.
EDIT: One of the good ideas in SRL is "if not followed by". There are too many [not]ahead-[not]behind combinations to warrant special syntax for each of them. I wonder if it could be streamlined, though?
If you don't know, SRE is a DSL in scheme that is essentially an alternate syntax for regex that does this, originally implemented by SCSH, and now most popularly by irregex. It looks like this:
(w/nocase (: (=> name (+ (or alnum ("._%+-")))) "@"
(=> domain (: (* (or alnum (".-"))) "."
(>= 2 alpha)))
eos))
I don't know if that's exactly what you were talking about: It's an alternate syntax, not a set of functions, which is what combinators usually imply.^^^ Regexp combinators in JS. 614 bytes minimized and gzipped.
Once you're proficient with the standard RegExp syntax, you can write pretty complex ones without external tools, but when the times comes to debug and modify them, using combinators makes the task much much more simple.
Here's a translation of the first SRL example in the parse dialect:
[
some [number | letter | symbol]
"@"
some [number | letter | "-" ]
some ["." copy tld some [number | letter | "-" ]]
if (parse tld [letter some letter])
]
And here's a full matching example: number: charset "0123456789"
letter: charset [#"a" - #"z"]
symbol: charset "._%+-"
s: {Message me at you@example.com. Business email: business@awesome.email}
parse s [
any [
copy local some [number | letter | symbol]
"@"
copy domain [
some [number | letter | "-" ]
some ["." copy tld some [number | letter | "-" ]]
]
if (parse tld [letter some letter])
(print ["local:" local "domain:" domain])
| skip
]
]
Some parse links:* http://blog.hostilefork.com/why-rebol-red-parse-cool/
* https://en.wikibooks.org/wiki/REBOL_Programming/Language_Fea...
* http://www.codeconscious.com/rebol/parse-tutorial.html
* http://www.red-lang.org/2013/11/041-introducing-parse.html
It'd be difficult to implement as tightly in another language that doesn't have Rebol's (or a somewhat Lisp-like) free-form code-is-data/data-is-code[2] approach. Rebol and its Parse dialect share the same vocabulary and block structure that amongst other things: makes it easy to insert progress-dependent Rebol snippets within Parse; build Parse blocks dynamically (mid-Parse if needed!); build Rebol code from parsed data; develop complex grammar rules very similar to EBNF[3]. The article linked above[4] (and now linked again below :) does a good job of fleshing out these ideas and why it may remain a unique feature for some time.
[1]: http://reb4.me/tt
[2]: http://rebol.info/rebolsteps.html
In more general terms, if a regex is complicated enough that something like this seems to make sense, the problem is that your regex is too complicated, and you should fix that.
I disagree. There is no such thing as a "complicated regex" in itself; it's all the same to the underlying engine. It's the maintenance of the regex that's the problem. You could, for instance, compile many small, maintainable patterns into a (technically) complicated regex with guarantees it will evaluate as expected—which is one possible way the outlined SQL approach could work. Don't confuse process with technology.
So what you're saying is that as long as I have a PCRE library linked into my brain, there's no problem. Gotcha.
To validate a user-supplied email address, you arguably just need something along the lines of ^.@.\..$ to help avoid whitespace and forgetting the @ sign. In a lot of cases, if the email address is wrong, all that happens is either a) user login fails or b) user registration fails. Hence there's no need for unreadably-complex regex to validate.
See the earlier discussion. It's not nearly so simple.
The word "either" implies only two choices, making your opening example confusing when the first "either of" was really picking from three possibilities.
I like the general approach, although the 2010-era BDD fake natural language is a turn off.
;o)
In both cases one looks up the documentation, in one I find that search for the line start requires a regex with "^" and in the other I find something like "begin with" of Simple Regex Language (SRL). I still need to read (or test) to find what "begin with" means and I still couldn't guess it - why not "start with", "open with", "first character", or a myriad of other possible options.
Whatever suits the user I suppose.
Whilst the conventional regular expression syntax is arguably overly compact, this is just too far in the opposite direction!
Something more PEG-like, or even Perl 6 regex-like, would make more more readable regular expressions whilst not completely throwing out everything we think things mean. Hell, even /x -- ignore whitespace and comments -- can make things much clearer:
/ ^
[0-9a-z._%+-]+ # The local part. Mailbox/user name. Can't contain ~, amongst other valid characters.
\@
[0-9a-z.-]+ \. [a-z]{2,} # The domain name. We've decided a TLD can never contain a digit, apparently.
$ /x
Tangentally, there's no point validating email addresses with anything more complicated than /@/. If people want to enter an email address that doesn't work, they can and will. If you want to be sure that the address is valid, send it an email! https://docs.python.org/2/library/re.html#re.VERBOSEWell once you have this in database, maybe you'll give it to some library, then maybe this poor little library will just paste it into the SMTP conversation. And maybe some user will be clever enough to exploit it.
I believe that if you want to get X from the user, it's always a good idea to make super sure that it is actually X before passing it further.
And I very much agree that whenever possible and necessary, simply splitting regular expression to multiple lines and using comments seems like a superior approach.
rx:i/^^
[ <+ alpha + digit + [._%+-] >+ ] ** 2 % '@'
'.'
<alpha>** 2..*
$$/
Notice that this avoids the repetition of a pattern.Of course, it'd make much more sense to write a grammar:
grammar Email {
token TOP { <name> '@' <domain> }
token name { <valid_char>+ }
token domain { <valid_char>+ '.' <alpha>** 2..* }
token valid_char { <alpha> | <digit> | <[._%+-]> }
}Note that your version is not actually the same, as it accepts all alphabetic characters, not just ASCII ones. Which put it much closer to RFC 6532, but still not exactly there due to the quoting rules in usernames. Which get pretty hairy. See [1] for an implementation of RFC 822, which is a simpler version of the modern standard, for a regex to validate email in Perl 5.
[1]: https://metacpan.org/source/RJBS/Email-Valid-1.200/lib/Email... trigger warning: bleeding eyes
http://www.ccs.neu.edu/home/shivers/papers/sre.txt
I think a better format for regex is long overdue, but this isn't it. It's way too verbose (other commentators also noticed the resemblance to COBOL). I'm picturing a Snort/Suricata rule with this format regex, and you've now doubled the amount of screen real estate per rule.
The real problems with regex readability are (1) the lack of easily grasped structure, so it's almost impossible to spot the level at which a sequence or alternation operates (PCRE's extended format and creative tabbing can help) and (2) the total lack of abstraction - so if you have a favorite character class or subregex you write it approximately a bazillion times.
http://www.slideshare.net/AurSaraf/re3-modern-regex-syntax-w... https://github.com/sonoflilit/re2
$ txr
This is the TXR Lisp interactive listener of TXR 147.
Use the :quit command or type Ctrl-D on empty line to exit.
1> (regex-parse ".*a(b|c)?")
(compound (0+ wild) #\a (? (or #\b #\c)))
2> (regex-compile *1)
#/.*a[bc]?/The examples give the gist: http://chubot.org/annex/cre-examples.html
More justification: http://chubot.org/annex/intro.html
doc index: http://chubot.org/annex/ (incomplete)
I showed it to some coworkers in 2013 and got some pretty good feedback. Then I got distracted by other things. One of the issues is that I learned Perl regex syntax so well by designing this language that I never needed to use it again :)
I plan on coming back to it since I'm writing a shell now, and I can't remember grep -E / sed -r syntax in addition to Perl/Python syntax.
SRL is the same idea, but I think it is way too verbose, which it appears a lot of others agree with.
If anyone is interested in the source code let me know! It was also bootstrapped with a parsing system, which worked well but perhaps wasn't "production quality". So I think I will reimplement CRE with a more traditional parsing implementation (probably write it by hand).
Bonus: COBORE -> (Japanese) kobore -> こぼれ -> 溢れ/零れ ("spillage").
"Overflowing spillage of verbosity."
Now if only I made this joke first...
https://www.slideshare.net/AurSaraf/re3-modern-regex-syntax-... https://github.com/sonoflilit/re2
If people like my direction, I may continue to work on it.
If you have fantastic tooling regex can actually be a pleasure
Unfortunately the best regex helper ever made seems to still be an old Windows app. But wow is it good: https://www.regexbuddy.com
I've seen online tools but they never seem to measure up.
It's the most disgraceful code style I've seen, and misleading: makes you mind think smth. async could be happening. I know Laravel popularized this, along with other ugly patterns, but let's stop cargo-culting this.
Example based on their example.
Is there a specific reason for the poor readability of regular expression syntax design?
If used as such, it'd be really nice to be able to go the other way - a regex explainer if you will.
I'm very interested in examples that extrapolate this idea to other areas of programming and even math. And also work in the reverse direction.
Most of the examples I've found are old or not open source.
another example of english to regex: https://people.csail.mit.edu/regina/my_papers/reg13.pdf https://arxiv.org/abs/1608.03000
English to dates https://github.com/neilgupta/sherlock
English to a graph (network representation) https://github.com/incrediblesound/MindGraph
C to English and vice versa http://www.mit.edu/~ocschwar/C_English.html
English to python: http://alumni.media.mit.edu/~hugo/publications/papers/IUI200...
English to database queries http://kueri.me/
This is a major step up in readability, so it's nice, and you have to invent a new syntax to do that, so I'll chalk that up as unavoidable. But did it have to be so verbose? SCSH/irregex's SRE had similar readability wins, with way less verbosity. You still have to learn a new syntax, though.
It's such a great resource.
They are subtle but they can cost valuable time while mentally context switching between them.
I'm not really sure what so different about the escape syntax in mySQL. But I'm sure you know.
So instead of writing this Hello World program in Brainfuck:
++++++++[>++++[>++>+++>+++>+<<<<-]>+>+>->>+[<]<-]>>.>---.+++++++..+++.>>.<-.<.+++.------.--------.>>+.>++.
You can instead have this much more readable version: increment byte increment byte increment byte increment byte increment byte
increment byte increment byte increment byte jump forward if zero
increment pointer increment byte increment byte increment byte increment byte
jump forward if zero increment pointer increment byte increment byte
increment pointer increment byte increment byte increment byte increment pointer
increment byte increment byte increment byte increment pointer increment byte
decrement pointer decrement pointer decrement pointer decrement pointer
decrement byte jump backward if zero increment pointer increment byte
increment pointer increment byte increment pointer decrement byte
increment pointer increment pointer increment byte jump forward if zero
decrement pointer jump backward if zero decrement pointer decrement byte
jump backward if zero increment pointer increment pointer output byte
increment pointer decrement byte decrement byte decrement byte output byte
increment byte increment byte increment byte increment byte increment byte
increment byte increment byte output byte output byte increment byte
increment byte increment byte output byte increment pointer increment pointer
output byte decrement pointer decrement byte output byte decrement pointer
output byte increment byte increment byte increment byte output byte
decrement byte decrement byte decrement byte decrement byte decrement byte
decrement byte output byte decrement byte decrement byte decrement byte
decrement byte decrement byte decrement byte decrement byte decrement byte
output byte increment pointer increment pointer increment byte output byte
increment pointer increment byte increment byte output byte include header file stdio.h, searching system paths first.
describe function main that returns a value of type
integer, and has argument of type integer argc, and
argument of type pointers to pointers to characters argv.
begin function body
call function printf with the single argument of type
string "hello world" with a newline appended.
begin new statement.
return the integer value 0. Wouldn't that be char ***argv? (with three stars) char **argv[]
It isn't. I've been away from C too long. Fixed in GP. /^(?:[0-9]|[a-z]|[\._%\+-])+(?:@)(?:[0-9]|[a-z]|[\.-])+(?:\.)[a-z]{2,}$/i
is a total strawman, needlessly obfuscated. How about writing it like this: /^[0-9a-z._%+-]+@[0-9a-z.-]+.[a-z][a-z]+$/i
which, while "scary looking", is at least immediately readable by anyone who knows even the basics about REs. If the argument for "verbose REs" is valid, it ought to stand up at least a typical standard RE.Also, it's not clear that "letter" and "[a-z]" mean the same thing. Does "letter" include uppercase? Does it include non-ASCII letters like "[[:alpha:]]" does? Don't forget the weird collation behavior "[a-z]" sometimes encounters.
/^[0-9a-z\._%\+\-]+@[0-9a-z\.\-]+\.[a-z][a-z]+$/i ^
[0-9a-z\._%\+-]+ # Local component
@
[0-9a-z\.-]+ # Domain name (subdomains permitted)
\.[a-z]{2,} # TLD
$The SRL documentation doesn't make it clear if they mean 'number' to be 'only numbers 0 to 9' or 'any digit in any language'.
"number" and [0-9] are even worse. That should have been called "digit" and, as another commenter already pointed out, in the age of Unicode, it still is confusing.
As to this attempt at simplifying regex writing and reading: nice try, but I think it needs more work. Apart from the Unicode thing, there's the fact that "letter" only is equivalent [a-z] because of the 'case insensitive' flag.
I think I would go for something that's less grammatical English and more programming language like (alignment of the colons optional)
Start of text.
1 or more : digit, lowercase or one of ._%+-
Literal : @
1 or more : digit, lowercase or one of .-
Literal : .
2 or more : lowercase
End of text.
All: case insensitive.
My default would be to have 'lowercase' mean the Unicode character class. 'ASCII lowercase' would handle [a-z]Adding capture groups, look ahead and look behind, comments, etc. is left as an exercise to the reader (they probably would make this look very ugly)
There's also the issue of nesting, like in this botched attempt to write a regex for URLs:
One or more of:
Once : letter or underscore
One or more : letter, digit or underscore
Separated by: /
Optional:
Literal: ?
One or more:
One or more : letter, digit or underscore
Literal : =
One or more : letter, digit or underscore
Separated by: ,
Of note here is that I think we need to digress from regexes a bit by introducing things like 'Separated by'. Without it, you often need to repeat potentially long phrases (programmatically building your regex can avoid that, but I think you still would need a serialization format, and I also think it makes sense for that to not use a full fledged programming language)Thinking of things of that complexity, I'm starting to think it would be better to have people write a BNF grammar.
You might be interested to learn that perl6 regexes have this, notated '%'. https://docs.perl6.org/language/regexes#Modified_quantifier:...
Nope, I'm mostly a DB guy very fluent in SQl and I use regex like two dozen time a year.But every time I nead to write something not trivial I must run to a regex cheatsheat website and spend long minutes trying to figure shit out.
It's not that I'm dumb and taking a MOOC about regex is definitely on my todo list... It's just that I haven't found the damn time yet to learn monstrosity and exceptions of regex.
And this is especially painful coming from PostgreSQL which have a good debugger and a clear syntax (even for non standard functions).
I was already cognisant with regex, so I'm perhaps biased, but a simple email-like search seems easy for a novice to read.
Perhaps you have a particular block when approaching regex. The book by Charles Severance that goes with the course (above) is freely available online.
for the same reasons I'm having mixed feeling about cucumber and similar testing frameworks (BDB), that also rely on semi-english language to do things. It looks cool, and enticing, but hard to sell (to others), even If I myself am super-excited to see it in action (just because how crazy it looked the first time I saw it).
I do wonder if having an EBNF compiler like ANTLR being more accessible would solve the readability & maintainability issues.
They might be ugly and everyone had a point where they thought they were just some sort of magic impossible to understand but I think that there is a point where they just... click. After that they can still be ugly and messy but you at least understand that there is a purpose to it.
Regular expressions are simple. It's just a matter of putting a bit of time in to learning them.
On the other hand, it did work for SQL...
TRANSFORM THE CURRENT OBJECT INTO THE DESIRED OBJECT.
Any other program statements are redundant."Literally, at sign."
"Like, literally, hashtag, guys."
I can't even.
14 LIKE, Y$KNOW (I MEAN) START
%% IF
PI A =LIKE BITCHEN AND
01 B =LIKE TUBULAR AND
9 C =LIKE GRODY**MAX
4K (FERSURE)**2
18 THEN
4I FOR I=LIKE 1 TO OH MAYBE 100
86 DO WAH + (DITTY**2)
9 BARF(I) =TOTALLY GROSS(OUT)
-17 SURE
1F LIKE BAG THIS PROGRAM
? REALLY
$$ LIKE TOTALLY (Y*KNOW)