RegExr: Learn, Build and Test Regex
regexr.com
regexr.com
I am also aware of RegexBuddy (https://www.regular-expressions.info/regexbuddy.html), whose author also publishes very good regex learning content on their site. It looks great, but it's a closed-source Windows-only application, which means it's something I'll never be able to benefit from.
Why does it matter if it's closed source?
I am a paying customer of RegexBuddy. Best $39 I've ever spent on software.
RegexBuddy's `debug` feature has no equal in open-source or commercial software: https://www.regexbuddy.com/debug.html
It has many more features than regex101 AND it keeps all data local. Maybe regex101 keeps things locally too, but I have to run it in a browser and I'm not about to put sensitive test data there.
That being said, regex101 is very well done, but I paid $39 for RegexBuddy 14 years ago (and price is still the same) and last year paid $19 for the optional upgrade from v3 to v4...but I really only did that to support the developer. v3 still met all my needs.
For me it's definitely been one of those "best money spent ever" tools as well.
Would make for an entertaining show as the various malformed regexii produce partially correct outputs.
HN discussion: https://news.ycombinator.com/item?id=26644110
I think one reason why most people have a hard time reading regex is because they don't use any indentation or linebreaks. Honestly, if a buddy came to you and asked you to help him debug a javascript method and all 15 statements were on the same line, would you offer to help him, or tell him to fix his shit first so you can read it? What if it was all on one line and his variable names were all "v1", "v2", etc. Would you help him then? fuck no. And yet, this is standard operating procedure with regex, except you don't even get "v1", "v2" because nothing is labeled at all. v1/v2/... would be an improvement!
This is how most people write a simple date regex:
\d{1,2}/\d{1,2}/(\d{4}|\d{2})
And mind you, this is a very simple scenario. Here is how you would write it if you treated it like actual code:
(?<month>\d{1,2})
/
(?<day>\d{1,2})
/
(?<year>\d{4}|\d{2})
First off, you can know what my intent is when I'm capturing each group. Maybe this code gets used by a european where the month and day switch places. They can figure out how to fix it in like two seconds. Secondly, the forward slashes are not lost in a sea of characters anymore because we use whitespace like a civilized developer, not a regex savage.
If you want to keep things simple with regular expressions:
* Be liberal with what your pattern matches and use a normal programming language for your complicated conditional logic to filter out crap you don't want
* Don't be afraid to break up the search with multiple regular expressions
* Ignore pattern whitespace and use it to visually break up your pattern. Nobody would agree to debug javascript that has been minimized, yet people do this all the time with regex
* For the love of all that is holy, USE NAMED GROUPS. It is a fantastic way to document your intent.
However it is closed source, and it sends your info to the server for processing. This is of course not (likely) nefarious, but it does give me a little pause, and means that you should never use it if the data you are putting into it is sensitive!. For example, don't test your API key parsing regex on it with real/active keys! Also for corporate dev you are (probably) violating company policy by using it.
Rubulex[1] is a neat open source clone that I use a lot. The only downside is I have to start it locally. One of these days I'll stand up a permanent instance, though I don't want to do that without auditing the code and I simply haven't had time to do that. If anyone has done so I'd love to hear about it. Scriptular[2] is an open source clone that uses javascript:
Side note: If anyone knows of or wants to build an Elixir regex tool in Phoenix/LiveView, I'd be willing to collaborate a bit and willing to host/maintain (and I'll pay for the VM/domain). I've already got some Phoenix apps running in prod so if you get it to work with `mix phx.server` I can take it from there (I know that operationalizing and devops isn't what usually interests most people). A non LiveView version would be cool too, but it seems like such a cool project to build with LiveView to be super responsive and show off what you can do (and also learn Elixir/Phoenix/LiveView with a simple app). Would be really neat to release an elixir-desktop[3] version too! My email is in my profile if you are interested.
[1]: https://github.com/ofeldt/rubulex
I’m sure that these resources are great. And it’s not their fault that the regex family of languages evolved in the way that they did. But the simpler regex languages (without backreferences and other stuff… that I might not even know about) seem simple at first glance. In a perfect world I want to just spend and hour internalizing them forever. But in practice it seems that doubt always grips me, mostly because of the meta-syntax problem: did I unintentionally use some metacharacter in this part of the string which I meant to be “fixed”? So then I feel I have to “validate” it with some external tool. And suddenly it feels like this seemingly terse and agile language is just making me second-guess myself.
When in doubt, just throw a backslash in front of it, which always means "the next character is to be interpreted literally," even in cases where it's not necessary.
(Well, not always; the backslash will invoke a special character when thrown before some letters; eg, "\t" means the tab character. But normal letters never need to be escaped; just punctuation.)
(There is ongoing discussion about relaxing this rule for some characters, since it is so common in some cases. For example, escaping / so common that folks try to do it with the regex crate and are surprised when it returns an error. / is rarely a regex meta character, rather, it tends to denote the start and stop of regexes, e.g., in Javascript or sed.)
In fact, I was just on it this morning!
Something like IMEI should be pretty easy, if Wikipedia [0] is to be trusted (e.g. in Python):
# Matches IMEI and IMEISV
imei_pattern = re.compile(r"\d{2}-\d{6}-\d{6}-\d\d?")
You could write a big monster pattern that sets up capture groups for all the different TAC and Check Digit variants, but why bother? Just slice off what you need from the result after matching.0: https://en.wikipedia.org/wiki/International_Mobile_Equipment...
What other regex tools do people use and why?
If so, RegExr does have that tool (at the middle/bottom, tools bar), and it probably is the functionality I most often use.
That said, that's also doable in a text editor, but I like that I can visualize the pattern and results more easily.
- more flavor support,
- a regex debugger,
- code generator, with support for a lot of languages,
- a complete quick reference with examples,
- an extensive regex library,
- a regex quiz for "golfing" and learning purposes,
Perhaps, most importantly, it runs entirely client side and does not submit any information to the server unless you hit save (which returns a delete link to remove all data). You can even run the website and (most) of its features offline.
Regexr submits all input to the server for processing.
But, I'm biased, since I wrote regex101 :)
It's what excel uses for power query's "column from examples" too.
https://www.regexbuddy.com/analyze.html
Or it has the best regex debugger I've ever seen in my 20 year career: https://www.regexbuddy.com/debug.html
Written by author of RegexBuddy and Oreilly's Regex Cookbook: https://learning.oreilly.com/library/view/regular-expression...