You Should Learn Regex
blog.patricktriest.com
blog.patricktriest.com
When I replaced Perl with Python in my toolbox I abandoned regexes altogether because I didn't know too many Python programmers using regexes.
Now a decade later I am realizing how often I do need regexes in my day to day tasks (using them in bash commands, in my editor, and even in Python sometimes).
It wasn't pretty to look at but it was still to this day the fastest site I've ever worked on.
There were other jobs that could rebuild indexes by scan, repair associations or break things down into parts of use on other areas of the site but the areas that needed safety were lumped together as a single record (similar to the NoSQL approach honestly).
It wasn't a general purpose setup, but it worked for his purposes.
https://blog.codinghorror.com/regular-expressions-now-you-ha...
That's easy with the `e` (eval) flag.
say '13 37 123 42' =~ s
{ ( \d+ ) # get all the numbers }
{ my $total += $+; # replace with running total }egrx;
__END__
output is 13 50 173 215
> I think they're ultimately more powerful than in PerlNo. Perl regex are more powerful than Python because
1. the programmer can pick from a large selection of delimiters
2. the full-featured engine is built-in and ready to use, but in Python you need to install the third-party "regex" module because the built-in one named "re" is hopelessly dyd¹
3. you may enable verbose (readable) character classes
4. the regex subsystem is pluggable with a common interface and you can within a lexical scope switch at run-time with more restricted implementations which have less features but better performance in some cases: https://metacpan.org/search?q=re%3A%3Aengine%3A%3A
5. you can embed arbitrary code with (?{...}) and (??{...})
> We lose first class patterns, but gain first class functions and classes :)
Perl had first-class functions since 1993.
You love having first-class classes in Python, but do not realise the downside. Python's one true meta object system (with its ageing design from the last millennium) is flawed: horizontal composition of methods with the same name does not emit a warning. Since the system is baked into the language, it cannot be changed. Contrast with Perl, where the language only provides some primitives on which to build a meta object system (incl. first-class classes). This enables bugs to be fixed, and the competition among several implementations on CPAN breeds excellence.
* `\b([01]?[0-9]|2[0-3]):([0-5]\d)\b` you are using both [0-9] and \d, is that intentional?
* `cat test.txt | grep -E "^[0-9]+$"` UUoC and double quotes subjected to shell interpretation, should be `grep -E '^[0-9]+$' test.txt` or `grep -xE '[0-9]+' test.txt`
* `(?i)` won't work with `grep -E` (or at least not for me on GNU grep). features of BRE/ERE and PCRE like regex are very different - see https://unix.stackexchange.com/questions/119905/why-does-my-...
* avoid parsing ls - https://unix.stackexchange.com/questions/128985/why-not-pars... , use glob/find
* `.<star>?` again, this regex feature is not available with BRE/ERE, your example just happens to work. you can check it with `echo 'abc foo 123 bar 123' | perl -pe 's/foo.<star>?123//'` and `echo 'abc foo 123 bar 123' | sed -E 's/foo.<star>?123//'` (using <star> to avoid formatting issues)
shopt -s nocaseglob
ls ~/Downloads/*.{png,jpg,jpeg,gif,webp}
----
for command line tools (grep/sed/awk/sort/etc), you can refer my ongoing project (https://github.com/learnbyexample/Command-line-text-processi...) as resource :)
It makes a point, in "8.3 - For Problems That Don't Require Regex" that RXes should not be overused when a simpler solution exists.
But there are also problems that are better solved without RXes, but with a "parser" instead. The article mentions parsers, fails to mention that self-made parsers also provide solutions in the same problem space as RXes. Some communities (like the Haskell community) prefer parsers as they allow to be inspected, tested and typed. Parsers provide a more robust and hackable solution, that is possibly faster than the equivalent RX.
Here a short tutorial for a popular parser library in Haskell, to get an idea of what it's like:
Guys, no. Send an email. https://davidcel.is/posts/stop-validating-email-addresses-wi...
If you use regexes (or any other method of that does not send emails) all you're saying is that you don't actually care whether or not the string points to a recipient (much less the correct recipient).
And don't get me wrong. It is absolutely okay to not care whether or not the string refers to a correct recipient. Most places that make me write my email have no business caring about it. But please also then make the field optional.
You are correct to say that those emails may still bounce, and the phone calls may also not go through. We completely understand that. For this reason, in very specific situations (like registering a new user), we do take that extra step to make sure the communication channel actually works. But there are plenty of situations where that makes absolutely no sense, and/or adds very little value for the cost. Knowing the difference between these two very different use cases certainly does not indicate that these people "don't actually care" about the accuracy of their data.
If you want to do some additional non-regex validation, like confirm the hostname exists and has an MX record, have at it.
Besides the practical issues mentioned, it's this no-brained "why the heck not" collection of personal data I'm turning against. Either you need the stuff and then you have to work for it, or you dont and then you have nothibg to do with it.
My logs show that it helps a lot. Bounced emails from orders dropped by more than half.
https://github.com/mailcheck/mailcheck/blob/master/README.md
Note - In a real-world application, validating an email address using a regular expression is not enough for many situations, such as when a user signs up. Once you have confirmed that the input text is an email address, you should always follow through with the standard practice of sending a confirmation/activation email.
If you have free-form content, and need to identify email addresses in it, then you can’t just email every subset of that text.
That's what validation is. You're talking about finding email addresses, which I agree that regex is relatively OK at doing, but there are still lots of edge cases in the spec that are irregular.
It's useful to ensure you're not matching things that you don't mean to, which is a problem I often see when code reviewing regex.
> The source code for the examples in this tutorial can be found at the Github repository here - https://github.com/triestpa/You-Should-Learn-Rege
has a `href` of `""` which is making it point to the blog post itself.
`[0-9] - Matches any digit between 0 and -`
Mistype the `9` for a `-`. :-)
see https://en.wikipedia.org/wiki/Cat_(Unix)#UUOC_.28Useless_Use...
>Beyond other benefits, the input redirection forms allow command to perform random access on the file, whereas the cat examples do not. This is because the redirection form opens the file as the stdin file descriptor which command can fully access, while the cat form simply provides the data as a stream of bytes.
commands like `sort` are optimized to handle large input file
and how would you do `grep -l 'foo' *.txt` if you use `cat`?
or `awk 'NR==FNR{a[$1]; next} $1 in a' file1 file2`
or `grep -Fxf file1 file2` and so on....