I found the 'greenery' project, which says, at https://qntm.org/greenery :
> Elementary automata theory tells us that the intersection of any two regular languages is a regular language, but carrying out this operation on actual regular expressions to generate a third regular expression afterwards is much harder than doing so for the other operations under which the regular languages are closed (concatenation, alternation, Kleene star closure).
The developer has code, which lets you do:
>>> from greenery.lego import parse
>>> print(parse("(ab{0,3})*") & parse("(abba)*"))
(ab{2}a)*
>>> print(parse("((ss*)t*)") & parse("((ss*)+(tt*))"))
s+t+
I have no other experience with the code, but it was nice to know that such a package exists.The author also wrote that it was "the most algorithmically complex thing I've ever implemented."
/(abc|def)(123|456)/
You can read that as "(abc OR def) AND (123 OR 456)". The string "abc789" wouldn't match, for instance.
/(\D\S)+/
/(\D|\S)+/
/(\D&\S)+/
If you look at the string "b5 ", the first regex matches "b5", the second regex matches the whole string because all of the characters are either not a number or not whitespace, and the third regex only matches "b", because that's the only character that is both not a number and not whitespace.Secondly, there are very few intersecting character classes (sets) that I'm aware of, and in all cases, you could achieve the desired result more clearly in other ways.
Said another way: "AND" would just make regexes even harder to understand/approach, and that is almost always undesirable.
A more practical example might be something like
/(10|22)(.*crab.*&.*apple.*)90/
in order to only match strings where the content between the numeric codes matches both "crab" and "apple" in any order.To be clear, I don't know that it's useful enough to warrant inclusion in a regex engine. I'm just trying to provide a useful illustration.
(?=...)(?=...)
http://stackoverflow.com/a/24102539