Regular Expression That Checks If A Number Is Prime
iluxonchik.github.io
iluxonchik.github.io
Because of back references PCRE regular expressions can match non-regular languages as well.
There's an interesting book that covers some unusual aspects of automata and regular expressions: https://www.amazon.com/Finite-Automata-Regular-Expressions-S...
But that won't get you a primality checker. You can't get a primality checker. Primes don't comprise a regular language, neither in unary nor in any nontrivial base.
You can invent a variety of notations to describe a DFA. But you can't change the formal definition of a DFA, or what kinds of things a DFA can match without making it not a DFA.
I'm not actually a fan of using "regular expressions" in a sense broader than that of Kleene's regular, DFA-accepted languages, to be clear. That is exactly what I've been questioning.
I'm not actually a fan of using "regular expressions" in a sense broader than that of Kleene's regular, DFA-accepted languages, to be clear. That is exactly what I've been questioning.
On the other hand, backreferences can allow matching non-regular languages.
However, most programmers are not mathematicians or computer scientists, and so are effectively using lay terminology which is similar to but not precisely the same as the mathematical terminology. Most programmers only care that their regex library allows them to concisely match strings, not that it allows them to match regular languages.
Thus there are two standards you could follow: the precise mathematical definition of "regular", or the lay definition.
(And this is part of why I find it confusing terminology and would prefer "regular expression" only to mean specifications of regular languages in the original formal sense. But, of course, I'm not the king of language, and it's natural that people would speak the way they speak, given the circumstances.)
And surely you don't think those implementing the regular expression engine are lay people. They should know better.
The point here is that users of (colloquial) regular expressions, however professional they are or aren't, are just engaging with the surface of a deeper set of mathematical principles that require greater nuances of meaning to discuss than their applications do.
What? No. Many programmers out there are just working on fun projects in their spare time. I'm sure there exist some aerospace engineers doing that, but it's not many. I don't see why your average bedroom open source programmer should be treated like an aerospace engineer.
Considering this writeup leads with a giant headline saying "Regular Expressions - The Theory"... I think the complaints about how it's not actually a regular expression are even more valid than usual.
Understanding REs (a little more) formally, one can exploit their properties for e.g. performance and/or concision. For an example of what that understanding gives you, I'd recommend checking out the Ragel state machine compiler.
The difference isn't academical. It has huge practical implications, from software that is just annoyingly slow to DoS attacks.
People need to understand that. And stop using non-regular regexes (or better yet: regex in general) in production code. It increases world-suck.
The same features have also been implemented independently many times. For example Ruby, Java and Python all have independent implementations of regular expressions with a lot of the same features.
> Grouping and backreferences: All versions of sed support grouping and backreferences on the LHS and backreferences only on the RHS.
If the former, then it predated perl by quite a bit, of course.
I've tried finding some source for early versions but all the links I've found are dead. It's an interesting question!
Perl used that as a starting place and began adding extensions such as the difference between .* and .*?.
So using "non-regular thing" as a regular expression predates Perl. But the dialect we standardized on with the features people now expect is due to Perl's influence.
The sad thing is that the libraries use the same algorithms even if the expression doesn't contain backreferences. A while ago, Stack Overflow had a brief outage because of regular expression performance, although the expression that caused it didn't even use backreferences or other non-regular features:
http://stackstatus.net/post/147710624694/outage-postmortem-j...
In contrast, Google Code Search - when it still existed - supported regular expression searches over world's public codebases. One key ingredient making this possible was to only use proper regular expressions:
I'd argue that this is useful, even if you use it as part of a parser that doesn't just parse regular languages, as you'll have a slightly easier time reasoning about the time and space complexity of the construction.
"Thus the question arises: Can regular expressions match only regular grammars, or can they also match more? The answer to this is both yes and no"...
You can think of this as the languages that can be matched with an algorithm that is bounded (i.e., O(1) ) in memory, no matter what is the string to be matched.
The following pseudocode is not bounded in memory -- can you guess why?
bool is_prime(Number n)
{
for (Number i = 2; i < n; ++i) {
if (n % i == 0) {
return false;
}
}
return true;
}It should be noted, so can the value n, but in defining regular languages in this way, we allow our O(1) memory algorithms to be fed the characters of an a priori unbounded input string one by one sequentially.
You can also do a separate test for 2, then start at 3 and use i+=2 instead of i++, that cuts the test time in half. Once you've done that, you can observe that all primes are 6k +/- 1 except 2 and 3, which makes your code 3 times faster than the naive version.
At this point, you've basically got a simple wheel primality testing algorithm going, so why not go all the way?
I don't think the parent poster was making the point that his was the most efficient algorithm for testing primes. I think his point was that regular numbers are O(log n) in CS terms, not O(1), which is not obvious at first.
The regexp in the article is perfectly compatible with UNIX grep:
$ s=.; for i in `seq 50`; do echo $s | grep -qE '^.?$|^(..+)\1+$' || echo -n $i\ ; s=$s.; done; echo
2 3 5 7 11 13 17 19 23 29 31 37 41 43 47
(Tested on Mac OS X)In this case, bashing PCRE is completely off-topic.
>but they are definitely part of the traditional (UNIX) definition of regular expressions.
This is not exactly true. In 1968 or ealier Ken Thompson wrote first grep implementation using NFA simulation, an algorithm which is now known as Thompson Construction [1][2] There were no back-references in that grep.
In 1975 Al Aho wrote an egrep which used DFA instead of NFA [3] (both NFA and DFA accepts regular languages, but in some cases DFA will have exponentially more states than the same regular language accepting NFA automata [4]) This added features such as alternation and grouping which was not supported by grep, but it not supported back-references.
Current GNU grep have -E switch which accepts extended regular expressions as described in Posix standard. Theses extended regular expressions supports back-references as do PCRE available in grep with -P switch
So no, back-references are not in traditional Unix definition of regular language.
[1] https://en.wikipedia.org/wiki/Thompson%27s_construction
[2] http://dl.acm.org/citation.cfm?doid=363347.363387
[3] http://dl.acm.org/citation.cfm?id=55333
[4] http://cs.stackexchange.com/questions/3381/nfa-with-exponent...
- convert your number to a string of that length (1 = 1, 2 = 11, 3 = 111)
- handle 0 / 1 separately
- ^(..+?)\1+$
- trick; the WHOLE thing has to match to return (^$)
- first \1 match is "11", ergo string must be "11", "1111", "111111" to match
- second \1 match "111", ergo string must be "111", "111111", "111111111" to match
- and so on. if you find before length of string, it was not prime.
Clever trick. Look forward to being asked it in your next google interview :)The difference is the sieve works on a large batch of numbers, and only tests divisibility with numbers not known to be composite, and less than sqrt(n).
newsgroup comp.lang.perl.misc, October 1997. Message-ID slrn64sudh.qp.abigail@betelgeuse.wayne.fnx.com
(I don't know if abigail "invented" that, but I remember discussing it on usenet when it appeared in her .sig back in the mid/late '90s...)
(I was curious about how it worked but didn't need a long article about the basics of RE's, prime numbers, and other things I understand well enough.)
http://neilk.net/blog/2000/06/01/abigails-regex-to-test-for-...
(check out Abigail's other JAPHs if you like stuff like this)
I explored this technique some years later to solve a bunch of other problems, coprimality, prime factorization, Euler's phi, continued fractions, etc. (see https://github.com/fxn/math-with-regexps/blob/master/one-lin...).
I was very confused until I realized the author's definition of 'in front' wasn't the same as mine...
(Which is also how I would normally think of "in front", for what it's worth!)
So that means "W", "e", "l", and "l" are in positions 1, 2, 3, and 4 respectively? The "W" comes first?
Isn't the first position always in front of the second by definition?
But English never conceives of text in this manner; we view text as being arranged in a chronological order, where text that occurs chronologically earlier comes "before" text that occurs chronologically later. This mirrors the application of "before" and "after" to time in the rest of the language. Whether you conceive of reading as the reader traveling through text from the beginning to the end, or as text arriving at the reader, the reader will always encounter text on the left before text on the right, and therefore the text on the left is in front of the text on the right.
(In the only other language I'm qualified to talk about this for, mandarin chinese, earlier and later time might be indicated by either of two spatial metaphors: "up" for the past and "down" for the future ["up" is also used as a metaphor for beginning things]; or "front" for earlier and "back" for later. When text is read from left to right, "front" is used to indicate text on the left.)
Here ( http://www.friesian.com/egypt.htm ) is someone writing about the ancient Egyptian writing system, inadvertently assuming that the front of text is its beginning and the back of text is its end:
> Note that Egyptian glyphs have a front and a back. All the images above and below face to the left, [...] which indicates that the text is to be read from left to right. This is conformable with the usage of English and other European languages. However, although this would be familiar and agreeable to the Egyptians, Egyptian usage was ordinarily to write from right to left, as today is done in Hebrew and Arabic. They indicated this direction by having all the glyphs face to the right instead of to the left
(Egyptian glyphs often depict a person or an animal with an actual face. They face towards the beginning of the text, not the end.)
You seem to speak English at a fully native level, based on your writeup here. (Although you don't seem to have picked up on the idea that if a quantifier "precedes" a '?', it must be "in front" of that '?'.) Do you have another native language? Are you based in a country that primarily speaks some other language? What is the metaphor that determines that later words are "in front" of earlier words?
On a tangent, I really wish Google would release a new algorithm that punished this stuff. It's killing articles about common search terms. :/
This is a bit more Perlish (not that I'm an authority):
sub is_prime {
(1x$_[0]) !~ /^(?:.?|(.{2,}?)\1+)$/;
}
But perhaps less clear to a novice reader. You can also leave off the semicolon.You are probably right, but this function was my first head-to-head encounter with Perl, I researched just a little bit what I needed to know in order to write it.
(My first exposure to that was in '97 or so as a .sig from abigail in comp.lang.perl.misc)
@sorted = map { $_->[0] }
sort { $a->[1] cmp $b->[1] }
map { [$_, foo($_)] }
@unsorted;
(Twenty years or so back, I could occasionally be "that asshole"... I'm better now, honest...)> So first, we’ll be testing the divisibility by 2, then by 3, then by 4 and then by 5, after which we would have a match.
>> As a heads-up, I just want to say that I’m lying a little in the explanation in the paragraph about the ^(..+?)\1+$ regex. The lie has to do with the order in which the regex engine checks for multiples, it actually starts with the highest number and goes to the lowest, and not how I explain it here. But feel free to ignore that distinction here, since the regular expression still matches the same thing
[1] https://en.wikipedia.org/wiki/Sieve_of_Eratosthenes [2] https://en.wikipedia.org/wiki/Unary_numeral_system
\(^1\{-,1}$\)\|\(^\(11\+\)\1\+$\)
The two parts match what you'd expect on their own but the OR-ing screws it up: it means the whole regex matches everything.Is this something vim gets right and every other engine wrong, the other way around, or...?
_____________
[1] I'm matching 1's rather than dots to avoid highlighting every bit of text ever anywhere all the time.
Also, that way it's so much more pleasing to the eye and easy to read, don't you think?
No, it's harder to read. For ease of reading, you need to match something that isn't already part of the expression, like 2s or Ks.
Don't worry, I had my humour removed at birth also. It grows back, eventually.
Imagine naive absolute-beginner-programmer trial division. This is worse. Now add the overhead of counting via regex backtracking and integer comparison via matching strings. A fair number of regex engines will also start using enormous amounts of memory.
AKS is of theoretical interest, but not really a "normal test." It's very slow in practice, being beat by even decent trial division for 64-bit inputs (it's eventually faster, as expected, but it takes a while). But it is quickly faster than this exponential-time method. The regex is in another universe of time taken when compared to the methods typically used for general form inputs (e.g. pretests + Miller-Rabin or BPSW, with APR-CL or ECPP for proofs).
As others have noted, "has been popularized by Perl" is because it was created by Abigail, who is a well-known Perl programmer (though almost certainly a polyglot). It's also been brought up many times, though it's a nice new blog article. I hope the OP found something better when "researching the most efficient way to check if a number is prime." In general the answer is a hybrid of methods depending on the input size, form, input expectation, and language. The optimal method for a 16-bit input is different than for a 30k-digit input, for example.
Also, what is the fastest way to test for primality that's practically feasible?
Lots of math packages have deterministic primality tests, but none use AKS as a primary method, because AKS offers no benefits over other methods and is many orders of magnitude slower.
For inputs of special form, there are various fast tests. E.g. Mersenne numbers, Proth numbers, etc.
The fastest method depends on the input size and how much work you want to put in. For tiny inputs, say less than 1M, trial division is typically fastest. For 32-bit inputs, a hashed single Miller-Rabin test is fastest. For 64-bit inputs, BPSW is fastest (debatable vs. hashed 2/3 test). The BLS methods from 1975 are significantly faster than AKS up to at least 80 digits, but ECPP and APR-CL easily beat those. ECPP is the method of choice for 10k+ digit numbers, with current records a little larger than 30k digits.