Regexper: Beautiful regexp visualizations
regexper.com
regexper.com
I just want to thank everyone for the feedback so far, I am looking into the issues that have been brought up (they'll have to wait until this evening to be fixed though...I have a day job).
A graphical tool to create them, built of units instead of just strings, and the output could still be a regex string.
A question though: a lot of people seem to encounter issues/50x errors with the server-side calculation of the regexp. Why not have a client-side script to interpret the regexp?
Would save server ressources and wouldn't limit the user to server capacity to interpret more complex regexps.
(?:[^\xFF\x0D]+|\x0D(?:\x0A|\x00)|\xFF(?:[\xFA\xFB\xFC\xFD\xFE].|\xFF|.))*
The graphic makes it obvious why Telnet requires NUL to follow any bare CR's: if it didn't, you would need a byte of lookahead!Suggestions
1) it would be really nice if you could allow the user to share a regex with a friend. IE http://www.regexper.com/shared/some_unique_slug will display my saved regex
2) Add a Favicon
Great work thanks for creating this
Real-time character-by-character output would be important. If you are working on a long regex expression real-time output can help you think like the FSA and see how it behaves.
Here's an example of this tool failing:
<a.*?>([^<]*)<\/a>
http://gskinner.com/RegExr/?339m7or
http://rubular.com/r/pCXZs0DnSS
take your pick, same results.
If you remove the "?" from the regex you'll get different results.
One version selects all anchor tags individually and make a selection group of their contents and the other version ends-up selecting everything between the first opening anchor tag and the last closing anchor tag into the selection group except for "<" characters.
The graph in Regexper does not show any difference between the two expressions even when any one of these forms is used:
#<a.*>([^<]*)</a>#msi
#<a.*>([^<]*)</a>#gsi
/<a.*>([^<]*)<\/a>/msi
/<a.*>([^<]*)</a>/gsi
It looks like the flags are not being processed.Also, please change the textbox to a fixed width font.
In addition to that, regex authoring tools, ultimately, are only useful if you can also enter some input text and see the result, preferably in real-time.
Other than that, it's an interesting concept.
http://search.cpan.org/~dconway/Regexp-Debugger-0.001011/lib...
That's the documentation for using the module in your code. If you're interested in the standalone tool, you'll want rxrx:
http://search.cpan.org/~dconway/Regexp-Debugger-0.001011/bin...
I'd really like to see some more creative visuals like this for making program structure visible.
Edit: It's called a Railroad Diagram.
The MV* family of patterns seems to be okay without the notion of a controller.
Use case for this feature: You encounter an undocumented regexp in someone else's code and want get a few quick examples of what it matches.
http://ex-parrot.com/~pdw/Mail-RFC822-Address.html
and the server 500'd. Just joking, it's a very good tool.
P.S. it crashed with three nested groups (((a*)))
I'm also going to look into the RFC822 one as well (even though rendering something that size isn't really a goal). I suspect I might not be handling massive inputs sufficiently and want to make sure to address that.
You're probably lucky it didn't render though...that SVG would be non-trivial...
I agree that the value derived from email validation for something like a new account registration is almost nil. Just do something like an email verification round trip, which not only validates the email could be real, but also provides assurance that the user has control of the email address used.
Note: My last name is O'Kelley, which is a bit of a pain when it comes to poorly designed computer services. Some sites will strip the ' and others will escape it so that I become Mr. O\'Kelley. The worst I've seen is that my school used my complete last name in my email address, so most sign up forms refuse to even try to accept it, even though it is a valid email address. It can be a pain if I need to use my .edu address for academic discounts or verification.
No validation = Gone. You lost them. You can't email them for a correction.
With validation = You catch the issue before the visitor leaves and you ask them to fix it.
Sure, it doesn't verify the 1 to 1 relationship between the email and the person. That requires a round-trip verification. I get it. At least you ensure that it isn't all garbage-in to begin with.
The other aspect of email verification is that you don't have to choose to bug the user with the results. Depending on what it is, if someone enters an obviously junky address you can simply tag that email as potential crud in your database. Someone would then manually look at these every so often for cleanup or re-categorization.
I don't like the idea of looking signups or customers in a transaction where both parties are interested in transferring the information accurately. That's a use-case where validation works well.
Now, regarding your last name. The issue is cause by programmers who simply go around grabbing code off the internet without vetting it in any way. There are email "validation" regex expressions out there that are horribly wrong, yet people post them on blogs and others use them without question. It's unfortunate.
`/[^0-9^\+]+/` -> `/[^0-9\+]+/` # Thought that I needed to negate "+" as well.
^([^\W]+)@((?:[a-zA-Z0-9-]+\.)+(?:org|com|net|gov))$
Though this works: ^([^\W]+)@([a-zA-Z0-9-]+\.)+(?:org|com|net|gov)$
And it doesn't seem to be the nested groups that are screwing it up because this works: ((?:tinker|tailor)+(?:soldier|spy)) ^(([a-z]+)+)$My only suggestion is that it would be nice to have each regex diagram be separately addressable, by either putting the regex as a query param and yielding a raw image/svg/whatever, or by using a shortened url like a gist after saving.
Very nice work.
I mean, obviously unless you are writing an MTA or whatever.
The first assumption is that your intent is to capture accurate user information and the user's intent is to give it to you.
With that established, it is a good idea to apply sensible measures at your end in order to ensure that the data is accurate.
Then there's the reduction of junk signups and the like. This is a case where your user doesn't necessarily have to know that you are validating. You can simply tag the signup as potentially invalid in the database and require human inspection or discard it if there's enough information to do so. In other words, the validation happens in the server and with no feedback to the client.
The RFC822 regex expression will pass something like this as valid:
joe @example.com
joe@ example .com
joe@example.ccom
joe@example.com-
joe@example.com----------------
joe@example.com/////
You can check it yourself here:http://mythic-beasts.com/~pdw/cgi-bin/emailvalidate
I haven't done exhaustive testing on that particular expression. I don't really know where it fails. And that's part of the problem.
If the "contract" is that both parties want this information to be correct the user couldn't possibly be annoyed if you point out a legitimate error in the email address. You can prevent a situation where someone fills out a form, clicks "send" and goes away thinking that the signed-up when they actually didn't because they made a mistake entering their email address. What are you going to do? Send them an email?
Obviously, if they enter
jjoe@somedomain.com
instead of joe@somedomain.com
The only hope you have to catch it is to make them enter it twice and hope they don't make the same mistake twice. This is ugly and bad for such things as landing page signup forms. People don't generally respond well to having to type their address twice.So, yes, even with validation you are going to loose a few.
What you are going to catch are cases where the email is entered with detectable mistakes:
jot dot@somedomain.com
jotdot@@somedomain.com
jotdot@somedomain.com.
jotdot@ somedomain.com
jotdot@somedomain..com
jotdot@ somedomain.comm
etc.
instead of jotdot@somedomain.com
With a good email validation approach --which includes DNS checks-- you can catch most of these and alert the user. I don't think this is a bad idea at all.While it is true that that super-large regex expression seems to validate most addresses correctly, there's a voodoo out there in the realm of regex for email validation. Buyer beware.
An absolute FQDN ends with a period:
apple.com.
A relative FQDN does not: apple.com
When a DNS resolver sees an absolute FQDN it has a pretty direct path to getting the corresponding IP address.If, instead, the resolver is looking at a relative FQDN it does not, and it tries to fix it, and it could go around in circles for a while depending on where the resolver is running. What the resolver does is add different forms of DNS suffixes it has available until it either figures it out, fails or times out. For example, if the resolver is running on www.example.com, it might try the following:
apple.com.www.example.com.
apple.com.example.com.
apple.com.com.
If it finds any CNAME records it will follow them and go around in circles some more. I ran across this recently while looking at the whole issue of email validation through DNS. So, this is fresh pain you are hearing about!Here's my post on SO:
http://stackoverflow.com/questions/14065946/what-would-cause...
I also have a DNS function test page on one of my sites here:
http://www.tommyteaches.com/test/test2.php
If you try a domain like "notarealdomain1.com" you'll see that it takes a while (20 to 30 seconds per test) for the DNS functions to resolve it. If, instead, you enter "notarealdomain2.com." (period at the end) the DNS resolver will come back almost immediately. This is due to the DNS suffix append mechanism described above.
Note that if you enter the same name twice, the second test will return faster due to caching.
Here's an interesting read:
The more trivial the better (something like .[star]@.[star] ([star] == *, HN formatting bites sometimes)); people going too clever with validating e-mails (like if there was a good reason for doing that) end up rejecting + sign in the address as invalid, which is incredibly annoying for GMail users.
Use something like `^([^\s]*)@([^\s]*\.[^\s]*)$` which will do most of the work for you, then check second group for common domain typos, and what have you.
I don't understand. How does this expression do anything even remotely close to email validation?
For example, how does it tell you that:
These are valid:
test@nasa.gov
~~~~@nasa.gov
joe+sometext@nasa.gov
test@bbc.co.uk
and that: These are NOT valid
test@example.com (no MX RR)
test@-nasa.gov
test"@nasa.gov
test@nasa.gov-
test
test@nasa.rockets
test@bbc.co..uk
test@bbc.com.uk
test@bbc.co.eu.uk
You'd have to write all the validation logic yourself all over again. And that's just a few examples.Barring anything else, the RFC822 expression isn't so bad that someone should replace it with the kind of thing you are suggesting.
Sorry if I don't see it.
The best way to validate an email address is to send an email to whatever address is supplied to you, if it is a true email address the user will receive an email and it will be validated, if not then their account or query will go unused/unanswered and that will be down to them.
Multiple reasons, and, yes, context is important.
Landing Page: You have one, and ONLY ONE, opportunity to capture a potential new customer's contact info. If they make a mistake entering their email and you didn't catch it you'll loose them forever. You can't send an email to let them know they entered two periods by mistake, can you? They are gone and you screwed-up.
Every single potential customer is sacred. Thou shalt not loose them by being careless.
Forum signup: In general terms, if someone is visiting a forum it probably means that they want to sign-up. In this case, it is OK to make them enter their address twice, make sure they match and send them a confirmation email. They'll probably try to log-on later on and discover something went wrong and re-register.
While I said "that's OK", I also think it is bad form not to at least do enough validation of all input data, including email, to catch innocent mistakes. I think people who are against email validation might have that position because they don't understand it or gat bitten by a crappy regex expression and that is that.
Now your forum sign-up user is angry because they have to enter all of their information again and go through the process one more time. Who knows, they might make a mistake once again. While I don't have any data to back this up I would venture to guess that the drop-off rate for making a visitor enter all of their data multiple times is significant.
Payment Confirmation: Must check as much as you can.
From my vantage point taking ANY action that might loose or annoy a visitor is simply --to be kind-- programming. There's no excuse for that in my book.
About the invalid cases, who cares? It's not a problem, you must send an e-mail to check for correctness either way:
your user may
* wrongly type his e-mail e.g. bil.gates@microsoft.com
* write on purpose a valid e-mail of another person e.g. yourbestenemy@gmail.com
* write a grammatically valid but nonexistent address
* forget how to access his own e-mail address
You must always send a mail to confirm his validity, so if you have some false positives there is no harm, and it's faster to validate too.
Sorry, that's not a good reason to use this. If you use the correct approach you will NOT filter out syntactically correct emails and you WILL catch all invalid addresses that can be detected syntactically.
It just isn't a good idea to use this expression in place of the RFC822 expression. And, keep in mind, I am not a fan of the RFC822 expression.
With regards to your other scenarios, please read my reply to "rawb92" here:
http://news.ycombinator.com/item?id=5003032
In a nutshell, if someone enters a malformed email address by mistake and you don't catch it, it's game over. What are you going to do? Send them an email?
The spammers and tricksters will always exist. You'll have to decide how to deal with them yourself. In other words, stuff like someone attempting to sign-up their buddy to a porn site. That's got nothing to do with email validity, that's a matter of identification, and, yes, in that case the first line of defense is to send out a confirmation email.
You can use something like Mailcheck.js [0] for that client-side; it'll help weed out a lot of domain typos.
It ends up saying "none of [0-9]" when that's not actually a good english description of what should go there. It should say something equivalent to "one or more things that are not [0-9]". The empty string fits the description "none of [0-9]" but does not match that part of the regex.
I found a couple of bugs: it doesn't seem to handle `?>` and inline modifiers (http://www.regular-expressions.info/modifiers.html), and with a couple of expressions it simply gave me a server error with no further information.
Note: if you type [b-a], you fail to say it's an invalid range.
^1?$|^(11+?)\1+$
I noticed two glitches:1) It gives me the same output image for these two different regexes:
^[a-x]*yo$
^[a-x]+yo$
2) It gives me a server error for this: a(b*(c*(d*)*)*)
Edit: I was wrong about glitch 1. See below.> ^[a-x]* yo$
> ^[a-x]+yo$
No, it does not.
When you do (something)* , pay attention to the additional path that is created that circumvents the (something) box. With (something)+, there is no such path.
\b(?:((?:https?|mailto|s?ftp): )(?:\/{1,3} |[a-z0-9%]) |w{2,3}\d{0,3}[.]|[a-z0-9\.\-]{1,40}[.][a-z]{2,4}\/)(?: \&[a-z]{2,8}\; | [\w\(\)\.\/\:\@\#\?\=\&\-\!\~\;\'\[\]] | \%[0-9]{2})+
This is after changing ?> to ?: since it doesn't understand the non-backtracking syntax.
[\s\S]*
Doesn't this match every possible character, 0 to infinity times?Also, this would be a very useful tool in a Formal Automata class.
also how do you put in the case insensitive flag?
document.location.href = 'data:image/svg+xml,'+encodeURIComponent(document.getElementById('paper-container').innerHTML)
it'll open just that element as a SVG document, where you can save it with ctrl-S.I'd like to have used window.open instead of document.location so it wouldn't replace the RegExper tab, but that gets caught in my browsers' pop-up blockers :) Inserting it as a regular link in the original document would've been the best solution, in a pinch, I am offering this quick hack :)
Also, poking at this puzzle, I just noticed the badass colour green he's using for the diagrams and header.
\b((regexs)[=]+([AWESOME])(am|i|right|\??))?
-> server error. Looks like 3 groups within each other are the limit
Next level would be to turn this around and build a (nice) visual regexbuilder. The would be fing nice.
Even better: make it directly a "tiny" URL.
I realize that then some kind of a database would be needed but it would be really sweet.
For the uninitiated, there's an isomorphic relationship between regular expressions (not PCREs which are way more complicated) and finite state automata proving that if you have a regexp you can generate a FSA for it, and visa versa.
What we have here is a system that generates a graph of the finite state automaton for any given regular expression.
Neat project, misleading headline.
edit: let me add more constructive comments. What is useful about things like rubular (http://rubular.com/ ) is the ability to see how a Regexp behaves with a given input. Where a finite state chart can be useful is giving users a better impression of the internal workings of a particular regular expression. Where rubular can indicate where a match can be found, it would be cool to have a tool like this if you can show the path a particular input would take (and possibly fail out on) through a given regular expression.
Quite useful to show junior devs who are usually afraid of RegExp. Example from jQuery/Sizzle's: http://cl.ly/image/321F0E2B222o (^ and $ removed since it's failing to parse)
> Neither beautiful nor a visualization really.
This is the least helpful thing I can imagine. "Here's a beautiful regex visualizier" "LOL NO IT'S NOT". Geez.
There's useless pedants like this all over the Internet. You should probably wear a hat.
I'd agree if this were the top two comments without many disagreeing replies or down-votes.
I see this comment in down-voted grey, somewhere way down the page and assume a lack of social skills, bad-morning-before-coffee-grumps.
Or simply Internet Asshat Background Radiation, it's everywhere. Legends say you can trace its echos back to the First Flame from which the Internet was born.