Show HN: my regular expression example match generator
regexio.com
regexio.com
Imagine using some rules to take something like a user driven entry, where they'll enter something like say...a phone number freetext, but then use some rules that turn that phone number into a regex that'll match many common ways to write that phone number, then generate all the possible matching versions of that phone number for submission against an indexed database of text documents to see if any match. It's a trivial example, but a similar concept could work for things like names that have been Romanized from a non-Latin alphabet....generate a regex to match some variants, then generate the variants and search an indexed database for those variants. It would be much much faster in many cases than scanning through all the documents with the regex.
Hint: Try using "1(0+)1\1" or "([a-zA-Z]+)\1" for some nifty examples of things that technically aren't possible with true regular expressions, but are supported by enhanced regex engines like Perl/PCRE.
Edit: too bad it only works with one back reference, eg, adding \2 apparently breaks everything, returning no matches and no code.
Edit2: after more playing, it seems that using \3 works as a backref for the second matching group... try "(a+)(b+)\1\3"...
Likewise, I imagine it might be useful to specify some background language that matches should be taken from. That is, you might want to tell it to show you example matches made from snippets of valid html, from email-style English text, ascii-only strings, arbitrary unicode, etc. Incorporating this kind requires a lot of original thinking about usability (and some prototyping to figure out what would even make sense).
(Same thing for HTML: intersecting a regular language with a context-free one is context-free.)
[\d]+
But it didn't work. Took me a few tries to get it to work. I'm guessing it simply doesn't work with special regex characters.You know what would be even more useful for me? I give you a string and highlight the important part I need to match, and you give me a good regex for it.
Example: I need to match the string "View conversation (5)", with the important part being that there is a number near the end of the string. So I type that and highlight the "5". Then you give me something like:
/[^\d]+\d+[^\d]*/
Or whatever. Obviously it won't be full-proof because you don't know all my use cases and edge cases, but it'd give me somewhere to start.The quality of the results depends a lot on the amount and quality of the input (how well it represents what you are really trying to match). A single example string is unlikely to get you far - you would certainly need multiple matching examples. So I think this approach works best when you already have an example corpus to work from, rather than providing input manually. If you're going to spend effort providing a lot of input, then you'd probably be better off spending at least some of that effort in providing hints or possible regex answers.
Further, there are many possible regexes that would match an input set, so the algorithm also needs some way to evaluate them and choose the better candidates. In my case I used some ideas from information theory (such as MML), which actually worked reasonably well. But this is a computationally hard problem, so even with an objective measure of the optimal regex you won't necessarily be able to find it.
A an iterative process where you start from a single positive example and then iteratively refine it with clarifying examples would work without being overwhelming. However, I'm really trying to get the user to better understand regexes to the point where they could do that process for themselves, mentally, and not depend on a super intelligent tool to do it for them.
I wrote one myself, but was never really satisfied that I did a good job on it. I'd love to see the code of an efficient implementation of such a thing.
The expression itself provides a nice summary of the generative space it implies. Perhaps for a particular application you could apply a kind of lazy generation. Say, given /[01]+/ it would first expand it into the subspaces of /0/, /1/, /0[01]+/, and /1[01]+/. One of these branches could be expanded next depending on what was actually needed in the next step.
The general idea is that it is a command line tool that reads a regex as input (along with a number of examples to generate). It parses the regex into an abstract syntax tree, then hands it over to an interpreter to "execute" the expression/program several times.
The 'simpleparse' python library does a nice job of easing the regex-to-tree mapping.
I can tell because it's doesn't understand scoped flags or atomic groupings.