Code + data here: https://github.com/nicholaslocascio/deep-regex
Code + data here: https://github.com/nicholaslocascio/deep-regex
In our case, we have no examples to test against, only a natural language (English) description of what the user wants the regex to do. This is an inference problem more than a search problem as we've got one shot to give our best guess without any tests to check against and modify our answer.
Some positive ones:
1) Spot-on prediction:
PROMPT: lines with 3 or more characters or lower-case letters
PRED: ((.)|([a-z])){3,}
GOLD: ((.)|([a-z])){3,}
2) Learned to generalize and produced a simpler regex: PROMPT: lines with a character and the string 'dog'
PRED: .*(.)&(dog).*
GOLD: .*((.)+)&(dog).*
3) Also learned to generalize and produced simpler regex without duplicate logic: PROMPT: lines not containing a letter
PRED: .*~(([A-z])+).*
GOLD: (.*)(.*~([A-z]).*)
4) Handling multiple references correctly: PROMPT: lines using 'su' after 'sun' or 'soon'.
PRED: .*(sun|soon).*su.*
GOLD: .*(sun|soon).*su.*
Though I find the mistakes interesting as well!1) Issues counting properly:
PROMPT: lines containing a 5 letter word beginning with 't'
PRED: .*\bt[A-z]{5}\b.*
GOLD: .*\bt[A-z]{4}\b.*
2) Misallocation of parenthesis (to be fair, the prompt is slightly ambiguous): PROMPT: lines with 'dog' follwed by 'truck' and a lower-case
PRED: (dog).*((truck)&([a-z])).*
GOLD: (dog.*truck.*)&(.*[a-z].*)Have you can considered generating something like a formal grammar, a recursive-descent parser or a software library?
So, the old ways are still better for this domain if it's a production system whose cost or results matter. These methods might be useful for search/query by casual users, though. Or people that come from a foreign language likely to express queries in a weird way.