14:51 [info] 51 some message
… more of 51 lines …
15:22 [error] 24 error!
… more of 24 lines …
^(\d\d:\d\d) \[(info|error)\] (\d+) (.+)$
Or maybe not a regex, but a structured pattern. 14:51 [info] 51 some message
… more of 51 lines …
15:22 [error] 24 error!
… more of 24 lines …
^(\d\d:\d\d) \[(info|error)\] (\d+) (.+)$
Or maybe not a regex, but a structured pattern.This could be built using set operations on deterministic finite automata (dfa). Every regex is equivalent to a dfa. You can now construct automata for every positive and negative example input. Then calculate the union for all positive examples and the union for all negative examples. And finally calculate the difference between the two unions. Convert the resulting automaton back to regex.
It shouldn't be hard to start with .* and resursively split it in two parts that still match the input strings, but I believe you will end up with matching but useless regexes.
There's research [5][6] as well as practical tools [7][8][9].
[1] https://en.wikipedia.org/wiki/Program_synthesis
[2] https://www.microsoft.com/en-us/research/project/program-syn...
[3] https://dl.acm.org/doi/10.1145/1836089.1836091
[4] https://royalsocietypublishing.org/doi/10.1098/rsta.2015.040...
[5] https://cs.stanford.edu/~minalee/pdf/gpce2016-alpharegex.pdf
[6] https://www.researchgate.net/publication/261794574_Automatic...
[7] https://regex-generator.olafneumann.org/
[8] http://regex.inginf.units.it/extract/
[9] https://stackoverflow.com/questions/6219790/need-a-regex-too...
% bundle exec bin/regexgen '14:51 [info] 51 some message' '15:22 [error] 24 error!'
(?-mix:1(?:4:51\ \[info\]\ 51\ some\ message|5:22\ \[error\]\ 24\ error!))
With enough inputs it should end up with something somewhat reasonable for the leading part, but it will never be smart enough to understand that the error message is "arbitrary" and should be matched with e.g. `(.+)`.EDIT: Ah no, sorry. Was thinking of the other way around[0].