^(?:http)(?:s)?(?:\:\/\/)(?:www\.)?(?:[^\ ]*)$
could be better written as ^https?://(?:www[.])?[^ ]*$
Or am I missing something? In that case, I'll readily admit this library is a good idea :) ^(?:http)(?:s)?(?:\:\/\/)(?:www\.)?(?:[^\ ]*)$
could be better written as ^https?://(?:www[.])?[^ ]*$
Or am I missing something? In that case, I'll readily admit this library is a good idea :)Writing regexes is not much of an issue usually (although the many dialects in common use are always a source of frustration) but reading them is always a pain, for me at least. For quick and dirty shell scripts or vim editing it's great, for stuff that's supposed to be long lived and actively maintained in a codebase I think this verbal approach is a great idea, at least in theory.
Regarding the optimization of the intermediate result it should only be a problem if you actually need to output these regexes for other uses or if you need to compile many of them at runtime with performance constraints. If your regexes are pre-compiled then the resulting DFA should look the same as far as I can tell.
If somebody makes a Rust crate with a similar concept I'll be sure to try it out next time I have to write regexes in a codebase.
It's actually a bad idea in this case because regex is mostly the same in every modern language, so if you know it, you know it everywhere. What you don't know is this.
I agree with the common complaint that regex is effectively write-only, but this is only half due to its terse syntax. A pattern can be pretty complex on its own, and complex things are hard to understand. Imagine what code matching behavior of a complex regex would look like.
I disagree, at least in my experience there are significant differences between multiple regex engines I'm used to use regularly. In no particular order: are parens and other operators treated literally by default or do they need to be escaped? Are character class like '[:alpha:]' understood, or do I need to write them explicitly? Similarly, do I have access to \w \W \s and friends? Can I use + to mean {1,} ? Can I use '?' to match 0 or 1 (common) or do I have to use = (vim)? Or maybe just {0,1}? But then should I escape the braces? Do I have recursion? Do I have named captures?
Those are not theoretical concerns, that's stuff I routinely end up getting wrong because I forget that this one feature that works in pcre does not work in vim or works differently in sed etc...
> Can I use + to mean {1,} ? Can I use '?' to match 0 or 1 (common) or do I have to use = (vim)? Or maybe just {0,1}? But then should I escape the braces?
I think that's just older tools like vi and sed. Perl, Python, Java, and Javascript use a similar modern version where + and ? work, and parentheses and braces don't need to be escaped.
Right, one language might have anythingBut(" ").endofline() and the next language might have a different . operator like anythingBut(" ")->endofline() or it might even require nesting calls. None of these things are a significant hurdle and if we standardize the names (endofline, anythingBut, ...) then you can make the same argument. It's a chicken and egg argument: just use regex because that works everywhere -> it's not universally implemented -> it won't work everywhere.
And aside from that, I have a similar experience to the sibling comment: when using some command line tool that I forgot (is it sed? Vim?) the default is that \( is a capture group whereas in normal regex ( is a capture group. Grep offers you three regex variants to choose from. I have to look up regex syntax or do trial and error every time I don't use a language that I use daily. And I don't know all of regex to begin with, I just know everything I ever needed but people posted examples here with (?:x) which I don't know. I once read it and remembered it for a few days I think... so anyway, consistent and descriptive method names seems a lot easier especially when you consider autocompleting IDEs.
https://github.com/VerbalExpressions/RustVerbalExpressions
Implementations for 36 different languages:
It's not just you. As you say this can only truly be used by people you understand regular expressions; and they would most likely prefer not to use this stuff.
It seems the whole IT industry is obsessed with helping us do all sorts of things, even simple things, which in the end often makes things more complex. Different query languages that translate to SQL to help us out, which often create super-complex SQL. All sorts of wrappers to avoid us having to deal with all sorts of formats (JSON/XML..). Hopefully those wrappers do something useful with those date-objects you know you have in there somewhere...
There's a niche where this might be useful, but by definition it's small. I understand regexes a moderate amount, and can construct arbitrarily complex ones when necessary. But I do it just infrequently enough that it can be painful and halting above a certain level of complexity, with lots of testing and reference-checking. It'd be nice to use something sane like this, and I think I fall squarely into the category of "people who understand regexes but would prefer to use stuff like this". Though as I said, this niche is almost by definition small, and on top of that I can't remember the last time I used Java.
Completely independently, in any non-trivial engineering system, readability is important, and this helps a lot there.
Doing it right is a delicate balancing act of being just powerful enough to express everything the user needs without devolving into an unreadable or repetitive mess. Some people manage to achieve neither.
That and UI SQL builders. What I want is typeahead column names, not a dropdown for the column, the operator, etc.
I know regex and I hate writing it. It's unreadable and I need to spend time remembering/googling/checking the exact syntax. And, of course, the syntax differs from implementation to implementation in subtle but important ways (ie: need to double escape in python, etc.).
These tools are great for letting someone build something they don't understand, and leaves them completely adrift when something goes wrong.
The next step is they bring this nonstandard thing to "the expert", who has to figure out their tool before they can figure out what's going wrong...
- SQL can already be made fairly readable by default, it's not just a long series of cryptic tokens. The main point of SQL builders is not to make SQL more readable, it's to make SQL approachable by people who don't know SQL.
- There can be several ways of achieving the same result in SQL, with sometimes deep performance implications, so it's really important to understand what is being executed and in what order. Regular languages are much simpler and while the string representation of the regex might end up longer than the handcrafted equivalent, the runtime performance should end up being the same since in the end it's all deterministic finite automatons.
- SQL builders have to be at least a little bit opinionated to be really useful, in general they make it easy to create simple queries but can quickly become limiting for complex queries, especially if you already know SQL. These "verbal expressions" on the other hand can easily map 1:1 with raw regex constructs, allowing somebody who already knows regex to express exactly the same logic, just in a more verbose and human readable way.
This verbose syntax operates at exactly the same level of abstraction as normal regex, it's just a syntactical transform effectively. It's like JSON vs. CBOR or something like that.
Which is also very true of regexes, especially the more feature-rich ones variants.
And the existence of variants was a large part of what I was getting at.
> it's just a syntactical transform effectively
Yes, it is tooling that helps people do things they don't understand.
Exactly, how can the same concept re-appear in all sorts of forms in this industry?
Another thing I though of was Mule, which I used a few years ago (it's hopefully better now). A horrid mess of an Eclipse-plugin that drew boxes with arrows between them, where "standard plugins" etc. could be plugged in to transform data and move it from here to there. The problem Mule solved (in our case) was comical, the complexity of the solution was also comical, or tragic; or maybe it was both.
I’m not sure this alone pulls its weight, though; most of the time, regular expressions are fixed at compile time. And I’d still prefer something that mostly preserves commonly-understood pattern syntax. Having to guess whether ‘anythingBut’ means (?!...) or [^...] is not encouraging.
(This was apparently ported from JavaScript, where it is even more pointless: template literals can take care of the escaping part without abandoning standard pattern syntax. But as far as I know, Java has no equivalent feature.)
If you use regexes a lot, you are better off learning regexes, if you use regexes a little, this is a lot to learn to avoid learning a little about regexes. But there is a moderate user sweet spot where I could see this useful.
Just escaping alone is a big selling point for me.
It may be under the hood, but there's no reason for it to be.
There's nothing inherent in our regexes that would imply they are 'the language' for that purpose, it just so happens we really only have one commonly used one.
Like most things invented forever ago, there might be opportunities for a 'cleaner, better way'.
But, regular expressions seem quite well optimize from my point of view.
Regular expressions are used for exact same task regardless of programming language -- using single expression language regardless of programming environment seems like a huge advantage. It can be embedded in configuration file, as a string in a database, on a web page or deep in backend code, and it will still work the same.
The "Java Verbal Expressions" already have "Java" in the name and so are complete loss when it comes to portability.
Then comes the fact that "Java Verbal Expressions" are many times more code that actual regular expressions. That isn't easier to scan, it is much worse.
Regular expressions are very succinct and you can express a lot in a single line of it. Comparable JVE-s would require many lines and wouldn't be more readable for anybody other than a person that doesn't know regexes at all.
It might be possible that s-expressions/style would would really well for regex, but nobody has really gone through the effort to do it.
Where I think regexes don't do so well is with UTF and (Grapheme clusters, true word boundaries) and also it's confusing the difference between match/capture/ignore etc..
I'll bet if you really put your mind to it creatively, you might be able to come up with a novel/new approach that wasn't really tried before ... but even if was 'better' it wouldn't catch on for a while (or never) unless there were some big, institutional backers.