Go-Restructure: Sane regular expressions with struct fields
github.com
github.com
grammar email-address {
token TOP { ^ <user> '@' <hostname> $ }
token user { <[ \w \d . _ % + - ]>+ }
token hostname { <domain> '.' <TLD> }
token domain { \w+ }
token TLD { \w+ }
}
if (my $email = email-address.parse('joe@猫.com')) {
say $email<user>;
say $email<hostname><domain>;
say $email<hostname><TLD>;
}
It's also possible to turn the captured parts into their own objects ($email has a type of Match; it's not just a hash). function matcher(obj) { "use strict";
let props = Object.getOwnPropertyNames(obj);
const re = new RegExp(props.reduce((p, c) => p + (c.startsWith("_") ? obj[c] : `(${obj[c]})`), ""));
props = props.filter(x => !x.startsWith("_"));
return function(pattern) {
let o = {};
const res = re.exec(pattern);
for(let i = 0; i < res.length; i++) o[props[i]] = res[i+1];
return o;
};
}
And example usage:```js matcher({ _ : "^", user : "\\w+", _2 : "@", host : "[^@]+", _3 : "$" })("user@ycombinator.com") ```
Normally I wouldn't expect getOwnPropertyNames to return them in the same order.
[edit] Also, the API makes less sense because JavaScript isn't statically typed so you can make up the result object on the fly.
match('^(?<user>\w+)@(?<host>[^@]+)$', 'joe@example.com') => {'user': 'joe', 'host': 'example.com'}But now you need to parse the string and correctly extract named groups before flogging that to the regex engine, whereas the other way around you do some string concatenation then match back on the groups by index.
Some benchmarks would be interesting.
There is a certain amount of trigger happiness around when it comes to third party libraries. People tend to use third party libs WAY TOO much. Often they don't even check what is in the standard library.
I've thrown out over 50% of the third party libraries used in iOS projects I've taken over while simplifying and reducing the amount of code. The reason why using the third party libs grew the code was that third party libraries are usually quite generic, which requires more code to adapt to them. If your needs are quite simple, custom tailored code can take less space than utilising a third party library.
I'd really rather see something that generates marshaling code based on a regex or something, though. This loads an awful lot of meaning onto things that most code assumes doesn't have meaning, like struct field order. It looks really clever in isolation but if you start playing multiple tricks like this in one code base they'll start conflicting.
For example, note how you now can't use encoding/json on these objects as currently written. Now, that's fixable... well... it's probably fixable. AFAIK the struct tagging system isn't actually specified, so, for instance, if you try to put a tag
regex:"[^\"]+",json:"quotefree"
there's no guarantee how anything will parse that. Should that backslash be there to get the "character class of everything but double-quote"? Will the regex code get the backslash? Will encoding/json see a field 'regex:"[^\"' and a field ']+",json:"quotefree"' and then fail because there's only a field named ']+",json' and not one named 'json'? Will you encounter one of those situations where the backslash is simultaneously required and not required? Beats me. Plus I don't guarantee stability on whatever the answer is between versions, nor do I guarantee it if you reverse the order of the two things, nor do I guarantee it if you try to add a third struct tag.I would suggest taking this code and converting it into a function
UnmarshalRegex(*regexp.Regexp, val interface{}) error
with no use of struct tags, because you get 80-125% of the value, while dodging all the previous problems. I'd go ahead and say the user of the code is responsible for proper grouping, just describe how it needs to be, and check it at runtime. Especially if you use named capture groups, further removing issues of order.If you use backticks, the tag must not contain another backtick. On anything but a struct tag, you could get by by concatenating literals, but that's not allowed for struct tags: https://golang.org/ref/spec#Tag You could use double quotes instead of backticks, but that leads to a dark corner of escaping hell and an utterly unreadable mess.
Thankfully, the language strongly nudges you to very, very simplistic and light uses of struct tags.
I meant the internals of a struct tag, not the grammar. The standard library implies some structure with things like `json:"name,omit_empty"` but that structure is not actually specified AFAIK. And I seem to recall finding a github issue where the core team said they don't intend to specify one, basically for the reason that they don't want struct tags to be used for things like this systematically, but I couldn't google it up. If they fully "specify" struct tags I think they fear massive metadata additions, instead of little annotations here and there. I'm not quite sure enough of this to state it without qualification, but I am pretty sure it is accurate.
https://golang.org/pkg/reflect/#StructTag
There's no need to parse the entire tag string yourself and it's safe to assume other libraries will work with the convention. If they don't, that's not your fault and it would break with other tags like "json" and "xml" anyway.
One alternative is to build up regexes dynamically concatenating smaller regexes, but that leads to whole other kind of hell.
Really makes me appreciate the value of Go's field tags.
Plus not all regex engines support named capture groups, JS's regex don't, and while Go's engine technically supports them it might as well not: you can't index into a match using that, you get a list of group names of size the total number of match groups, and the name at index i is the name of the i'th match group.
Great concept and implementation.
Since we're discussing the performance footprint of regular expression libraries, it's also worth mentioning how slow even just running the base regexp library can be. For example, the email example could also be written without regex:
indices := strings.Split("joe@example.com", "@")
fmt.Println("Name:", indices[0])
fmt.Println("Domain:", indices[1])
( https://play.golang.org/p/ZezcoBjc9v )Obviously the above code doesn't do any sanity checks - which is where regular expressions can often make things easier. But the above would run a lot faster than a regexp pattern match.
Nice idea.