Regex engine internals as a library
blog.burntsushi.net
blog.burntsushi.net
This article as a whole is also a treasure. There's a fairly limited number of people who have written a ton about regular expressions, but they all do so extremely well. Russ Cox's article series (as mentioned in this one) is really great. I used it in college to write a regular expression engine one summer after I had started to fall in love with the perfect cross between theory and practice that regular expressions are.
The changes for more in depth testing here are also very interesting, especially for a crate that I'm sure if critical so a lot of the ecosystem, and I appreciate the writeup on such a deep dive topic.
Are regular expressions hard to read? Sometimes they can, especially if you haven't gone out of your way to take deep dives. Are they not perfect for a lot of parsing tasks? Sure, there are definitely things people shouldn't use regexs for that they do, like validating emails. At their core though, regular expressions are one of the most power dense tools we have in pretty much every language. The insane stuff you can get done (sometimes in a way that would be frowned upon) in so little time if just incredible.
I'd love for there to be more content on regular expressions as well. Right now I'm only familiar with one book that really does a great job (at a practical level), and that's Mastering Regular Expressions by Jeffrey Friedl. On a theory level, a lot of compiler books will talk about them to varying degrees. The Dragon Book has some good content on regular expressions at an implementation level too.
Does anyone have any other book recommendations for regular expressions?
And if BurntSushi or any other regular expression engine contributors see this, I really appreciate the effort y'all put in to building such powerful software.
A regex is a very small concept with a very large syntax :) It's "just" literals, alternation, repetition, and concatenation. So it's mostly being able to see through the syntax IMO.
Nice animated diagrams in this recent post - https://news.ycombinator.com/item?id=35503089
(On the other hand, from the implementation POV, there is a wealth of material online, but you already know where to find that :) I still think someone should "cliff notes" the Cox RE2 articles, since the ideas are important but there is a lot of detail. There are also dozens of references in those articles for people to read; it was a comprehensive survey at the time.)
---
BTW Friedl's book covers backtracking engines, which makes sense because PCRE, Perl itself, Python, Java, JavaScript, etc. all use that style
But as far as I remember, most of the book is irrelevant for users of automata-based engines. It's a whole bunch of detail about optimizing backtracking, reordering your clauses, and lots of detail about engine-specific and even version-specific differences.
Personally I just ignore all of that, and it doesn't cause me a problem. I use regexes all the time in every language, and stick to the features that an automata-based engine can handle.
Once you go into the backtracking stuff, then you might want the Friedl book, and that's a huge waste of time IMO. It's way easier to do that kind of thing outside of a regex engine, with normal parsing techniques.
And honestly once you go outside, you'll find you don't even need to backtrack. Perl-style regexes can be an astoundingly inefficient way to write a pattern recognition algorithm.
---
Eggex is my attempt to separate the two kinds of things, since it's not obvious based on the syntax: https://www.oilshell.org/release/latest/doc/eggex.html
It makes the syntax smaller, more familiar, and more composable. It's closer to the classic lex tool (e.g. GNU flex). It's funny that lex has a way better syntax for regular expressions for Perl, but it's used probably 10,000x less often because most people want an interpreter and not a compiler to C code. But lex has 90%-100% of what you need, and it's way simpler, with a nicer syntax.
Literals are quoted so you don't have the literal/metachar confusion, and you can reuse patterns by name. The manual is worth skimming:
https://westes.github.io/flex/manual/Simple-Examples.html#Si...
I use re2c though, not flex, and its manual is also worth skimming: https://re2c.org/manual/manual_c.html
The real crowning use case for regular expressions (at least in terms of parsing-like tasks) is when you've got formats with varied delimiters. So one format I was parsing this weekend was basically header:field1,field2,field3"data"hash (with a fixed number of fields). Or another format I work with is suite~split/test1,test2@opt1:opt2^hw1^hw2#flags1#flags2 (most of the elements of which are optional). Regular expressions shine here; using primitives like split don't quite cut it in these formats.
And I think this also is a major cause of why regular expressions tend to become unreadable quickly. If you look at parsing via regular expressions as a whole, there's basically three things that are being jammed into one expression: what are the delimiters between fields, what is valid in each field, and which fields are optional. These are basically three separate concerns, but most regex APIs are absolutely terrible at letting you separate these concerns into separate steps (all you can do is provide the string that combines all of them).
In a language where regex is just right there, like Perl, I agree that the natural way to parse this is always a regex, although on the other hand in Perl the natural way to do almost everything is a regex so maybe Perl was a bad example.
In a language with split_once (and a proper string reference type) I actually rather like split_once here, splitting away header, and then field1, and then field2, and then field3, and then data, leaving just hash.
I guess this gets to your concern, by writing it with split_once we're clear about a lot of your answers, each delimiter is specified with what it's splitting, none of the fields are optional (unless we write that) and (if we write code to check) what is valid in each field.
let (y,m,d,h,m,s) = break_str!(s, year-month-day hour:minute:second)?;
instead of this: let (y, split) = s.split_once('-')?;
let (m, split) = split.split_once('-')?;
let (d, split) = split.split_once(' ')?;
let (h, split) = split.split_once(':')?;
let (m, s) = split.split_once(':')?; use regex::Regex;
fn main() {
println!("{:?}", extract("1973-01-05 09:30:00"));
}
fn extract(haystack: &str) -> Option<(&str, &str, &str, &str, &str, &str)> {
let re = Regex::new(
r"([0-9]{4})-([0-9]{2})-([0-9]{2}) ([0-9]{2}):([0-9]{2}):([0-9]{2})",
).unwrap();
let (_, [y, m, d, h, min, s]) = re.captures(haystack)?.extract();
Some((y, m, d, h, min, s))
}
Output: Some(("1973", "01", "05", "09", "30", "00"))
That gets you pretty close to what you want here.(The regex matches more than what is a valid date/time of course.)
>>> import re
>>> pattern = re.compile(r"([0-9]{4})-([0-9]{2})-([0-9]{2}) ([0-9]{2}):([0-9]{2}):([0-9]{2})")
>>> results = pattern.match('1973-01-05 09:30:00')
>>> results.groups()
('1973', '01', '05', '09', '30', '00')
>>> (y, m, d, h, min, s) = results.groups()
(`results` will be None if the regex didn't match)Regexes are one of those things where, once you understand it (and capture groups in particular) and it's available in the language you're working in, string-splitting usually doesn't feel right anymore.
There is a lot of reasonable room to disagree about when and where regexes should be used.
When using an interpreted language like Perl or Python, it's a slightly different story, since often there's no way to get as good of performance out of a custom parser written in that language as you can get from the optimized regular expression library included (if you use it competently). The overhead in the interpreted language can make some tasks impossible to do as quickly as teh regular expression engine can. In a way, a regular expression library is to parsing what Numpy is to computation.
I remember a task I had to parse a bunch of very simplistic XML in Perl a decade ago, where it consisted of the equivalent of a bunch of records of dictionaries/hashes (solr), and needed them as Perl hashes to work on them. I surveyed every XML parsing module available to me in Perl, and couldn't quite get the performance I needed, given the file was a few GB and it needed to be done as quick as possible multiple times a day. I implemented a simple outer loop to grab a working chunk of the file, a regex to split that chunk into individual records text chunks, and another regex to return each key/value as an alternating list I could assign directly to a hash. It was 5-6x faster than the faster XML streaming solution I found that I could use as a Perl module (which were all just wrapping X libs).
Could I have coded a custom solution in C that parsed this specialized format faster? Undoubtedly, but I'd still be marshaling data through the C/Perl interface, so would have taken a hit to get the data easily accessible to the rest of the codebase, which is less of an issue for regular expressions in Perl since they're a native capability of the language and deeply integrated. Using regular expressions as a parsing toolkit yielded very good results at a fraction of the time and effort, and honestly in my opinion that's what regular expressions are all about.
I think this exercise is valuable for anyone writing regexes to not only understand that there's less magic than one might think, but also to visualize a bunch of balls bouncing along an NFA - that bug you inevitably hit in production due to catastrophic backtracking now takes on a physical meaning!
Separately re: the OP, https://github.com/rust-lang/regex/issues/822 (and specifically BurntSushi's comment at the very end of the issue) adds really useful context to the paragraph in the OP about niche APIs: https://blog.burntsushi.net/regex-internals/#problem-request... - searching with multiple regexes simultaneously against a text is both incredibly complex and incredibly useful, and I can't wait to see what the community comes up with for this pattern!
If not, then perhaps this is a case where JavaScript beats Rust.
I would be interested in experimenting with a JIT in the future, primarily as a way to speed up capture group extraction. But this regex library uses a lazy DFA, which has excellent theoughput characteristics compared to the normal backtracking interpreter (and perhaps even a backtracking JIT).
> If not, then perhaps this is a case where JavaScript beats Rust.
Haha. No. Follow the link to rebar to see performance comparisons. A JIT isn't everything. https://github.com/BurntSushi/rebar
I dont have the stats of how often the other engines would be used, but if most of the programming languages out there are using PikeVM, then I can see why Google has not only written their own OS for their servers, but also saved a few more clock cycles by getting other engines into action for specific situations where the PikeVM would be too slow and/or heavy on the clock cycles. I know all to well how a few extra characters in the search string can drastically slow up the pattern matching.
The proverb "take care of the pennies and the pounds will take care of themselves" definitely applies to RegEx and clock cycles. I suspect when looking back at some conversations from the 90's, its made some coders I know very rich when dealing with millions of records per second processing.
These days, various web-based tools offer superior functionality. But in 2001, Komodo’s Rx Debugger was absolutely state of the art and so much fun to work on.
So should be able to burn something like [1] to a disk.
[1]: https://github.com/ibaaj/Regex101.com-offline-app/pull/1/fil...
https://flathub.org/apps/com.felipekinoshita.Wildcard
Haven't tried yet...
Let's say I have a list of login attempt dates, and I want all sequences of 5+ failed login attempts followed by a success. With a regex it's trivial, but I can't use that; I have to roll my own loop, flags, and temporary lists.
Sure, I could convert my list to a string, process it, and (try to) convert it back, but the downsides are obvious.
Even if the performance is not as good as string-based regex, why shouldn't we have regexes for arbitrary list types?
---
Edit: I realized this is not the first time I had this idea, and found an old Python prototype: https://github.com/boppreh/listregex
It's horribly slow, but as an API experiment I'm quite happy. It also gives some tools not available to regexes, like inverting and intersect patterns, and matched pairs.
# A sequence of 5 or more failed login attempts followed by a successful one.
obj_regex.search(repeat(is_failed, min=5)+is_success, my_list)
[1]: https://en.wikipedia.org/wiki/Parser_combinatorThe kind of regex engine you want, especially one where you don't care about perf, is not hard to build. You could take the `regex-lite` crate I published, for example, and code it up to be as generic as you want. In so doing, I expect you'll run into a number of interesting challenges.
Anyway, it's not like these things don't exist. They do. People attempt to build them[1]. They just usually don't gain much traction because I suspect you're overstating their general utility. :-)
[1]: https://docs.rs/automata/latest/automata/trait.Alphabet.html
Anyway, super appreciate your work and your writing!
Someone should go out and built it. It won't be me though, that's for sure. Believe it or not, such a library wouldn't leverage much of my experience. The vast majority of the complexity inside the regex crate is about optimizing for sequences of bytes.
let is_even = |x: &i32| *x % 2 == 0;
let is_odd = |x: &i32| *x % 2 == 0;
let pattern = sequence(&[
is_even,
is_odd,
repeat(or(is_even, is_odd), 2, 4),
]);
let mut matcher = Matcher::compile(pattern);
let elements = [4, 1, 23, 5, 6];
for element in elements {
matcher.step(element);
}
assert!(matcher.matches());
I'll get there one day.Amazing work on the Regex crate by the way!
Example pattern: https://github.com/icsharpcode/ILSpy/blob/1100d64e4bbd878164...
Implementation: https://github.com/icsharpcode/ILSpy/tree/1100d64e4bbd878164...
I find myself desiring a DFA over LLM tokens right now, which are 16 bits each, at least for this model.
I, personally, am very pleased with the idea of attempting to build a dense DFA with that vocabulary. Maybe your computer would explode!
Thank you, I'll look into it.
Last time I tried to figure out if there were existing tools to solve a problem like this, I came across Event Calculus: https://en.wikipedia.org/wiki/Event_calculus
I'm sure there's some interesting CS theory to be uncovered here.
I have no doubt that its utility can be great in niche use cases. I've never come across one in my decades of programming that I know of, but I'm sure they exist.
The login attempt example was convoluted, so here are more common scenarios:
- Grouping a list into pairs by matching /../
- Finding or collapsing repeated sequences with /(.)\1+/
- Parsing tag+value binary formats.
- Searching structured logs.
- DSL's for unit tests.
Are there better ways to do each of those? Yes, but either because someone implemented those specific functions, or it's a much longer solution.
Though I'm not recommending list-regexes for production code anytime soon. Prototypes and code golfing, sure, but not until the community understands it better.
But like I said, someone should go build it.
how can this be true? Not trying to pick any type of religious war, but if you had a C library based on char vectors, you could global replace with short or long and it would still work, pointers included, except for any places where you relied on the knowledge that a char was implemented as 8 bits.
or if you're saying "but I rely on the string class", is it somehow impossible to write a utf-32 based string class? would that string class be required to suppress anything that wasn't a valid code point at this particular time? I know professors like to teach that knowledge that a character is stored as an 8bit number is undefined behavior, but it sure is nice to know you have an 8-bit-clean data type, where 0 can in some contexts considered a terminator, and utf-8 is considered a special case on top of that.
I'm just trying to grok the type of impossiblity you're talking about here. Yes, I could read the code instead of asking :)
> but if you had a C library based on char vectors, you could global replace with short or long and it would still work, pointers included, except for any places where you relied on the knowledge that a char was implemented as 8 bits.
Right, and that last sentence is critical. Another way to say "rely on char being 8 bits" is "rely on an alphabet size of no more than 256."
The regex crate and all of its non-standard-library dependencies (of which I wrote all of them) corresponds to about 100K source lines of code. Everything from the concrete syntax on up to the matching engines itself would need to be parameterized over the alphabet. You can't just replace `u8` (unsigned 8-bit integer) with any other type and hope that it works, because the entire regex crate makes use of the fact that it's searching a sequence of `u8` values specifically. It doesn't treat `u8` as some opaque type with integer-like operations. The literal architecture of the code embeds assumptions about it, such as the fact that every possible value can be enumerated as [0, 1, 2, ..., 255]. The answer I wrote seven years ago touches on this. A concrete place where this occurs is in the lazy DFA (described in the blog). Very loosely, the logical representation of an individual state looks like this:
struct State {
next: [*State; 256],
}
The `u8` type appears nowhere there. Instead, the length of `next` is hard-coded to the total number of unique possible elements of the `u8` type. One could of course define this as part of some generic alphabet type, but this entire representational choice is only feasible in the first place because the alphabet size of `u8` is so small. What's the alphabet size of `u16`? Oops. Now your logical representation is: struct State {
next: [*State; 65536],
}
You can see how that would be quarrelsome right?Of course, you could use something like `HashMap<u16, *State>` instead. But holy moly you really do not want to do that when searching a sequence of `u8` values because a hashmap lookup for every byte in a string would absolutely trash performance. So now your generic alphabet interface needs to know about DFA state representation in-memory so that when you use `u8` you get the fast version and when you use anything else you get the slow-but-feasible version.
And then there's the whole UTF-8 automata stuff and various Unicode-aware algorithms sprinkled in at various places. This thread is talking about a generic alphabet, and that goes well beyond Unicode. So now all that Unicode stuff needs to be gated behind your generic alphabet interface.
At some point, you realize this is incredibly impractical, so if you want to expose a generic alphabet API, you do something like, "with a `u8` alphabet, go do all this fast stuff, but with any other alphabet, just do a basic simplistic generic regex engine." But that basic simpistic regex engine is going to wind up sharing very little with the `u8` specific stuff. And thus you get to my conclusion that it has no place in a general purpose regex engine designed for searching strings. For these niche use cases, you should just go build something for what you need and specialize for that. If you want a generic alphabet, then use that as your design criteria from the start and build for it. You'll wind up in a completely different (and likely much simpler) spot than where the regex crate is.
Basically, this line of thinking is a sort of "disease of abstraction" that is common in programming. We take what we think about the theory of regular languages and maybe some cute use cases and think, "well why not just make it generic!" But the thing is, coupling is generally how you make things fast, and making something more general usually involves de-coupling. So they are actually and usually exclusionary goals. But it takes a lot of domain knowledge to actually see this.
There's also a whole mess of code that uses low level vector instructions for accelerating searches. Thousands of lines of codes are devoted to this, and are just yet another part of the regex crate that doesn't really lend itself to being easily generic over arbitrary alphabets. You could probably write different versions of each of these algorithms for u8, u16 and u32 types (and that would in turn require using different vector operations because of the lane size differences), but they certainly aren't generalizable to arbitrary alphabets.
> I know professors like to teach that knowledge that a character is stored as an 8bit number is undefined behavior, but it sure is nice to know you have an 8-bit-clean data type, where 0 can in some contexts considered a terminator, and utf-8 is considered a special case on top of that.
I don't think regex engines have been tightly coupled to NUL terminated C strings for quite some time. Not even PCRE2 is. It's just a sequence of bytes and a length. That's it.
There does exist a gaggle of general purpose regex engines that, instead of working on UTF-8, works on UTF-16 code units. These are the Javascript, Java and .NET regex engines. None of them work on anything other than `u16` values. All of them are backtrackers. (Well, .NET does have a non-backtracking regex engine. I don't know its internals and how it arranges things with respect to alphabet size.)
PCRE2 does expose compile time configuration knobs for supporting UTF-8, UTF-16 and UTF-32. I'm not an expert on its internals, but PCRE2 is a backtracker and backtrackers tend to care less about the alphabet size than automata based engines such as the regex crate and RE2. But I'd guess that PCRE2 likely has some UTF-8 specific optimizations inside of it. They just don't show up as much (I assume) in fundamental aspects of its data structure design by virtue of its approach to regex searching.
Still, just supporting UTF-8, UTF-16 and UTF-32 is a very different thing than supporting generic alphabets.
> I'm just trying to grok the type of impossiblity you're talking about here. Yes, I could read the code instead of asking :)
It's really all about coupling and being able to make assumptions that the coupling grants you. If you remove the coupling---and that's what abstraction over arbitrary alphabets actually is---then you have to either remove the assumptions or push those assumptions into the abstraction. And whether that's even possible at all depends on the nature of how those assumptions are used in the chosen representations.
This is why Unicode is such a beastly thing to support in a regex engine. It runs roughshod all over ASCII assumptions that are extremely convenient in contexts like this. Something like `\w` that is ASCII aware can be implemented with a single DFA state. But as soon as you make `\w` Unicode-aware, you're now talking about hundreds of DFA states when you're alphabet is `u8`. You could switch to making your alphabet be the space of all Unicode codepoints (and thus collapse your DFA back down into a single state), but now your state representation basically demands to use a sparse representation which in turn requires more computational resources for a lookup (whether it's a hashmap or a linear/binary search). If your alphabet size is small enough to use a dense representation, then your lookup is literally just a pointer offset and dereference. That operation is categorically different than, say, a hash lookup and there's really nothing you can do to change that.
The bottom line here is that if you want a regex engine on a generic alphabet, then you probably don't care about perf. (As folks in this thread have come out and said.) And if you don't care about perf, it turns out you don't need something like the regex crate at all. You'd be able to build something a lot simpler. By at least an order of magnitude.
Why hasn't someone built this "simpler" library? Well, they have. I've linked to at least one elsewhere in this thread. It just turns out that they are often not productionized and don't catch on. My thesis for why is, as I've said elsewhere in this thread, that the utility of this sort of generic support is often overstated. I could be wrong. Maybe someone just really needs to do the work to make a great library. I doubt it, but it's not really a question that I want to put the work into answering.
[1]: https://old.reddit.com/r/rust/comments/4a8pbv/the_regex_crat...
E.g. it would be very nice to derive doing unicode regexes on u8 (as you say) from the serial composition of `user_regex . unicode_regex`. Fusing that composition into a single automaton is a handy generic thing to have in the library bag of tricks.
I am not disagreeing with your assessment, I am saying a better world is possible where these hurdles are not so insurmountable.
What you're talking about is categorically different. Your idea is so out there that I can't even begin to dream up a programming language that would resolve all of my concerns and challenges with building a general purpose regex library with Unicode support that is one of the fastest in the world, supports arbitrary alphabets and is actually something that I can maintain with reasonablish compile times.
IMO, I'm the optimist here and you're the pessimist. From my perspective, I see and celebrate what Rust has allowed me to accomplish. But from your perspective, what you see (or what you comment on, to be more precise) is what Rust has supposedly held me back from accomplishing. (And even that is something of a stretch, because it isn't at all clear to me that programming language expressivity and the concept of "zero cost abstractions" is arbitrarily powerful in the first place. See, for example, the premise of the 2014 film "Lucy".)
I wouldn't describe the difference between our views here as optimism vs pessimism. I like Rust decently enough to use it a fair amount and contribute to it somewhat. But I'm hungry for something more. It's fine if you aren't.
> But the thing is, coupling is generally how you make things fast, and making something more general usually involves de-coupling. So they are actually and usually exclusionary goals. But it takes a lot of domain knowledge to actually see this.
The biggest difference is that I don't believe this is true. The interface may be extremely gnarly and hard to express, but I do believe the interface can always in principle be sliced. It may not be practical (the interface is far more complicated than the thing before cutting) but it is still in principle possible.
---------------------
Let me put it this way, imagine some alternative world Unicode which was pathological in all the same ways (case, variable width encodings, all the other things I am forgetting). Imagine if someone forked your work to handle that. Clearly there would be lots of small things different --- it would be hard to abstract over --- but the broad brushstrokes would be the same?
I am not interested in supporting things which are not like Unicode, agree stray to far towards large alphabets etc. and all the algorithms one would want would be different anyways. But supporting things like Unicode or strictly easier than Unicode should be doable in some language, I think.
Well "practicality" matters. And I used weasel words such as "generally."
It's one thing to make concrete suggestions about how to actually achieve the thing you're looking for. But when you're way out in right field picking daisies, it just comes across (to me) as unproductive.
> but the broad brushstrokes would be the same?
This is basically impossible to answer. I could just as well say that the "broad brushstrokes" are the same between the `regex` and `regex-lite` crates.
Unfortunately performance goes out of the window. And it's likely to work as well as storing arbitrary non-character things into a custom std::basic_string.
However, trying to use and read through regex-automata and regex-syntax (even in 2018) was such a very valuable and helpful learning resource. I ended up modeling the project at work off Lucene APIs, but it was all only after learning the basics from the regex crates.
Best set of libs! Thanks sushi!
Future work is definitely trying to make the regex engine scale better with more patterns. Currently, it will crap out well before 10 million regexes, and I'm not sure that goal is actually attainable. But it can certainly be better than what it is today.
And of course, Hyperscan is really the gold standard here as far as multi-pattern searching goes. I'm not sure how well it will handle 10 million patterns though.
Quick notes: to scale past X regexes in the set.
1. Don’t make a single RegexSet. Create multiple internally.
2. Ideally colocate similar partners together.
3. Use the immutability of a RegexSet to your advantage by generating an inverse index of “must match” portions of each regex.
4. When you get an input to match, first check the inverse index for a “may match set”, then only execute the internal automata that may possibly match on the input.
The inverse index can be as complicated (FSTs, or all of Lucene) or simple (hash map ) as you’d like it to be. The better it can filter the regexes in the set the more performance and scale you end up with. And tune the sizes of the internal automata for your usecase for some Goldilocks size for extra perf.
There were many different strategies at play (eg neural nets) but the false positives can be very expensive. But the one I described above is an expert system of manual overrides that made the final decisions.
Amazon don’t have much structured content, eg “title”, “description” and bad actors are constantly trying to obfuscate using language based grammar tricks or N1ke type attacks.
Also just the law is very nuanced. You can say “a case compatible with iPhone, that is ..” but it is infringement to say “iPhone case that is ..”
Also there are many nuances around licensing. For example, “BurntSushi T-shirts” as a company cannot have the Disney logo on it, however, a “Lego tshirt” might have the Disney logo legally on it. So a lot of overrides.
Even if you compress all the nuances into a single regex, which you can’t: there are 10 MM estimated brands worldwide (Amazon at the time hosted 1 MM) with on average 10 trademarks each that is 10 MM regexes.
Then you also need to multiple the 25 countries Amazon operates in for legal nuances, and multiple languages (eg even in the US Amazon store someone might use Spanish or Italian to sell their fake product). Additionally there is more fanout for reasons I’ve long forgotten.
Point of story 10 MM was a conservative upper bound on the RegexSet and it had to be fast and cheap. I built it on top of Lucene with some hacks here and there but it did (does) the job - fast and cheap :)
I hope now generative AI can help them more, but I’m not holding my breath.
You cant, standalone hw isnt that capable, distributed computing/cluster computing is.
> I hope now generative AI can help them more, but I’m not holding my breath.
I'm not from what I have seen.
>Also just the law is very nuanced. You can say “a case compatible with iPhone, that is ..” but it is infringement to say “iPhone case that is ..”
And Amazon would be up against users of Lexis Nexis who have a head start in all of this RegEx malarkey, by virtue of what they sell, without giving too much away about their code inner workings.
(I haven't read the post yet because I have an important call in a few minutes, but it looks like a very interesting and also conveniently-timed blog post.)
(Edit a few minutes later: looks like the answer might be yes, but since this is a polished release, I might be able to simplify my code massively. Wish me luck~)
(Edit 2 about 10 minutes later: well that was pretty painless and the new Builder::patch method is a total upgrade. Awesome~)
P.S. I'm still blocked from all your GitHub repositories and I think that's kind of unfair considering how widespread a lot of your crates are. I don't remember the original incident anymore. I believe the regex crate itself is under the rust-lang organization now, but there are still others I can't interact with.
> I'm still blocked from all your GitHub repositories and I think that's kind of unfair considering how widespread a lot of your crates are. I don't remember the original incident anymore.
I don't either. I block a lot of people for a lot of reasons. I unblocked you.
I didn't listen, and it paid off, because a bunch of the knowledge I learned on it is still perfectly applicable to 0.3.0; I already have the same NFA as before. :)
Now my appointment got canceled so I get to read your blog post and play with the new version. Fun!
Not to take away from Rusts regex which are far more advanced than Automa, but I just take issue with that library being the first time regex internals have been exposed as a library :)
What I'm talking about in this blog is very different from that. This is about taking the internals of the regex library itself and turning them in a separately versioned library that others can compose.
This makes somewhat less sense for a backtracker because most backtrackers just have one engine: the backtracker. But an automata based library usually has many different engines that can be composed in different ways. Still, backtrackers have things they could expose but don't in practice, such as a regex parser and an AST.
My time horizon is very long. It takes me a long time to do things these days.
It has never been true that I don't want to support it. Merely that it is difficult to verify and test if I can't do it myself. There is also the problem that the port from x86 to arm is not straight-forward, due to both my own ignorance and what I believe are important missing vector operations such as movemask.
This is discussed a bit more here (including the bit about movemask): https://github.com/BurntSushi/memchr/issues/76
Tangential (since he answered already), but in this era of cheap cloud instances, I'd be surprised if "not owning a <blank>" is any hindrance to support.
When you own a device, it doesn't cost you more (other than electricity) if you take longer while playing with it, and there's no extra cost (other than space) to having it ready to be used when necessary. When you rent a cloud instance, there's a direct monetary cost for every extra minute you take while investigating an issue, and it's less guaranteed that it'll be available when you need it.
Or, to put it more simply: unless he has his own personal device to play with, he'll have to spend extra money for every issue he investigates and every change he makes to the ARM support.