A New RegExp Engine in SpiderMonkey
hacks.mozilla.org
hacks.mozilla.org
The regexp grammar portion of the ES spec is different, in that it's not sufficient. You can implement JavaScript from the ES spec, but NOT regexp. The grammar doesn't really make sense [2]. The Chakra team found the same thing [3].
I have not confirmed this but I have been told the regexp portion is tightly controlled by Google, more than most parts of the spec. When v8's regexp implementation is found to deviate from the spec, the spec gets changed. This makes sense (why risk breaking websites), but it illustrates how little weight the regexp spec actually has.
With Firefox using irregexp, it further cements this unfortunate reality.
2: Example: the NonemptyClassRangesNoDash production may include a dash
3: https://docs.microsoft.com/en-us/archive/blogs/ie/chakra-int...
I haven't tried to implement the RegExp section, so I can't speak to that, but this is largely why we have Test262[1].
Also that Chakra blogpost is from 10 years ago, talking about ES5. Things have improved considerably since then, though of course we have to live with historical baggage.
> With Firefox using irregexp, it further cements this unfortunate reality.
Is this a specific unfortunate reality? I don't spot any RegExp entry in the last spec Incompatibilities section[2] that seems like a bad change or even having to do with regex grammar at all.
[1] https://github.com/tc39/test262/tree/master/test/built-ins/R...
[2] https://www.ecma-international.org/ecma-262/#sec-additions-a...
Some things have indeed improved. The 'u' flag introduced in ES6 provides a much cleaner grammar for example. However the process is still busted.
Here is a concrete example. /{1}/ is permitted by the ES6 spec (including Annex B). But in practice, "web browsers" rejected it. Rather than fixing the browsers, the spec was changed in ES8 to reflect "web-reality". https://github.com/tc39/ecma262/pull/303
Browsers could have become ES6 compliant here, with zero compatibility risk. But when there is non-compliance, it is the spec that moves.
A few responses:
1. This was not at all my experience working with the V8 team. For example, when I noticed that Irregexp did not properly implement the Canonicalize algorithm for non-unicode case-insensitive comparisons [1], they were happy to accept my patches [2][3]. Nobody ever suggested changing the spec.
2. Nothing changes in the JS spec without consensus from all the major stakeholders. See, for example, the second sentence in the TC39 process document here [4]: "The committee operates by consensus and has discretion to alter the specification as it sees fit." Mozilla can block (and has blocked) proposals that it feels are bad for the web platform. Google has no special power to bend the committee to its will.
3. Another way of saying "when V8's regexp implementation is found to deviate from the spec, the spec gets changed" is "when the spec is found to deviate from the consensus of implementations, we correct the spec". The Chakra team's post notes: "In practice, all browsers accept regular expressions such as `/]/` and web developers write them." The purpose of the spec is to make the web work better by maximizing compatibility between implementations. In cases where all implementations agree, and real websites depend on that behaviour, the spec should change to match reality.
Don't get me wrong: Chromium monoculture is a real problem. If we thought that Google was going to add a bunch of new regexp features without going through the standards process, or if we had plans to prototype regexp features that we feared Google wouldn't accept upstream, we might have made a different decision. But regexps are not a plausible battleground, and the compatibility benefits of sharing code outweigh the harms.
PS: regress looks really sweet, and I am excited to see how it turns out.
[1] https://tc39.es/ecma262/#sec-runtime-semantics-canonicalize-... [2] https://github.com/v8/v8/commit/3fab9d05cf34a7f0bc0e9405729a... [3] https://github.com/v8/v8/commit/b65fcfe92566415c8aae373d5fb5... [4] https://tc39.es/process-document/
First, I confess to some skepticism of the standardization process. For example, consider ES2018 lookbehinds. At the point of standardization, only one browser implemented them: v8. And in the standardized form (arbitrary width) they are very difficult to retrofit onto existing engines. Moddable had to throw their engine out and start over [1]. JSC's YARR still doesn't support them, neither does SpiderMonkey. Did all stakeholders really agree on this feature and then not implement it? Or was this just v8 in the driver's seat?
Second, regexps actually do have a role in driving monoculture. You can't polyfill regexp syntax. At one point, Steam's website only worked on Chrome because only v8 implemented lookbehinds [2] - again because it's hard to retrofit.
Lastly, it remains a problem that one cannot implement a conforming JS regexp engine from the spec alone. The rest of JavaScript does not have that problem, only regexp. And I disagree with the "real websites depend on that behaviour" qualifier. No website depended on /{1}/ being invalid syntax, yet it was made so in ES8.
Incidentally, it's amusing that v8's Canonicalize was buggy. Non-Unicode case-insensitive char ranges is THE ugliest part of regexp (QuickJS doesn't even try [3]) and I just assumed it was codifying v8's implementation. Guess it was someone else's, hah.
1: https://blog.moddable.com/blog/regexp/
2: https://github.com/webcompat/web-bugs/issues/51385
3: https://github.com/ldarren/QuickJS/blob/mod/libregexp.c#L259
They reached that point because: a) everybody agreed that they were a good feature, with precedent in other languages; b) the spec text had been scrutinized by the people who like to spend time scrutinizing spec text; and c) two implementations (Moddable and V8) had successfully implemented them. That's the process, and it was followed to the letter.
Does that put a burden on other engines to keep up? Sure. But that's what we signed up for. The webcompat bug you linked is from April 2020, more than two years after lookbehind assertions were standardized. That's nobody's fault but ours, and we did this project so that we don't end up in a similar situation in the future.
I'm not sure what your source is for JS regexps being impossible to implement from the spec, but if you have specific points of concern, you should open an issue [6]. Writing a spec is hard work, and things get missed, but the people working on this are genuinely trying their best.
(As far as I can tell, Canonicalize is what everybody thought they were implementing. V8 tried to get clever with ICU and ran facefirst into some of the dark corners of Unicode.)
1: https://github.com/tc39/notes/tree/master/meetings
2: https://github.com/tc39/notes/blob/master/meetings/2016-11/n...
3: https://github.com/tc39/notes/blob/master/meetings/2017-03/m...
4: https://github.com/tc39/notes/blob/master/meetings/2017-05/m...
5: https://github.com/tc39/notes/blob/master/meetings/2018-01/j...
Second, I cast no aspersions against any of the people involved in the ES spec. I recognize they are engaged in hard and unrewarding work driven by a sincere effort to push JavaScript forward, while everyone's a critic, especially myself. It's awesome that the meeting minutes are available. Thank you for the links.
Third, I honestly believe the meeting minutes support my point. I risk overstepping, because I had no involvement, but from your third link:
> We have two implementations. One in V8 and one in the Dart VM. It's not a JS implementation, but it does implement this feature.
This supports my speculation that it's all about v8. Later:
> I would assume Chakra or V8 would have said something by now if they had issues
SpiderMonkey and JSC are chopped liver? This feature really seems to have been driven by the implementors.
And this one feature (lookbehinds) was so disruptive, so damn hard to implement, that Mozilla abandoned their implementation and now just uses v8's. The linked blog post cites this feature as an impetus to switch.
Was that the goal? Did SpiderMonkey engineers give the nod to lookbehinds with the intention of abandoning their engine in favor of irregexp? I have to believe not; I speculate it was the very human response of conflict avoidance. Easier to say "yes" especially if you do not appreciate how much work you are signing up for. But it means ES2018 regexp advanced the monoculture. One less regexp engine in the wild.
Fourth, I really do know that the ES regexp spec is not useful for implementors, because I am an implementor. Ok, it's not really a question, huh; I should just make a PR.
On the other hand, the more dependent Mozilla is on the Chromium base, the more power it gives Google (even if at this point Google already acts like they own the internet and do things that break in FF at will).
https://www.rand.org/content/dam/rand/pubs/research_memorand...
It's incredible that biology served as inspiration for this.
https://github.com/DmitrySoshnikov/regexp-tree
A fun bug I found in it: https://github.com/DmitrySoshnikov/regexp-tree/issues/69
[regexp] Remove trivial assertion
The assertion in BytecodeSequenceNode::ArgumentMapping cannot fail,
because size_t is an unsigned type. This triggered static analysis
warnings in SpiderMonkey.
Does that mean Chromium doesn't use any static analysis tool? Or one that does not work for this trivial assertion?Or one that does work for this trivial assertion, which is to say not emit a diagnostic.
What ends up happening in a situation like this is that somebody removes the assertion to quiet the compiler diagnostic. Then the types change and all of a sudden it can fail, but the assertion is gone.
You might say, "well, now that one of the operands is size_t how could the types ever change to reintroduce an issue?" That's the wrong question to ask. The whole point of the assertion is to avoid having to answer such a tricky question, or at least to encode your answer in a way that if the unforeseen happens things break loudly instead of silently. Anyhow, you'd be surprised by how things can break. I personally religiously use size_t for anything related to object size, but many other developers don't (including at Google and Mozilla), and so you often see a mix of, e.g., size_t and uint64_t, size_t and uint32_t, size_t and int, etc, and regular tweaks back-and-forth, which can easily introduce regressions.
I understand why compilers emit a warning--the idea is that if the assertion couldn't possibly be false, maybe it has a bug. But, IME, the opposite is usually true--it's well-written and deliberate, because the developer is trying to catch spooky action at a distance where the type, which is defined far away, is changed, accidentally or intentionally. I don't know where to draw the line in terms of second-guessing the code to help catch bugs, but GCC and clang need to provide far more succinct constructs to tell the compiler to shut up. Currently you're stuck with inline pragmas, __extension__, statement expressions, and other weird convolutions that require far more code than the assertion or operation itself. The issue makes writing arithmetic overflow-safe code more tedious and error prone.
(Other languages simply prohibit mixing integer operands of different types, so if a type is changed far away then code will break loudly even without any assertions. But in codebases like SpiderMonkey and V8 that use a wide range of integer types for space and performance optimizations, that tends to encourage casting, which has the exact same problems.)
Asserting things which are given from the types is a waste. Or do you think we should assert that a value of type bool is really either true or false and fail the program otherwise?
I love Rust. I've lost track of the number of times I've stared at a chunk of C++ code, considered how much nicer it would look in Rust, and sighed. At the beginning of this project, we did consider whether it was the right place to use Rust.
There are no existing JS-compatible regexp crates. BurntSushi's regex crate is great, but finite automata don't do backreferences, which is a dealbreaker for JS. After this code landed, somebody brought https://github.com/ridiculousfish/regress to my attention. It looks promising, but still has a long way to go before it's production-ready, and it didn't exist when we made our decision.
If we had written our own replacement, it would likely have been Rust. SpiderMonkey has a lot of cross-cutting issues (GC especially) that make it hard to replace individual C++ components with Rust, but the regexp engine has a pretty clean API boundary. It's the same reason that it was feasible to swap in Irregexp in the first place.
Ultimately we decided that writing a new engine wasn't the best use of time. A regexp engine is a complicated beast. Writing a new high-performance, JS-compatible engine in Rust could have been person-years of effort, with a long tail of corner cases and performance issues. We haven't had many memory safety bugs in regexp code. As sstangl points out in a sibling comment (hi Sean!), doing JIT compilation undermines some of Rust's safety guarantees.
When it comes down to it, the regexp engine is not a place where SpiderMonkey is looking to push the state of the art. We have to be reasonably fast and feature-complete, but beyond that nobody is going to notice marginal gains. There are higher-leverage opportunities elsewhere.
And the partial memory safety over the metadata around the actual JITed code is a big win as well.
Yes, Rust doesn't get you there 100%, but IMO it gets you closer than C or C++.
C++ has both, too, so in this case Rust's only advantage would be memory safety.
Rust gives you that more or less by default and for free wrt tooling. It's sort of the classic "Rust makes you write the C++ you should have been writing all along", which makes it a net win IMO.
C++'s ADTs are easy to subvert even accidently; Rust's can't be without explicitly calling it out as unsafe.
It allows you to describe transformations of state in a formal set theoretical way. You should check out formally verified software like CompCERT and sel4 and their heavy use ADTs internally to achieve that. Rust obvs isn't full formally verified but it's a neat 80/20 in that direction.