The original implementation was grotesquely inefficient, it's a 172kB file! To fix that, they ended up manually inlining and duplicating 70 lines of code. And the optimized version is still 50x slower than grep --color would be.
Why was the original version so slow? Can't be due to the searching, given they were 4 orders of magnitude off from where they should be. It can't be due to this kind of batch oriented programming having more efficient memory access patterns. The data set is so tiny that it's all going to fit in caches. Something in the editor's internal data structures working better when the stages are not interleaved? That would be my first guess. Maybe making changes to the editor data structures invalidated the search state, forced each search step to restart from the start of the buffer, and you ended up with an accidentally quadratic algorithm?
There's clearly a real performance problem in their system that was papered over by copy-pasting some code, but that they'll inevitably keep hitting over and over as people write the obvious code, it seems to work fine, but then has bad performance in practice. The right thing to do is to either fix that underlying problem, or to make it as easy to write the well-performing version as the bad one. (E.g. have some kind of scoped "batch context" abstraction, which queues up the edits and then applies them in a batch when leaving the scope.)
Anyway, all of this makes the conclusion quite odd. The real moral of the story should be that CPUs are really fast, not that Rust is magical and will change the way one thinks.