Regular expressions are powerful, and are used in plenty of production software.
* RegEx is not very readable
* RegEx can be (very) slow
* It's not trivial to write RegEx code that achieves your goal in a high-quality way. Often quircks and edge-cases are missed.
I'm not saying that you should never use them, but oftentimes a (much) better alternative of achieving your goals is present.
See https://blog.codinghorror.com/regular-expressions-now-you-ha...
(I'm not sure what you mean by DoS attacks - are you referring to the exponential case of backtracking? If so, don't use a regex engine with that problem, and don't use lookbehind/lookahead assertions, which aren't needed to solve this problem.)
No, it is not a good solution to the problem; you ignore my earlier comments. English or latin text is not comprised of the sole ASCII characterset; it contains characters outside this set (quoting other languages, names, imported words for example).
Honestly, I do get your point about inappropriate use of regex, but this kind of simple text manipulation is well suited for regex. The biggest argument against using regex for this kind of problem is performance verses writing the same code programmatically in the host language (assuming you're using a fast AOT compiled language). However even that is a non-issue given the small quantities of text you're decoding.
Also I'd bet the regex in this instance would actually work out more readable because the transformations are basic so you're localising the text manipulation to simple rules rather than multiple lines of byte array reading and thus also potentially having to manually build in your own rudimentary unicode support too.
If you can do the same manipulation in three ones of code it’s more likely to be correct and stay correct. And when you look at it again in six months you won’t have to stare at it. All those little time sucks add up as the code grows.
Edit: autocorrect got me twice.
Firsly, for relatively simple regex expressions, any competent developer should be able to grok them very quickly - at least as quick as the equivelant C#/Java/whatever code.
Secondly, it may be that regex is the most performant solution, and sometimes that matters quite a bit.
Honestly, I just don't get why some people are intimidated by regex.
1/ the regex engine
2/ host language
3/ problem you're trying to solve
But I've found generally I was better off not using regex for performance critical code.
HOWEVER (!!!) where regex consistently wins is development time. Not just writing the code, but testing (it's trivially easy to test regex) and updating the pattern matching (Vs updating the equipment character matching in an imperative language).
Yeah regex can get ugly quickly, but then so can any language if misused.