The dangers of single line regular expressions
greg.molnar.io
greg.molnar.io
Obviously if you pass data into the variable side of the engine, you hardly have to worry about it at all, since it's already going into a place that was designed for handling arbitrary and possibly-hostile input and been battle-tested at doing it correctly in Production for many years. If you pass it into the template side, you're betting that you can be as good as dozens of templating engine writers working for a decade at doing that, in exchange for, well, I can't really think of any possible legitimate advantage for doing that.
Do it in a sandbox and have aggressive timeouts.
Not practical given a large amount of documents.
> Do it in a sandbox and have aggressive timeouts.
Sure! I was just replying to this:
> the danger of passing anything derived from user input into the TEMPLATE side of a templating engine. Why in the world would you ever do that?!?
Ruby seems to be in multiline mode all the time?
$ python -c 'import re; print "yes" if re.match(r"^[a-z ]+$", "foobar") else "no"'
yes
$ python -c 'import re; print "yes" if re.match(r"^[a-z ]+$", "foo\nbar") else "no"'
no
$ python -c 'import re; print "yes" if re.match(r"^[a-z ]+$", "foo\nbar", re.M) else "no"'
yes
$ perl -le 'print "foobar" =~ /^[a-z ]+$/ ? "yes" : "no"'
yes
$ perl -le 'print "foo\nbar" =~ /^[a-z ]+$/ ? "yes" : "no"'
no
$ perl -le 'print "foo\nbar" =~ /^[a-z ]+$/m ? "yes" : "no"'
yes
$ node -e 'console.log(/^[a-z ]+$/.test("foobar") ? "yes" : "no")'
yes
$ node -e 'console.log(/^[a-z ]+$/.test("foo\nbar") ? "yes" : "no")'
no
$ node -e 'console.log(/^[a-z ]+$/m.test("foo\nbar") ? "yes" : "no")'
yes
$ ruby -e 'if "foobar" =~ /^[0-9a-z ]+$/i then puts "yes" else puts "no" end'
yes
$ ruby -e 'if "foo\nbar" =~ /^[0-9a-z ]+$/i then puts "yes" else puts "no" end'
yes
EDIT: this is documented behavior for Ruby. What other languages call multiline mode is the default; you're supposed to use \A and \Z instead. They do have an `/m` but it only affects the interpretation of `.`https://docs.ruby-lang.org/en/master/Regexp.html#class-Regex...
I think the simplest fix would be to use "\Z" rather than "$", which means "match end of input" rather than "end of line." This is also Perl-compatible. So weird that the "$" default meaning is different in Ruby.
I guess one could argue that Ruby's way is better since "$" has a fixed meaning, rather than being context-dependent.
> Ruby seems to be in multiline mode all the time?
Ruby does have a "/m" for multiline mode, but it just makes "." match newline, rather than changing the meaning of "$", it seems.
[1] https://ruby-doc.org/3.2.2/Regexp.html#class-Regexp-label-An...
if !m=/^[a-z0-9 ]+$/match(str) return "Bad Input" end str=m[0]
$ python3 -c 'import re; print("yes" if re.search(r"^foo$", "foo") else "no")'
yes
$ python3 -c 'import re; print("yes" if re.search(r"^foo$", "foo\n") else "no")'
yes
$ python3 -c 'import re; print("yes" if re.search(r"\Afoo\Z", "foo") else "no")'
yes
$ python3 -c 'import re; print("yes" if re.search(r"\Afoo\Z", "foo\n") else "no")'
no
Even if the newline is not problematic, using \A and \Z makes your intentions clearer to the reader, especially if you add re.X and place comments into the pattern.Asides:
1. Based on syntax, you appear to be testing with python2.
2. With python, re.match is implicitly anchored to the start, so the ^ is redundant. Use re.search or omit the ^.
If so, it is not "matching the end of a string" at all. Just end of line. That's exactly as expected in single-line mode, so it's good. May mismatch your expectations in multi-line mode though.
A $ does mean end-of-string in Javascript, POSIX, Rust (if using its usual package), and Go.
I'm working with the OpenSSF best practices working group to create some guidance on this stuff. It's a very common misconception. Stay tuned.
If anyone knows of vulnerabilities caused by thus, let me know.
I wrote a compiler from Java regex to JavaScript RegExp, in which you'll find that particular compilation scheme [1].
Edit: also quoting from [2]:
> By default, the regular expressions ^ and $ ignore line terminators and only match at the beginning and the end, respectively, of the entire input sequence. If MULTILINE mode is activated then ^ matches at the beginning of input and after any line terminator except at the end of input. When in MULTILINE mode $ matches just before a line terminator or the end of the input sequence.
[1] https://github.com/scala-js/scala-js/blob/eb160f1ef113794999...
[2] https://docs.oracle.com/javase/8/docs/api/java/util/regex/Pa...
> If MULTILINE mode is not activated, the regular expression ^ ignores line terminators and only matches at the beginning of the entire input sequence. The regular expression $ matches at the end of the entire input sequence, but also matches just before the last line terminator if this is not followed by any other input character. Other line terminators are ignored, including the last one if it is followed by other input characters.
Looks like I have some code to fix.
[1] https://docs.oracle.com/en%2Fjava%2Fjavase%2F21%2Fdocs%2Fapi...
I don't think Java should be in your first list, though? Pattern.matches("^foo$", "foo\n") returns false.
If that's true, then I fear the answer for Java may vary. The O'Reilly book on Regular Expressions, and the JDK documentation for version 21, say clearly that $ permits an optional \n at the end. The Java 8 documentation is murky, and maybe Java 8 is different.
$ php -v
PHP 8.2.7 (cli) (built: Jun 9 2023 19:37:27) (NTS)
$ php -r 'var_dump(preg_match("/^[a-z0-9 ]+\$/", "hello world"));'
int(1)
$ php -r 'var_dump(preg_match("/^[a-z0-9 ]+\$/", "hello\nworld"));'
int(0)
$ php -r 'var_dump("hello\nworld");'
string(11) "hello
world"
...
$ php -v
PHP 7.2.26-1+0~20191218.33+debian8~1.gbpb5a34b (cli) (built: Dec 18 2019 16:09:52) ( NTS )
$ php -r 'var_dump(preg_match("/^[a-z0-9 ]+\$/", "hello world"));'
int(1)
$ php -r 'var_dump(preg_match("/^[a-z0-9 ]+\$/", "hello\nworld"));'
int(0)
$ php -r 'var_dump("hello\nworld");'
string(11) "hello
world"
I'm not sure which version of PHP had the behavior you describe, or whether it misbehaves under more specific conditions, but preg_match() is one of the more commonly-used regex functions, all of which share the same engine. The behavior here seems to be "correct" for at least the last 5 years, for varying interpretations of "correct".edit: https://3v4l.org/N4o8D suggests that the behavior here is identical for all versions of PHP from 4.3 to 8.3.6.
This is easily demonstrated with an example.
<?php
var_dump(preg_match("/^[a-z0-9 ]+\$/", "hello\n"));
int(1)
versus <?php
var_dump(preg_match("/^[a-z0-9 ]+\$/", "hello\nworld"));
int(0)
versus <?php
var_dump(preg_match("/^[a-z0-9 ]+\$/", "hello\n\n"));
int(0)
Which all makes sense, as by default PHP doesn't operate in multiline mode. So, by default, PHP is not going to fall prey to the same problem being discussed here. In addition, the first \n would be apart of the first line it's on, so including it as a part of the string would make sense. More to the point, in this context, $ does mean end of the string in PHP. You can prove otherwise by getting the 2nd and 3rd example above to output a 1 instead of a 0 without going into multiline mode.In PHP, the following is considered true:
> var_dump(preg_match("/^[a-z0-9 ]+\$/", "hello\n"));
That is clear proof that "$" does NOT just match the end of the string; it also accepts an extra newline at the end of the string. In PHP you need to use \z if you want to match the end of the string, or use the "D" flag when using "$".
That definition of "$" is often reasonable when you read files a line-at-a-time from a file, which is why Perl changed its definition. However, PHP is often used for server-side web applications. In this case, you are often NOT reading a line-at-a-time from a file. In such cases, allowing an extra newline at the end could be disastrous. The MediaWiki code (written in PHP) deals with this by adding the "D" flag when it uses "$", but I'm not sure it always uses it, and I doubt all PHP programs use this flag when they should.
(\Z allows a trailing newline, \z does not)
Even better, compare that to the original and fail validation if they're not identical, but that requires maintaining a higher level of paranoia than may be reasonable to expect.
Parse don't validate https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-va...
But philosophically I agree, that's exactly the relevant advice.
What would work is having a small object holding a readonly string which parses the original on creation, then becomes immutable.
There's a related concept of "failing open vs failing closed" (fail open: fire exit, fail closed: ranch gate)
In Jurassic park (amazing book/film to understand system failures), when the power goes out, the fence is functionally an open gate
In this case, we shouldn't assume that we can enumerate all possible bad strings (even with a regex)
But I guess this is why Python has so many ways of matching a pattern against a string (match, find, findall, I think - they are hard to remember)
This will guarantee that you’re safe no matter how a piece of content is used tomorrow (just need a new escaping function for that content type), and prevent awkward things like not letting users use “unsafe” strings as input. JSX and XHP are example templating systems that understand context and escape appropriately.
If a user wants their title to be “hello%0a%3C%25%3D%20File.open%28%27flag.txt%27%29.read%20%25%3E”, so be it.
Use input validation / parsing to ensure data types aren’t violated, but not as an output safety mechanism.
that's a good way to horizontally propagate/reflect XSS and other Code As Data vulnerabilities.
better to strip the known-bad/problematic characters
Ku' 'Laangah't is valid.
When the user hands you a string and you then pass this down to other bits of code, you can't know if it will be used in an SQL query, a regex, in an error message that will be rendered into HTML, etc.
Ideally all layers of your code would handle user input with the utmost care, but that is often very hard to achieve. If you take user input and use it in a regex, it's easy to regex-escape it, but it's much harder to remember that now this whole regex is user input and can't be safely used to, say, construct an SQL query. And even if you remember to properly escape it in the SQL query, it may show up in the returned result, and now if you display that result, you need to be careful to escape it before passing it to some HTML engine.
But then none of this works if you did intend to have some SQL syntax in the regex, or some HTML snippets in the DB: you'd need to make all of these technologies aware of which parts of the expressions are safe and which are tainted by user input.
And this is all just to prevent code injection type attacks. I haven't even discussed more subtle attacks, like using Unicode look-like characters to confuse other users.
I guess I would tweak my first comment and say input filtering is not enough. You must do output filtering to truly be safe.
It should be just fine to pass in any character to the site, so a regex deny list is the wrong approach.
It's also how you end up with apps that people can't use because they reject their perfectly valid legal name, address etc.
For example, even if your name is officially 鳥山 in Japan, you will have to spell it out as Toriyama when you leave Japan, both in formal and informal settings, on paper just as much as in electronic forms, since no one would be able to understand it otherwise. And similarly, if your name is Smith, in Japan you will often have to spell it (and sometimes even pronounce it) スミス.
~ > raku -e 'say "foobar" ~~ /^ <[a..z ]> +$/ ?? "yes" !! "no"'
yes
~ > raku -e 'say "foo\nbar" ~~ /^ <[a..z ]> +$/ ?? "yes" !! "no"'
no
~ > raku -e 'say "foo\nbar" ~~ /^^<[a..z ]>+$$/ ?? "yes" !! "no"'
yes
- ^^ and $$ are the raku flavour of multiline mode- ~~ the smartmatch operator binds the regex to the matchee and much more
- character classes are now <[...]> (plain [...] does what (...) does in math)
- perl's triadic x ? y : z becomes x ?? y !! z
We can have whitespace in our regexen now (and comments and multiline regexen)
my $regex = rx/ \d ** 4 #`(match the year YYYY)
'-'
\d ** 2 # ...the month MM
'-'
\d ** 2 /; # ...and the day DD
say '2015-12-25'.match($regex); # OUTPUT: «「2015-12-25」»The same thing is available in many other languages. They copied it when they copied from Perl. For example Python's https://docs.python.org/3/library/re.html#flags documents that re.X, also called re.VERBOSE, does the same exact thing.
The fact that people don't use it is because few people care to learn regular expressions well enough to even know that it is an option. One of my favorite examples of astounding people with this was when I was writing a complex stored procedure in PostgreSQL. I read https://www.postgresql.org/docs/current/functions-matching.h.... I looked for flags. And yup, there is an x flag. It turns on "extended syntax". Which does the same exact thing. I needed a complex regular expression that I knew my coworkers couldn't have written themselves. So I commented the heck out of it. They couldn't believe that that was even a thing that you could do!
PS. raku has added quite a lot to the regex facilities we are familiar with, not least a straight line to using them in Grammars with rule and token methods that give you control over handling of whitespace in the target
The bigger problem here is executing user input.
There's a whole lot of faulty expressions out there for validating email addresses. I prefer to do less validation and let it fail. If the email address is wrong, whatever service you're using for sending emails will just reject it. If you really do need to validate email addresses, use something somebody else wrote that does it properly.
If you're working with some exotic format for which there isn't already an open source library, do what this guy says: parse it, don't try to validate it with regex: https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-va...
However, the conclusion of "use something somebody else wrote that does it properly", while valid, is asking a lot. As regex is hard to get right, don't assume the code you find on the web or book or via some other means works correctly.
My rule is if I didn't write it and can't wrap my head around the code to convince myself it is the right solution, I don't use it. And as I think others have written, there are some interactive online tests for regex expressions that can help.
Email addresses in particular are surprisingly complicated and far from being regular languages. I don't know how commonly real servers support the full feature set, but even if they just support non-ascii names they quickly become a pain.
The fact that Ruby has this behavior at all is a major security issue.
0. There is no universal regex language but many.
1. Perl-like ones (Ruby, Perl, and PCRE1/2) contain additional hidden traps.
2. You must vigorously match untrusted input to assume it to include invalid unicode, control characters, and other oddities.
3. You should replicate frontend and backend validations to ensure they are always exactly consistent and correct, preferably through fuzzing and/or property testing.
@neon = "Glow With The Flow"
erb :'index'
What exactly is `@neon = ERB.new(params[:neon]).result(binding)` even supposed to be doing?Why wouldn't it just be:
@neon = params[:neon]
erb :'index'When does the blogspam end?
"Consider every ambiguity of technology as a personal marketing opportunity."
The true topic at hand is that text substitution in scripted services is an eternal hazard of code injection.
The point that regex-based input sanitization doesn't work because everyone misunderstands the token semantics for string termination is made to look like a marvelous mitigation, but this teaching on regex is distracting from an unavoidable hazard of scripting.
Good news for the contractor: he appears like Jesus to shine the Lord's light on the sin of the fathers while dancing by the moral hazard of the priesthood.
Elsewhere another instance of the OP is a service provider pushing business solutions based on the ease of use of scripted service frameworks ("Input sanitization is as simple as a regex!)
These hazards are going to get much worse as AI merges the causes of and solutions to these ambiguities into the same semantic mush.
If you read the early papers, you get a very clear language for pattern matching on sequences. They have really nice properties - the compilation to finite automata gives you decidable equality and decidable minimisation. As in you can compile equivalent regex to exactly the same state machine however they were expressed.
At some point perl happened and that seems to have sent us down a path to encoding the regular expression in an illegible subset of ascii. The backtracking implementation cost us negation and intersection. What should be linear time matching becomes exponential.
Emacs will let you write regex in s-expressions at which point they're much easier to read. Everywhere else has gone with "looks like Perl but has different semantics, which we kind of document, be lucky".
I started writing tests to check that regex I'd begrudgingly converted to the perl style behaved the same under different engines and the divergence is rough. Granted I was parsing regex with regex which is possibly a path to insanity but things like a literal [ were a real puzzle to match on different implementations.
I don't know that the horrible syntax on semantic beauty is due to perl but it looks likely from a superficial standpoint.
Easy to explode into a lot lines, but I'd rather have a 50 line RegexBuilder implementation than try to keep track of what the equivalent single-line version is doing. Especially if you ever have to come back to it later and understand it again.
And if you ever make revisions in RegexBuilder you have useful diffs instead of "the one line that does everything is different than before."
https://developer.apple.com/documentation/regexbuilder
Are there similar tools in any other languages?
Parsing regex then pretty-printing the parse tree as s-expressions is very legible. You can also print the parse tree as the original syntax. Postfix will work better for some people, I like the lispy look for parse trees.
Most regex are similar syntax over a parse tree with different parts missing, if you keep track of roughly what features the current engine has in your head the sema checking a real compiler should do could be deferred or incomplete.
Some coding standards will want redundant escapes because that is considered more readable, could put that logic in the pretty-printer.
That's sort of suggesting using your IDE to translate the thing back and forth on the fly instead of persuading colleagues to stop writing in the obfuscated format.