Tiny C Compiler
bellard.org
bellard.org
I used it as a scripting backend for a long time, but eventually you will realize that you are not the only one that could write scripts, but you also really need to have a trust boundary (or just a different address space), so that errors in the script don't drag the whole thing down.
That's where Lua and LuaJIT shines: It's simple, and it's sandboxed. However, there is still one thing missing. These days, several programming languages have several targets that they can compile to, and so imagine if you could emulate some platform at high performance with low overhead, in a sandbox. You would then be able to script in whatever language you so desire, provided it has that target.
Unfortunately, those languages tend to be system languages and not the absolute best for scripting. With one big exception: Compile-time programming support.
I don't really recommend doing it, unless you want to do it as a learning experiment. It was a fun thing to do.
It is more complicated than just embedding a scripting language (or a scripting compiler :)) but more general and secure.
Especially if WASM evolves to a point where you can compile C# and other high level languages to it (you can to a point already but it's not on par with native runtimes) - it would be the most general embedding runtime.
Mike Pall told me this in no uncertain terms: http://lua-users.org/lists/lua-l/2011-02/msg01582.html
> The only reasonably safe way to run untrusted/malicious Lua scripts is to sandbox it at the process level. Everything else has far too many loopholes.
I have been benchmarking against LuaJIT for the longest time because I thought it was an equal in that respect. Guess it should have been against regular Lua.
Lua is commonly used for scripting games. If I load an untrusted script into the game I’m playing and it can access all of my files, it’s potentially catastrophic. OTOH if the worst an untrusted script can do is use 100% CPU (and cause the game I’m playing to lock up so I need to restart), it’s a minor annoyance.
AFAIK in the absence of bugs (and regular Lua is pretty close to that [0]), Lua’s sandboxing features are sufficient to protect against the former but not the latter. Am I right to think that’s often good enough? It seems a huge improvement over embedding, say, Python - I wonder how it compares to a embedding a JavaScript engine like V8?
The answer in the case of LuaJIT is definitely no, because the JIT engine is sufficiently complex that exploits are inevitable. Note that this is also the case with JavaScript. Many browser exploits start with some codegen bug in V8 or SpiderMonkey. And there are many more eyeballs looking to fix bugs in V8 and SpiderMonkey than LuaJIT, so in the case of LuaJIT the prudent answer is that you should never trust it to run potentially malicious code.
The case for PUC Lua is more nuanced. Lua used to have a bytecode verifier, but it was removed for 5.2 because too many bugs were found in the verifier, and because the VM relied heavily on the verifier to filter bad opcode patterns, that led to sandbox breakouts. This made the PUC developers believe that the verifier made Lua worse off as compared to a more wholistic emphasis on correctness and robustness, so they dropped the verifier. They also dropped any pretense that you could safely sandbox bytecode (i.e. precompiled scripts); if you want a sandbox in Lua you should only load untrusted code as plain Lua scripts into the sandboxed environment. To that end Lua 5.2 added a parameter to all APIs for loading code that specified whether to accept text scripts or binary bytecode. In other words, the bytecode verifier was considered a third wheel, so they removed it and redirected attention to the compiler and the rest of the VM.
So for PUC Lua the issue really comes down to how prudent it is to draw a trust boundary around a pure Lua sandbox; or rather, what your adversarial model is, precisely. PUC Lua is committed to Lua's sandboxing features, but many developers are fairly of the opinion that the only way to run untrusted code, if you're to run untrusted code at all, is using either a hardware VM or a very strict seccomp jail. If you're of the latter opinion, the language is irrelevant--you shouldn't trust Lua, JavaScript, Java, or any other language environment, period. In practice, however, even people of the latter opinion generally apply the principle of defense in depth. That's why browser JavaScript APIs and capabilities are still relatively limited, even though browsers execute JavaScript in OS-based sandboxes. In most practical contexts Lua's sandboxing features still provide great value; it just needs to be understood that they're a complement to rather than substitute for process sandboxing.
The same analysis applies to WebAssembly, FWIW, especially JIT'ing WASM environments. Anybody who thinks WASM is a magic cureall for running untrusted code is mistaken.
Presumably a formal approach could also work. CompCert exists, after all.
- 5.2, 2016: https://apocrypha.numin.it/talks/lua_bytecode_exploitation.p... (9MB PDF)
- 5.2, 2016: https://github.com/erezto/lua-sandbox-escape
- 5.1, 2015: https://www.corsix.org/content/malicious-luajit-bytecode (warning: dense)
There's a "luarop" link (boop.i0i0.me/blog.lua/luarop) referenced in the PDF, but the link sadly seems to have died (IA never crawled the domain).
Do any of them work without needing to load arbitrary bytecode, which is known to be insecure?
And presumably not.
My understanding is that this is what Google's PNaCl was about (https://developer.chrome.com/native-client/nacl-and-pnacl) : leveraging LLVM bitcode for both portability (source and target) and native performance, being able to sandbox properly the resulting executable.
Understatement of the century! :-) The breadth of his work, all non-trivial, is truly humbling. All done without fanfare or self-aggrandizement. I can't even find a good interview of him to understand his mind and thought-process.
And the interview should be purely on the technical side i.e his educational background, how he approaches design and programming, how he learns new technical domains, his thoughts on languages, advice to young programmers etc. Given his off the charts productivity, i suspect he has a highly efficient way of learning and doing things which i would like to copy :-)
//usr/bin/tcc -run $0; exit
main() { printf("Hello\n"); return 0; }
Just `chmod +x` and run it...EDIT: hah, I should have realized that "//" is the C comment and "//usr/bin/tcc" is equivalent to "/usr/bin/tcc". Clever!
So it's running the file as a shell script, where the first line runs tcc on the current file and quits.
@rem = '--*--Perl-*--
@echo off
echo Hello from a command script
perl -x -S %0 %*
goto endofperl
@rem ';
#!perl
print "Hello from Perl!\n";
__END__
:endofperl
By far very less common is to stuff some Perl code in a C# source file. A dev that was way too clever for his own code did this to have a Perl script that would update a C# script file and be self contained. This is a simple example of what it looked like, minus the instant headache anyone got that had to look at the real version: #if PerlScript
$csharp=<<WhyJustWhy;
#endif
using System;
namespace Example
{
public static class Program
{
static void Main()
{
Console.WriteLine("Hello from C#");
}
}
}
#if PerlScript
WhyJustWhy
print("Hello from Perl!")
__END__
#endif #if 0
set -e; [ "$0" -nt "$0.bin" ] &&
cc "$0" -o "$0.bin"
exec "$0.bin" "$@"
#endif
#include <stdio.h>
int
main(int argc, char *argv[]) {
puts("Hello world!");
return 0;
}
It works, because by default system shell will be spawned. #!/usr/bin/tcc -run
#include <stdio.h>
int main()
{
printf("Hello World\n");
return 0;
} #!/usr/bin/rdmd
void main() { import std.stdio; writeln("hello"); }Now I carry on and would like to make a TCC rival. I also wondered if I can make a TCCPP until Iearnt not only C++ syntax is ambiguous (and so we need some form of context sensitivity, and GLR is one good technique to handle C++ type of thing) C++ templates are Turing-complete effectively meaning it will be hellish hard. I abandoned this idea thereafter.
I believe a C parser would also need some form of context sensitivity. "T * x;" can be a declaration or a statement. However, C++'s syntax is larger than C's and I believe that there may be more cases where context sensitivity is needed to resolve ambiguities.
https://eli.thegreenplace.net/2007/11/24/the-context-sensiti...
Clang basicaly moves the semantic processing to another pass, and tokenizes both types and variables as ‘identifiers’.
I’d agree that C/C++ have some sort context sensitivity. However I think this is not an ambiguity. You can deduce a single logical solution when you see T *x; . If the T in scope is a variable you do multiplication, if the T in scope is a type definition you make a variable declaration.
Real ambiguity happens when there are two logical solutions to the same statement and the language specification has to prefer one to remove ambiguity and undefined behaviour.
For example a C# code example from spec that is ambigious and therefore doesn’t compile:
“ static void F(bool a, bool b) { Console.Writeline($”{a} and {b}”); }
static void Main(string[] args) { int G = 2; int A = 1; int B = 2; F(True, False); F(G<A, B>(7)); ”
The compiler prefers to interpret the last line as a function call with one argument, which is a call to a generic method G (that doesn’t exist in the code above) with two type arguments a one regular argument. Instead of one function call with two boolean arguments.
template<size_t N = sizeof(void*)> struct a;
template<> struct a<4> {
enum { b };
};
template<> struct a<8> {
template<int> struct b {};
};
enum { c, d };
int main() {
a<>::b<c>d;
}
This is not ambiguous, but the meaning of the line inside main() depends on the compiler and the target - it can be either a declaration: a<>::b<c> d;
or an expression: a<>::b < c > d;
depending on which b gets selected. This in turn affects future references to d etc.But yeah, it looks like a C++ compiler has to instantiate templates in lock-step with parsing to make this work, but possibly there are some tricks that are applicable here. In any case, this looks like a hard problem.
Nevertheless I think the C# way of choosing generics over comparison without semantic processing is strange. The example above should run as intended (as a simple function call with two arguments) and not fail because of some spec rule.
What I mean is something like "Foo<Bar<Baz>>", the ">>" part can be a little bit challenging because we don't know if this is a right shift or not without context
C++ is also Turing Complete that its template system can implement boundless recursion and infinitely expand on its own without termination...essentially it means it may not even halt. That's why template metaprogramming is a (hard and horrible) thing...it's a sublanguage that uses C++ syntax and generates C++ code making the parsing even more non-deterministic.
2016 https://news.ycombinator.com/item?id=13249851
2017 https://news.ycombinator.com/item?id=15272894
2018 - obfuscated! https://news.ycombinator.com/item?id=17335856
> On 31 December 2009 he claimed the world record for calculations of pi, having calculated it to nearly 2.7 trillion places in 90 days. Slashdot wrote: "While the improvement may seem small, it is an outstanding achievement because only a single desktop PC, costing less than US$3,000, was used—instead of a multi-million dollar supercomputer as in the previous records."[10][11] On 2 August 2010 this record was eclipsed by Shigeru Kondo who computed 5 trillion digits, although this was done using a server-class machine running dual Intel Xeon processors, equipped with 96 GB of RAM.
https://andrewkelley.me/post/why-donating-to-musl-libc-proje...
:(
Both compilation and link are generally finished before your fingers have relased the enter key...
(In all seriousness, I am curious as well.)
DSL's goal was to provide most of the tools you might need in a tiny footprint. The platform target was the old "thin-client" style machines which might pack a 500 MHz AMD Geode or a VIA C3 processo and have either no internal storage or maybe a few hundred megabytes for preboot environments.
They were able to fit X (Xfree86, iirc) w/ Fluxbox, Firefox, Ted (a word processor), and a few other things into a 50 MB "bizcard" image (PCMCIA cards & compact flash).
Another thing it was designed to do was allow you to install most of the common things to a small internal storage device, while running the rest off of optical media.
It dates to the time before common support for USB booting. Also kinda crazy to realize you could fit Firefox, a windowing environment, and a kernel into 50 MiB while today the standard vim install is 33 MiB...
of course I could compile it, but really, it doesn't have to, and to develop/try that sort of small C tools, tcc is just unbeatable.
[0]: https://gist.github.com/buserror/227a7e64c92acece821ec8ee587...
https://web.archive.org/web/20110726063943/http://www.freear...
How old is this project?
http://www.doc.ic.ac.uk/~phjk/BoundsChecking.html
Someone needs to add this to llvm.
Clang supports bounds checking as part of -fsanitize=address, (though with a few more flags you can _just_ have bounds checking instead of the other sanitisation options). (Since around 2015?)
GCC supports bounds checking and others, depending on which frontend you're using the options can change. (Like -fbounds-checking for C, and -fcheck=all for gfortran). (Since around 2013? GCC 4.71)
Even Intel has -check=bounds. (Though not under macOS). (Since around 2015?)
Is there a distribution that offers bounds checking for all linux software ? I wonder how slow a typical LAMP stack will be. My guess is no more than 5x. I think thats an acceptable tradeoff. I'm guessing there's a way to add global compiler flags in source distributions like Arch linux / BSD.
(Arch isn't actually a source distro, it's binary, though rolling release. Gentoo is, however).