I wrote a 231-byte Brainfuck compiler by abusing everything
briancallahan.net
briancallahan.net
$ uname -a
Linux personal-1 4.19.0-6-cloud-amd64 #1 SMP Debian 4.19.67-2 (2019-08-28) x86_64 GNU/Linux
$ as -g -o bf.o bf.S && ld -o bf bf.o
$ size bf.o
text data bss dec hex filename
167 0 0 167 a7 bf.o
$ size bf
text data bss dec hex filename
199 0 0 199 c7 bfIt's still an impressive feat, and very informative on the assembly details - but doesn't feel as incredible as the headline makes believe as the core logic seems to be a string search-and-replace of the strings in the "reviewing brainfuck" table.
Seems to me, you could write the next brainfuck compiler in sed.
I'm more impressed by a brainfuck compiler written in brainfuck.
https://news.ycombinator.com/item?id=697501
https://github.com/matslina/awib
Awib is a brainfuck compiler entirely written in brainfuck.
Awib implements several optimization strategies and its compiled output outperforms that of many other brainfuck compilers
Awib is itself a 4-language polyglot and can be run/compiled as brainfuck, Tcl, C and bash
Awib has 6 separate backends and is capable of compiling brainfuck source code to Linux executables (i386) and five programming languages: C, Tcl, Go, Ruby and Java
The original was 240 bytes, and created an executable directly.
https://wiki.tcl-lang.org/page/Brainfuck
Brainfuck-to-Tcl transpiler:
Also, a parser implies having an internal, structured representation. As the C language is neither internal nor structured, I'd say the program is more likely a compiler.
I can write a tiny C compiler in bash...
The definition of this problem strikes me far more as either a transpiler or a pre-processor for a Domain Specific Language.
That said though, I do agree that calling an assembler like this (brainfuck is a kind of assembly) a compiler stretches things a bit, for an entirely different reason: usually there is some sort of a complexity threshold of the input language. The compiler has to at least maintain a symbol table for named entities (traditional assemblies has variable-like entities, macros and subroutines, the assembler has to keep track of all that). Brainfuck is completely linear, with no named entities at all. The "compiler" looks like a dumb string processor, just iterates over one buffer to transform it into another buffer, and doesn't maintain any sort of structures on the code being translated.
By the strict "compiler : program->program" definition, this doesn't matter. But my intuition holds that dumb string processing is a bit short of "true" compilation.
That is arguing semantics and, while not wrong, I think it muddies the waters: By that logic, sed and awk are compilers.
In the end, only machine code can be executed, so you'll need something at the end of your chain that produces machine code or at least executes your DSL in an interpreter loop.
So I think a distinction between "compilers" that generate machine code and "compilers" that don't is worthwhile.
https://en.wikipedia.org/wiki/Source-to-source_compiler#Inte...
This doesn’t seem wrong to me; after all, assembly is already an abstraction layer over machine code, and outputting assembly would hardly be “uncompilerish” behavior. I suppose it depends on whether you view C as a low-level language.
C-- is not a subset of C.
Bill Joy’s Law: 2^(Year-1984) Million Instructions per Second
https://donhopkins.medium.com/bill-joys-law-2-year-1984-mill...
> The name of the language is an in-joke, indicating that C-- is a reduced form of C, in the same way that C++ is basically an expanded form of C. ("--" and "++" mean "decrement" and "increment".)
It can optionally also generate assembly code if you want to look at it, but it does not do this during normal compilation.
That used to be the case in the dark ages of autoconf.
clang does everything in a single address space which is a lot faster now that you aren't doing a million filesystem calls.
btw I remember seeing a talk a few years ago about integrating clang with the build system so that the same compiler process can be reused to compile multiple files. Startup time is significant when you have a lot of source files
Past: >1 process and many fs calls per cpp. Present: 1 process and a few fs calls per cpp. Future: <1 process and a few fs calls per cpp (on average)
EDIT: clang can tell what's being done:
$ clang -ccc-print-phases -x c t.c 0: input, "t.c", c 1: preprocessor, {0}, cpp-output 2: compiler, {1}, ir 3: backend, {2}, assembler 4: assembler, {3}, object 5: linker, {4}, image
Machine code is generated "directly from the source code" in the sense that there are no intermediate languages produced. There are multiple stages of compilation, of course, but these all involved in-memory data structures, not textual languages.
The only external binary clang invokes is, optionally, the linker. Otherwise, everything is done in one executable invocation without temporary files.
With a bit of effort and the right add-on (https://github.com/JuliaComputing/llvm-cbe) you can make any LLVM-based compiler produce C code!
98 bytes, the smallest!
This is my favorite answer to why this should or should not qualify as a compiler.
If translating to c is a compiler then anything is a compiler and the term has no meaning at all.
I've created a new language called Brainfuck2 that is exactly the same as Brainfuck except the + symbol is replaced with a t. Clearly my compiler from Brainfuck2 => Brainfuck does not implement the identity function, and is less than 50 bytes.
[1] if you did it right, it's idempotent, which is about as close to trivial as a nontrivial function gets under the lens of symbolic dynamics, but I'm getting a bit far afield