Not being familiar with this sense of “thread,” I found a good explainer here: https://www.complang.tuwien.ac.at/forth/threaded-code.html
And, as always, Wikipedia: https://en.m.wikipedia.org/wiki/Threaded_code
However Forth does allow you to "inline" words, basically form a new word by chaining together existing words, which avoids this overhead. This doesn't optimise adjacent words, but it still gets semi-reasonable performance. (Inlining in JONESFORTH: https://github.com/nornagon/jonesforth/blob/d97a25bb0b06fb58...)
Modern Forths just have regular optimising compilers so none of this stuff applies, but they don't have the simple purity of one that you write and fully understand yourself.
In my benchmarks that are definitely not suitable to infer much, OpenJDK without JIT can perform 1.5-8x better, though of course the whole implementation is different so don’t read too much into that.
So many techniques that would improve performance on a 1980s processor can be woefully inefficient on a 21st century one. It's so easy to be still holding the computing model of the former in your head.
Super wrong. You're far better off precomputing the whole sprite bitmap -- even if you don't end up using it and doing a bulk operation to display it or not display it. Because doing that "if sprite is here?" in a loop is super super expensive, more expensive than just blitting the splits and not using them.
Most efficient ended up being precomputing things into boolean bitset vectors and using SIMD operations to do things based on those.
Even if not doing GPU stuff, the fastest way to compute these days is to think of everything in bulk bulk bulk. The hardware we have now is super efficient at vector and matrix operations, so try to take advantage of it.
(After doing this I have a hunch that a lot of the classic machine emulators <VICE, etc.> that are out there could be made way faster if they were rewritten in this way. )
For example to compile : SQUARE DUP + ; all the compiler has to do is copy the machine code of DUP to the place where SQUARE is being compiled, remove the ret instruction at the end of it, and copy the machine code of + after it. It can also do some small optimizations to remove redundant instructions.
You can do this with other languages, but concatenative languages can make it as simple as literally concatenating bits of code.
ERROR: DUP + is not SQUARE. Did you mean DOUBLE?Relatively easily, you can get rid of the interpreter overhead by writing blocks of machine code that do each Forth word's action, instead of bytecode to dispatch the interpreter to each of those routines. Forth can be adapted for that easily enough, and some Forths do support this on a per-word basis, allowing you to pick how you want your words compiled.
Forth can also get the whole optimizing compiler treatment. Some optimizing Forth compilers have been released, but I don't know how good they were/are. Certainly never needed that kind of speed myself. I don't know if any Forth yet benefited from it, but a lot of work was put into Java, on approaches for optimizing stack machine-style code to reasonably fast native code for register machines.