Forth: The Hacker’s Language
hackaday.com
hackaday.com
Part 1: http://git.annexia.org/?p=jonesforth.git;a=blob;f=jonesforth...
Part 2: http://git.annexia.org/?p=jonesforth.git;a=blob;f=jonesforth...
(github mirror: https://github.com/AlexandreAbreu/jonesforth/blob/master/jon... https://github.com/AlexandreAbreu/jonesforth/blob/master/jon...)
Previous HN comments: https://news.ycombinator.com/item?id=10187248
Thanks!
It's an excellent codebase; thank you very much! I've changed the design a slight bit so far -- but the vague signature of jonesforth is really still there, I think!
Once I got a better idea of Forth, I also realized that Jones stays in assembly for rather long. He builds up the entire interpreter from raw assembly words, because you need the interpreter to start parsing textual Forth source. But you if you could somehow write Forth before the interpreter is put together, it would be quite natural to switch to Forth much earlier. And it turns out you can do that. Jones even defines the `defword` macro. But for some reason he makes very little use of it. Rewriting most of the code leading up to INTERPRET using defword was a fun exercise.
The next step would've been making the whole system bootstrapping by rewriting the assembler in Forth, but x86 machine code generation is hairy enough that I bailed out at this point.
Of all the alternatives to C, Forth is the only one which groks hardware. This is particularly important for embedded development. Any realistic alternative to C must understand interrupts, memory mapped registers etc.. Forth is the only one in which this is possible. Every other example I've seen always uses C for "low-level" access or a libraries. If you need to do that, then you won't win over embedded devs like me. It's much easier to write good embedded C code (no mallocs etc) than having to debug FFI.
Having said all that, my mind doesn't understand Forth. I'm too used to seeing code in C or even assembly to fully understand Forth. Nevertheless things like https://github.com/jeelabs/mecrisp-stellaris are really cool and are a realistic alternative to C.
Forth is a great little language, and a perfect example of how to to do a surprising amount with very little. For more complicated embedded work, however, I've been deeply impressed with Philipp Oppermann's tutorial on writing a kernel in Rust: http://os.phil-opp.com/ You have to limit yourself to Rust's "core" library (instead of the usual "std"), and you won't have a heap until you write one, but it's surprisingly nice. Well, except for debugging double faults. That's never nice. :-/
Here's my toy PIC 8592 interface in pure Rust: https://github.com/emk/toyos-rs/tree/master/crates/pic8259_s... Be careful, it's a subtle chip and I only deal with the basics.
But if you're working in a raw microconroller, you're going to be touching metal constantly, and you won't want anything you don't use slowing you down.
In short, Rust is great, but it doesn't beat Forth in Forth's biggest problem space: programming embedded systems under exceptionally tight size/speed constraints.
Comparing it to Rust is missing this point entirely. Finally, let's not forget that Forth has been empirically validated multiple times in this domain (e.g. NASA has used Forth on board satellites and spacecraft).
Rust is entirely unproven in the hardware/microcontroller space (some would say in general), so I tend to view posts like ekidd's ("for more complicated embedded work") as projecting/wishful thinking.
But this is a huge draw for Forth, speaking as a Lisp user.
Not really.
https://www.mikroe.com/mikrobasic/
https://www.mikroe.com/mikropascal/
http://www.astrobe.com/default.htm
http://www.adacore.com/gnatpro/embedded
http://www.ptc.com/developer-tools/apexada
http://www.ghs.com/products/ada_optimizing_compilers.html
Zero lines of C.
Some of these do look neat: I've had an interest in both Ada and Oberon, although I haven't been able to get over the verbosity and unpleasant syntax of either (it's not COBOL, or anything, but it's not nice).
OTOH, I am immediately suspicious of any product that claims it's professional and also has BASIC in the name...
For many years the Elektor magazine and a few other competing ones, had the listings of electronic stuff done in Assembly, Basic and Turbo Pascal.
Before Raspberry PI was a thing, many developers were using Basic STAMP.
https://www.parallax.com/catalog/microcontrollers/basic-stam...
Although I suppose it does depend on the dialect, to an extent...
www.esmertec.com
Of course there are other systems programming "monolithic executable" programming languages that allow straightforward access to memory mapped hardware registers.
In what way does C understand interrupts?
And yeah, it's a given that it's going to be non-portable.
Given asm, implementing a Turbo C-style int86() function seems pretty trivial. It might even be doable as a macro.
asm {
mov ax, 0x0013
int 0x10
}
Can't really like the way clang and gcc asm work, even it it means giving more info to the optimizer. int foo(int c)
{
int r;
asm {
mov ax, 0x0b00
int 0x21 /* poll stdin for input */
mov ax, offset c /* get address of `c` */
call bar /* call other C function */
mov r, dx /* save the result in `r` */
}
return r;
}
Or something like that — it's been a while. Those were the days! GCC's (and I'm assuming clang's is much the same) inline assembly is a joke in comparison. I remember, when I first encountered it, flipping back-and-forth through the GCC docs looking for how people actually do inline asm with GCC...Also, my time with Lisp has taught me that interactive environments are A Good Thing, and Forth is exceedingly interactive, and has better native introspection than many Lisps (when your guts are strewn across the floor, it's hard to hide them).
But for all that is good about it, Forth does have its problems: depending on how well your problem maps to the stack, Forth can be exceedingly painful.
A discussion of Forth's flaws, of course, would not be complete without a link to Yossi Kreinin's now-famous post on the subject: http://yosefk.com/blog/my-history-with-forth-stack-machines....
Unfortunately, learning Forth was a real let down. Forth turned out to be way too much of a write-only language for me, with lots of convoluted, hard to understand code. The Forth community preaches simplicity and using short, clear, well documented words (functions). But the reality, at least in most of the open source Forth code I've seen was the opposite, with long, convoluted, poorly documented words being the rule rather than the exception. I've been told that commercial Forth code is a lot better in these respects, but I haven't verified that.
I've also heard even Forth fans tell me that the whole point of Forth is to write yourself something more pleasant than Forth to work in as quickly as possible. Since I don't have an interest or need to write programming languages myself, I'd just rather start with something more pleasant from the beginning, and skip the painful stage of having to work with Forth.
In very resource constrained environments where your only alternative to assembly is Forth, Forth might be the best choice. But since I don't do most of my programming in such environments, I really don't see the point of using Forth at all.
Finally, after learning Forth, I went on to learn Lisp and then Scheme, and fell in love. Those (especially Scheme) really were super clear and easy to both write and read, and felt far more powerful and not at all painful. I felt far more productive in them, and would now far rather use a tiny Scheme or even Lisp on a resource constrained system, if at all possible.
https://github.com/s-macke/starflight-reverse
It turns out to be very portable because the assembler code is only about 1000-2000 lines. The rest is implemented in Forth itself. They use indirect threading to produce very dense code. They kept even most of the debugging symbols inside the code and kept the forth interpreter itself in the code.
http://web.archive.org/web/20030214111810/http://www.sonic.n...
https://github.com/s-macke/starflight-reverse/blob/master/sr...
is based on these documents and I filled some gaps of the symbol table. To start the game they used the word "LET-THERE-BE-STARFLIGHT", in the binary you see only part of it: "LET-TH".
I wish this "someone" would provide a full copy of the archive :)
[1] http://forthworks.com/retro/People on reddit seems to hint at the fact that Forth people stay under the radar because it allow them to design and solve problems in better ways.
I wonder how true it is ..
> But Forth is also like a high-wire act; if C gives you enough rope to hang yourself, Forth is a flamethrower crawling with cobras. There is no type checking, no scope, and no separation of data and code. You can do horrible things like redefine 2 as a function that will return seven, and forever after your math won’t work. (But why would you?) You can easily jump off into bad sections of memory and crash the system. You will develop a good mental model of what data is on the stack at any given time, or you will suffer. If you want a compiler to worry about code safety for you, go see Rust, Ada, or Java. You will not find it here. Forth is about simplicity and flexibility.
When it comes to working on a large project in a team, I suspect Forth doesn't scale.
Think of why Go works so well for Google: it's designed for code bases that are maintained by many people. And almost every design decision that was made for that goes against the spirit of Forth.
It's too flexible; you can change anything about the system so any Forth code ends up deeply personal, and specific to the current task. It's a language for individual artisans, which is another reason why it works well in the embedded space. But it's not hard to see how it all goes up in flames when you're in a large team.
But I'm still curious about which company is using Forth, who are the devs, etc etc
That's the point though. You create a domain specific language using Forth and solve your domain specific problems with it.
DSLs come at a cost.
If, in any language, flexibility is a problem for a larger team, fire the lead developer. The article does a good joint pointing out that using that flexibility is usually a sign of code smell.
There's another dimension of this that works for Google, though, and it's one that people particularly here should take note of.
As far as I can tell, Go is a language where the solution to a given expressive problem is usually "write more code." There's nothing subtle or clever to leverage into a multiplier, like there is with Forth or LISP. You don't write a DSL. You don't factor out largely common implementation details. You may well not create composable subcomponents for some intermediate level. Instead, you put in the time, turn the crank, and create whatever straightforward code solves your concrete problem at the level it's written, straightforward if verbose. It's similar to how MJD describes working in Java (http://blog.plover.com/prog/Java.html ), except you're missing some of Java's expressive power and you have some of Pascal's to make up for it.
There's a degree of simplicity to it.
An organization like Google is not lacking in any way resources to have developers produce more code. Or understand it, if any issues with expressivity to volume ratio ever makes a given codebase difficult to grok.
Does your organization have those resources?
Do you?
The answer to that question (along with who you're competing with) might tell you whether you should be using a Google language or a hacker language.
https://www.infoq.com/presentations/power-144-chip
Glad this piece popped up, time to finally watch it.
http://www.greenarraychips.com/
which is actually a truly message-passing based arch with 144 clockless "cpus". Each cpu is actually so limited that forth is one of the few languages that makes programming this tricky platform viable.
http://w3-o.cs.hm.edu/~ruckert/compiler/ThinkingInPostScript...
tutorial and cookbook (blue book): https://www-cdf.fnal.gov/offline/PostScript/BLUEBOOK.PDF
language program design (green book): https://www-cdf.fnal.gov/offline/PostScript/GREENBK.PDF
Back in the 90s the most popular train route web site in Germany used to send you a PS file for printing on Linux (they did something more native for Windows). And that was basically a serious of quite easy to understand line drawing and text positioning instructions.
And I think it was jwz who had some cassette cover sheets where you put in the song and artist data by editing the PostScript file itself.
It's long out of print, but used copies are cheap, and scans are readily findable on the net.
I wrote a comment a while back outlining how to do a FORTH-like language from scratch by starting with a simple calculator and expanding it, writing in C with optional assembly optimizations. This was aimed at people who have less knowledge of FORTH-like systems than you do, I think, but if you are curious here it is [3].
[1] https://www.amazon.com/Threaded-Interpretive-Languages-Desig...
[2] http://www.retroprogramming.com/2010/03/threaded-interpretiv...
https://docs.google.com/presentation/d/1SJQGow_fnMSt5PsMgESu...
Some discussion on linked slides: https://www.reddit.com/r/concatenative/comments/4bj4lu/slide...
Extensions that saved me the most agony:
1: Use of stack frame comments to actually define local variables and enforce the in - out stack transform. While at this I also added the comment string to the word definition in an additional help link. Thus:
[:] do_it ( in1 in2 -- out : return in1 * in2 * 2 ) local lv : lv_helper 2 * ; in1 @ in2 @ * lv_helper out ! [;]
> help do_it in1 in2 -- out : return in1 * in2 > 3 6 do_it . 36 >
Note the [:] and [;] redefs for compiler words, had to do this to accommodate nesting, locals, stack frames, while not compromising regular FORTH.
all variables are local in scope to the containing word. Defining variables inside ditto. Had words for listing contents of these too.
I later extended this to include nested word definitions, dynamically allocated variables, and stack frames that will unwind these on execution error. Gone were stack frame size errors, a missed dup or drop no longer fatal, indeed dup and drop not really necessary any more.
2: Structured data debugging aides: a few words and a C header parser allowed me to dump or modify readable contaents of application structs, and to call into application functions.
3: along the same lines of nested definitions, I was able to associate sub-dictionaries with graphical views along their heirarchy. This was a later addition, applied to VNOS. Lots of little things became manageable with this paradigm. In particular, allowed me to do a scoped local console (as a debug object instantated in a View), and to insert message overrides into the VNOS messaging mechanism, effectively adding what turned out to be a really powerful scripting system. Done as an object too - current version includes Perl 5 and Lua (almost ready).
And yes, concurrency is a concern, and being sure not to call non-re-entrant (errant:) functions, but hey, ypu'd likely be surprised at how good a pattern recognizer ther Human Brain is here.
Over-all, I think I used this system with a good half-dozen different CPU architectures and their OS's, also in the high level VNOS GUI (originally put my FORTH into it for debugging purposes, but that was rapidly out-grown).
Forth was quite popular in Europe thanks to Jupiter Ace and ZX Spectrum extensions.
By the way, I found a rather neat page about building a Jupiter ACE: http://searle.hostei.com/grant/JupiterAce/JupiterAce.html
Later ('77), TIS APL ran on Z80 systems, and was aparently quite a bit better, although it wasn't a full implementation.
But I'm just quoting Wikipedia at this point, so look it up yourself.
And then I got my hands on a Forth. Functions. Control structures. Maybe they weren't fancy by modern standards, but they were easily an order of magnitude better than Basic's, and you had the ability to write new ones fairly easily. Quick compiles, code which was both small and fast. Easy inline assembly code for when you needed even more speed. It was a dream come true.
Or at least, it was last I used it.
It wasn't until the 2nd or 3rd time I'd used it that I actually figured out how to make sense of it and run something (for Scratch's definition of "run").
To be honest I've progressed extremely slowly with CompSci/programming over the past 18 years I've been using them (got my first computer around 7-8) - I started with QBasic, been shouting at PHP for way too long, I have a basic understanding of C I badly need to develop, and I'm moving toward playing with Lua next - and I hardly consider myself a dyed-in-the-algorithms academic type with a brain that's unable to understand Scratch. (In fact, I'd argue that the best programming teachers would be precisely those types of people, and if they were unable to understand Scratch that would be a major problem.)
Rather, I firmly believe Scatch's UI is a disaster, and horribly unintuitive to use. Other languages are beset with grammatical idiosyncrasies; with Scratch you have to learn the UI before you can learn the... few parts of the language that are actually there.
I'm concerned that systems like Scratch are so widely used; I fear that it's an even worse mind-scrambler than the bad sides of BASIC. Of course, like BASIC, there are good sides, and it teaches the basics without presenting a Mt. Everest-sized learning curve. Perhaps https://en.wikipedia.org/wiki/Dartmouth_BASIC was the Scratch of 1964, and I'm just griping about the dilutory effects of educationally-targeted software in this day and age and "modern" GUI design.
Scratch is also really slow/laggy on my old laptop (Thinkpad T43), I can't imagine how bad it is for schools with limited hardware.
If you want a Real Language presented the same way, Snap! (descended from BYOB) is essentially a Scheme in Scratch's clothing.
But the two real draws of Scratch were its hackability and its community. Back before Scratch 2.0 ruined everything, Scratch was written in Smalltalk, and using a widely-known hidden feature, you could examine the source code and make whatever changes you wanted with relative ease, resulting in a healthy community of mods and derivatives which explored new features and ideas, or those that the official team had dropped by the wayside (like Mesh, a fully-featured networking system).
Scratch's community was likewise excellent: I spent a lot of time lurking in the Scratch Advanced Topics forum - a sort of off-topic general programming section, where people far smarter than I discussed modifying Scratch, improving the website, and whatever programming projects they happened to be working on (usually web programming in PHP - it was the mid 2000s, after all).
But that's enough nostalgia for one day...
I suspect the reason Scratch felt slow to me is that 2.0 is some kind of HTML5 and/or Flash mess now - you're right, Smalltalk is really fast. I spun up Squeak to check something on this T43 yesterday, and everything was really snappy. I have no reason to expect Scratch will be any slower.
Also, I wouldn't be surprised if a reasonable bit of the exploration everyone did was motivated by the fact that they were "hacking" the platform :P
There seems to be a sad lack of excellent online communities nowadays; I've long looked for sites to complement HN, but without success.
I just had a look at Snap! which is interesting. It definitely flatlines this laptop though, I had to try it on a faster machine. But I'm running the tree animation demo right now, and it looks awesome....
Nor would I. Even at the age of 8, before I was really able to understand the code, there was a thrill to it, in a cracking-open-the-toy kind of way.
And it helps that Smalltalk does exploration better than just about any other language/environment. You can just open up any Smalltalk app and extend/take it apart using the same tools the developers did to build it.
>There seems to be a sad lack of excellent online communities nowadays; I've long looked for sites to complement HN, but without success.
Lainchan (a sort-of cyberpunk/whatever chan) is quite popular with some of the people who are here on HN. It's a very different atmosphere, but it does emphasize actual good discussion. And it's got a containment board for politics, which always helps.
At the very least, their magazine (https://lainzine.neocities.org) is worth looking at, if not for the generally interesting articles, than for the outright strangeness of a lot of it.
I was recommended Lainchain a couple months ago, actually, but nobody mentioned the magazine, which is really cool. I am not impressed that the ASCII art generation paper in Vol.1 has any associated source code!! The rest of the magazine content and design is really interesting too.
And it is dead simple to implement. I did it in common lisp 3 years ago for an IRC-bot - a lazy Sunday afternoon and 50 lines of code later I was done.
I would not want to write code in Forth on a regular basis, but reading (and making an effort to really understand it) definitely made me a better programmer.
The August 1980 issue of BYTE magazine was all about FORTH. I read it cover to cover, then sent to the Forth Interest Group for a copy of the implementation guide and the 8080A listing of Forth.
I typed it all in, including comments but changed the 8080A mnemonics to Z80 mnemonics. After the typos were corrected and the I/O routines modified to run under TRSDOS it ran as advertised. Then I optimized the code for Z80 which made it both smaller and faster.
My job then took me from Tandy's manufacturing plant to the R&D department where I wrote assembly code for another product but continued to work on Forth as a side project. Management said that they would be willing to release it as a product if I could get a signed statement declaring Fig-Forth to be in the public domain. I tried repeatedly to get people at the Forth Interest Group to sign but was never successful. This is the reason Radio Shack never released Forth, though many people in R&D used my code internally. They especially liked the ease of number base conversion.
From Tandy, I went to an industrial equipment manufacturer where I wrote a lot of assembly for embedded applications. I ported CP/M to run on TRS-80 model 12. The need arose to read input from a graphics tablet via serial, process the input and output a bitmap. My estimate on the length of time it would take to write this for the CP/M machine extended past the deadline.
I combined my Forth code with the drivers for the Model 12. The boot procedure for Model 12's consists of reading 26 128-Byte sectors from Track 0, Side 0 into memory and jumping to the loaded code. I changed to format of the remaining tracks (both sides) to 9 1K sectors and a 256 byte sector (double density). My boot code prompted to press the F1 or F2 key, then read all the 256 byte sectors on side 0 or side 1 (depending on F1/F2) into memory. I added a FSAVE word which wrote memory to the 256 byte sectors. I had a Forth operating system. I wrote the code for the application in Forth, performed the necessary scanning and conversion and output the desired bitmap with time to spare.
I highly recommend learning assembly for at least one processor. Without it, you can only know about Forth, you don't really know Forth. I understand how directly threaded code (DTC) works but I prefer the indirectly threaded code (ITC) model of Fig-Forth as it seems more elegant to me.
I still have an x86 machine running Windows XP which will run the 16-bit code but it will not run on my Windows 7 system. I am looking at ciforth [1] which is a 32-bit implementation. I intend to get it running under Windows 7.
Forth is a really good system for a command line environment. As far as readability goes, the code itself has such high information content that it is difficult. It is up to the programmer to add abundant comments which make reading the code unnecessary in most cases.
I'll conclude with a Forth joke:
: Decompose ROT ROT ROT ;
Starting FORTH — Online Edition:
https://www.forth.com/starting-forth/
I had read it and played around with FORTH on a microcomputer some years ago. Fun language.
Thinking FORTH, a more advanced book, also by Leo Brodie, is also linked to from the above URL.