Arthur Whitney's One-page Interpreter (1992)
jsoftware.com
jsoftware.com
https://gist.github.com/anonymous/f72e5c4a432492abce59
Some basic observations:
- He uses all the obscure C features, like the ability to not declare the type of an argument or return value and let the compiler fill in, and also the old school style of declaring functions:
foo(x,y) int x, y { return 0; }
Instead of: int foo(int x, int y) { return 0; }
Since the compiler will infer the int, and since he uses R for return this example eventually becomes: foo(x,y){R(0);}
- Variables, or really registers, are only accessible as letters from a to z, and the "st" array stores the values of all registers. Numbers entered in the REPL have to be between 0 and 9, and hence he avoids the dirty work of making a proper lexer. It's also super easy to trigger a segfault since any error handling is non-existent.- DO(n,x) is a C macro that evaluates the given expression "x" for all numbers between 0 and "n"
- V1 is a C macro that defines unary operators for the interpreted language, and V2 defines binary operators. In V1 definitions the operand is called "w", in V2 the operands are called "a" and "w".
- For example ",", which calls the cat function, is a binary operator that creates vectors:
1,2,3,4
4
1 2 3 4
- The vt, vd and vm arrays map ascii symbols to the functions defined with V1 and V2. { is the second symbol in vt, so when used as a unary operator it calls "size" (second non-null element of vm): {5,6,7,8
4
and when used as a binary operator it calls "from" (second non-null element of vd).- wd is a parser that goes from the original input string to a weird intermediate form that is an array of longs. Each input character gets mapped one-to-one to an item in this intermediate form.
If the input character was a number between 0 and 9:
Value type instance gets allocated
Intermediate form for this input character consists of the address of the allocated instance
If the character is a letter between "a" and "z":
Intermediate form consists simply of this character
If the character represents an operator
Intermediate form consists of the index of the operator in the vt array
In other words, the intermediate form is an array where some elements are ascii characters, others are memory addresses and yet other indices into some array. This part is really something.- The ex function executes the intermediate form. Since everything in the input is fixed length, and there is no syntax checking, it just indexes into the intermediate form assuming everything is well formed, while the parser did not check that so it's not really guaranteed - again a source of easy segfaults. The execution goes from left to right and consists of looking at the first position in the intermediate form and then making recurrent calls if necessary (let X be the current item in the intermediate form):
If X is a character
Lookahead one item
If it is a '=' char
Assign the result of executing everything after the '=' to the register indicated by X
Assign to X the value of the register named by X
If X is not a character and is a small integer
We are applying a binary operator
X is the index into the "vm" array
Fetch the function from "vm", apply it to the result of executing the rest of the intermediate form
Otherwise:
If there is any more input remaining other than the current item, we are applying a binary operator
Lookup the function in "vt", apply it to the result of executing the intermediate form to the left and to the right of the operator
- I have the biggest problem with understanding that "a" struct, that represents all values in the interpreted language, which are arrays. ga is clearly the basic allocation function for it, "plus" obviously adds two arrays, so it's clear the "p" field holds the actual contents, but that's where things get very shady.Edit: the structure is at https://github.com/openj/core/blob/master/jtype.h .
So the "r" field in the "a" struct is the number of dimensions of the array that "a" holds, and the "d" field is the array that holds the length of the array "a" in each dimension.
Can we replace some of the one-letter codes with better namings?
This isn't an obscure feature, the default C89 type is "int". Leaving out "int" has long been considered bad form, and since C99 omitting return type has not been allowed (but generally supported in nonstrict modes).
(WRT weird names: IIRC, some time ago kdb featured a number of functions named with digits, like '2' being the function to connect to a socket or something.)
That being said, for more production-heavy applications and interfacing with other languages, variable/function naming could be more descriptive :)
Would you rather write/read:
DIVIDE X BY 5 GIVING Y
or
y=x/5;
The first is COBOL (designed to make code easier for "normal" people to read. The second is C (which looks more like the math that we learn in grade school).
k/APL/J simple moves further in the direction of math.
As a far as descriptive names go, in k, names can always be aliased with longer names (assuming no conflicts). I do this when exposing a library k code as a web interface.
I agree, I would very much like a ansified very that doesn't crash on unsupported features, but still in this style and in a single file.
3+7
a=3+7
a,a,a
b=a,a,a
{b
b+9
p=2,3,5,7
2{p
{p
~8
<~8
#~8
4#3
c=b#2
c,c
It's interesting that what he calls "Parsing" is more typically known as term reduction by pattern matching in the functional world.
noun(c){A z;if(c<'0'||c>'9')R 0;z=ga(0,0,0);z->p=c-'0';R z;}
would be
noun(c){
A z;
if(c<'0'||c>'9')
R 0;
z=ga(0,0,0);
*z->p=c-'0';
R z;
}
In this forms, it's still cryptic, but not quiet as inscrutable.N.B. - k/APL/J functions are often written on one line because that are often very short.
Mashing things up without line breaks isn't a functional vs. imperative thing [1]. Its a density vs. negative space thing. There's a quite a bit of research in many contexts indicating that effective use if negative space generally improves readability of most things for most readers, so while I know Whitney and his disciples like to say that J is "more readable" because of its density, I don't think that's generally the case for most readers.
[1] idiomatic Haskell is definitely functional, but does not eschew line breaks; mashing C programs into cryptic single-line monstrosities doesn't make it any less imperative, it just makes it unreadable.
Similarly, J programmers prefer to write
a =: a + b
instead of
(for i=0; i<a.size(); ++i) a[i] += b[i];
which obscures the idea quite a bit. Or saying
(# %: */) a
to calculate geometric mean of vector a and use that in next expression.
> a =: a + b
> instead of
> (for i=0; i<a.size(); ++i) a[i] += b[i];
Well, sure, most people prefer not to write code with syntax errors.
But, nitpicks aside, your post is mostly frequently-made commentary in favor of functional programming (both FP languages and functional style -- which, again, has nothing to do with avoiding line breaks or whitespace -- in other languages), but really has nothing to do with the characterization uphtread of the mass of minimally-whitespace, non-line-broken C code presented upthread as being a matter of "functional vs. imperative", which remains completely wrong. So its a complete nonsequitur to the post it "responds" to.
Not using whitespace, including line breaks, has nothing to do with functional style, and using it has nothing to do with imperative style. (If it did, your C++ example, minus the syntax error, would be just as "functional" as the J example.)
[1] For an example of the large whitespace for factoring-out, see the factoring out of J-sub-k on the top of p. 460 of http://titan.fsb.hr/~venovako/dist/tensor.pdf
And I always use "--lint" with "-f".
"noun" in k: noun:{[c]$[(c<"0")|c>"9";0;c-"0"]} / $[boolean; true; false]
The C version is effectively the k version expanded with declarations, memory allocation and initialization. I would not be surprised if the interpreter was written in APL and transliterated to C.
Arthur had learned how to think a problem in APL and was a wizard at transliteration.
I'm saying that it being in functional style has nothing to do with the absence of whitespace/newlines. Both the use of the functional style and the absence of whitespace no doubt reflect Whitney's language-independent programming preferences, but they are distinctly different things.
BC: Do you ever look at your own code and think, "What the hell was I doing here?"
AW: No, I guess I don't.
Of course, maybe he was answering the first 8 words.Of course, it is understandable. NOT.
"J is conducive to thinking about computation because a lot is expressed by combining even a few components. With J it is natural to think in code, not just about code."
What's an "interpreter fragment"? Most of the Google results for that term point to the same story.
I know this is Hacker News and everyone is supposed to know everything about everything already, but on obscure topics it's nice to take a moment and explain what you're talking about.
I can compile and run pretty much any K&R code from the C Programming Language in 1978 with no problem. Finding the appropriate compiler, I can do the same with BCPL, the Language and its Compiler (1981). And the code is actually understandable.
This, on the other hand, is an obfuscated mess that segfaults on any input, J or otherwise. I'll be damned if I spend 5 minutes debugging it, ex() segfaults with infinite recursion. It's not worth the trouble.
Not saying this isn't worth sharing as a curiousity or an artifact, but this is not a work of genius and it is not defensible or to be emulated. Do your colleagues a favor and don't write like this.
I'd second this - in any normal situation.
If you're really aspiring to form the future, strive to be extremely productive, demands that from your colleagues - that's another thing.
Sigh - industry as a whole isn't nearly on the level to use those advanced tools.
Is this a parody?
Is J a language for writing personal code with personal meaning only that you can only run in your head?
There is nothing professional about this. Sounds more like a mental illness.
In my experience, J is closer to thought process than many other languages. Really. You just have to try - and get used a little bit.
There are far better tools available these days for rapid prototyping; interpreted languages with REPLs and libraries for just about everything--without resorting to a write-only line-noise language and actually ending up with a usable product.
For that matter, I suppose it's possible to write understandable, maintainable TECO, APL and J with some effort if you don't value obscurity and cleverness over simplicity, correctness, maintainability and understandability.
If your business logic or product is written in line noise written by some lone J genius, your company is fucked when your guy inevitably gets hit by a bus, goes nuts, or leaves for greener pastures.
Do you think mathematical articles are understandable and maintainable?
What if there is not a lone J genius, but a whole - small - department of people, who spend time actually thinking and talking about computations, not the ways to express it?
APL was widespread in operations research in the 80s, and since the early 90s is usually found in and around finance, even though it isn't really common anymore.
Part of it is a mindset thing; People used to consider longer employment periods, both from the company's perspective and from the employee's perspective. Although APL is extremely useful when you use it properly, it is definitely not a "programmer-is-a-replaceable-cog" language that Java strives to be and the most firms now assume.
"For that matter, I suppose it's possible to write understandable, maintainable TECO, APL and J with some effort if you don't value obscurity and cleverness over simplicity, correctness, maintainability and understandability."
Yes. I do separate long functions in k onto separate lines. However, long functions in k are rare (and usually) an indication that I haven't arrived at a simple solution.
Writing in J/k/APL allows to think in chunks where the chunk is an algorithm - not all of the tiny steps in an algorithm. It is a higher level of abstraction - not unlike programming in C versus assembly language or machine code.
E.g.
The sum of a list (of any type) in k (APLis similar):
sum:+/x / read "plus over x" or "reduce x with plus"
versus (C code fragment) which computes the sum of integers only:
int sum=0: for(int i=0; i<n; ++i) { sum = sum + x[i]; }
Btw, John Von Neumann once chastised a grad student for writing an assembler. Perhaps real software engineers only write in machine code.
"If your business logic or product is written in line noise written by some lone J genius, your company is fucked when your guy inevitably gets hit by a bus, goes nuts, or leaves for greener pastures."
I think the larger issue is documentation. Undocumented code (of any kind) is a liability. When I write in k, Every line of k is documented (as text past column 80). In addition, there is header documentation that described the purpose and approach of the code in the file/module. The number of lines of code in a file is typically no more than 10 lines.
Here's a working adaptation that is quite close to the original. https://github.com/tangentstorm/j-incunabulum/blob/master/mj...
A current version of the k language still has traces of it inside. And there are still people on Wall Street running mission critical systems with it. http://kx.com/q/c/c/k.h
The kona project attempts an open source implementation. And they give some rationale on the paricular coding style. https://github.com/kevinlawler/kona/wiki/Coding-Guidelines