Ode to J
zserge.com
zserge.com
I also follow tangentstorm and took some of this J-talks gui code[2], and mashed it up with Tyler's as a J learning exercise. I just put this up on my github. [3]
I highly recommend tangentstorm's YouTube videos too![4]
The concise J code, and the way it makes me approach a problem brings me joy! Given all of my coding projects in J fit on one page, I can revisit them later, and get right back in without worrying about having too many comments or documents to understand my code. I like refining the code over time for fun and learning even after it works, or somebody much smarter than myself shows me a different way to look at it.
[1] https://datakinds.github.io/2020/03/15/modeling-the-coronavi...
[2] https://github.com/tangentstorm/j-talks
It is definitely a different way to think. My programs are taking a long time to write as I train my brain to think in this domain, but the code is very concise and generally not hard to follow afterwards. Everything is so interactive it's a joy.
In particular, I enjoy how the poster broke the code down as I can only understand some of C code. It's cool to think I could create my own little J interpreter this quick.
The next time I write C, I'm tempted to do it Whitney style like this. There really is something to be said about seeing everything at once.
They say that APL was developed to mimic mathematical notation, and I see the resemblance. But real works of mathematics consist of one or two lines of expressions separated by sentences and paragraphs explaining the goal of those expressions.
So far in K "advertising" I only see the expressions and not the explanations. The explanation is there, but it's not a comment in the codebase; it's in a blog post that explains it. The code has many lines of expressions on top of each other, and nobody writes math that way.
I feel like I'm missing something, basically. It's clear how to write K, but what are the best practices for making a large codebase readable? I shouldn't have to invent that myself, right?
I wrote a toy simulator for the domain I work in using something like 6 lines of APL code and sent it to someone good with APL. They replied back to the email pretty quickly with some tips on how to improve it and could generally get the gist of what I was trying to do while knowing nothing of the problem domain. It was really an interesting experience. I think with zero doc, a colleague of mine would immediately understand what I was doing after going through a short APL tutorial. It isn't magic, but not nearly as cryptic as people make it out to be.
Edit: people do comment their APL code. I think a million loc project would be pretty difficult in APL, but the beauty is that you should be able to do some pretty powerful work in a few pages I would imagine.
The co-dfns source I've seen[1] is much longer. Do you know if there's a write-up somewhere of the differences between that and Appendix B from the dissertation?
I like APL and J as a scratchpad where arrays are the basic unit and not scalars. J is functional and it turned me on to that world before I touched Haskell or F#.
Aaron Hsu has a lot of great videos that speak to a lot of the usability and scaling out you mention:
https://www.youtube.com/results?search_query=aaron+hsu
I particularly like this one: https://www.youtube.com/watch?v=z8MVKianh54&t=2857s
I am able to grasp concepts or own them after coding them in APL or J even if the code isn't as fast such as how well APL applies to Convolutional Neural Networks [1,2]. I really understood the mechanics of CNNs better after working through this paper a lot more than books I had read on ANNs in general since the late 80s/early 90s. By contrast, I have coded ANNs in C and Python, and I get lost in the PL, not the concept, if that makes sense. Anyway, I am a polyglot and find people criticize J/APL/k etc. from a brief look without really trying to learn the language. I learned assembler and basic back in 1978 to 1982, and I felt the same way when I first looked at opcodes.
Yes, the takeaway is that with APL or J is that you can see the mechanics in a paragraph of code, and it is not a very trivial example. If the libraries or verbs are created to deal with some of the speed or efficiency issues, it is promising as a way of understanding the concept better.
The dataframes of R and Python (Pandas) were always a thing in APL/J/k/q, so it is their lingua franca or basic unit of computation upon which the languages were built - arrays, not a library.
More importantly, almost along the lines of the emperor has no clothes, is a tack to get away from the black box, minimal domain knowledge, ML or DL that cannot be explained too easily - see newly proposed "Algorithmic Accountability Act" in US legislature. Differentiable Programming and AD (Automatic Differentiation)applied with domain knowledge to create a more easily explainable model, and try to avoid biases that may creep into a model and affect health care and criminal systems in a negative way [1][2].
And then there are those who use DL/ANNs for everything, even things that are easily applied and solved using standard optimization techniques. Forest from the trees kind of phenomenon. I have been guilty of getting swept up with them too. I started programming ANNs in the late 80s to teach myself about this new, cool-sounding thing called "neural networks" back then ;)
I suspect the things that work ok in a <50 line application start to work less well when you're approaching 500 lines.
Funnily enough, the framework is actually a more expansive version of a tick system developed by Kx (the company that makes kdb) https://github.com/KxSystems/kdb-tick
The Kx one is incredibly concise. When I first started working with it, it took me a while to figure out what was going on.
The larger K/q programs get, the more they tend to look like "normal" code, but you still see a lot of these clever one liners hidden away in there.
Yep, that's what I suspected. TorQ confirms it.. giving everything a single-letter name will no longer do :-)
I'm a bit surprised at the number of comments that explain what the next line does. I'm not sure what to think of that, but it reminds me of beginner tutorials explaining code for people who can't yet read it confidently; for obvious reasons, not a popular style among more conventional languages.
https://github.com/AquaQAnalytics/TorQ/blob/master/code/proc...
I could imagine working on a codebase like that.
Because the code is concise and includes all the important information, you can view most useful programs in one page, so it doesn't take much to work through the logic again if you have not touched it in a while. I put comments inline for attribution to a source, or a quick mnemonic to unravel some tacit J code.
I loved J the first time I saw it back in 2011/2012. I had played with APL in the 80s. I have played with many languages (asm, Basic, C, Haskell, Joy, Forth, k, Lisp, Pascal, Ada, SPARK 2014, Python, Julia, F#, Erlang, R, etc.), and each paradigm shift has taught me to approach problems from many different angles. I use the language that suits my current need at hand. Frink is on my desktop at work for all of my engineering, unit conversion, small input program stuff. R/RStudio is there for my statistics work. Julia is replacing MATLAB for me. I wrote Blender 3D scripts in Python in the early 2000s to make 3D wood carvings from 2D photos.
J is always open on my desktop, and is more than a desktop calculator. It is my scratchpad for mathematical and whimsical ideas or exploration. See Cliff Reiter's "Fractals, Visualization & J" for fun [1], or Norman J. Thomson's "J - The Natural Language for Analytic Computing" [1]. I just bought Thomson's book three month's ago for about $35. There's now a crazy $925 posting on Amazon! Somebody's creating sales from HN!
"Mr. Babbage's Secret: The Tale of a Cypher and Apl" was also a fun book. It's not really an APL or programming book!
APL was developed to replace mathematical notation. Making it a computer language was almost an afterthought. Indeed, there have been some trivial mathematical proofs in j[1][2][3]. Perhaps those give you a better idea of the connection?
1: https://code.jsoftware.com/wiki/Essays/Trains#Proof_of_Compl...
2: https://code.jsoftware.com/mediawiki/images/8/82/A_Proof_in_...
3: https://code.jsoftware.com/wiki/Essays/Ackermann%27s_Functio...
See this Bret Victor talk, at 25:03. https://vimeo.com/67076984
The problem Cobol has is that uses some natural structures, but then becomes very verbose because it only uses a few, and they often can't be combined, so it is very repetitive.
It's not natural to say,
I will buy eggs at the grocery store.
And I will buy ham at the grocery store.
And I will buy cheese at the grocery store.
> The point of programming languages is to be more productive...But you should be productive not only in initially writing the code, but also in reading it later when you do maintenance.
And what you see with most languages is extensive commentary to help future authors understand what's going on, and that's partly because terse languages tend to be cryptic.
SQL isn't even one language. It's a family of incompatible dialects. I've never written more than the most trivial SQL that would even be valid (much less the same result) on any other SQL dialect. Even though it has ISO and ANSI standards, every database in the world requires proprietary rules and conventions and extensions. How do you quote identifiers? Are strings case sensitive? Is '' the same as NULL? How do you create an index? What data types are there? We can't even agree on the most basic aspects of syntax.
It's "incredibly successful" in the same way that early HTML and JS was. People wanted access to the underlying platform so badly they'll put up with a nutty design and gratuitous incompatibilities. They don't really have a choice, and many are going out of their way to build their own alternatives because the vendors won't.
I don't want "close to natural language" because it's too verbose. It's like reading an article in The New Yorker - it takes so long to get anything said that you forget what the point was.
Of course, "too terse" isn't the answer either. There's a sweet spot. I suspect that the sweet spot varies, depending on the person and the kind of code.
[1] http://www.eecg.toronto.edu/~jzhu/csc326/readings/iverson.pd...
And like the symbols in math, symbols in K (and presumably other APLs) have names and a programmer who learned the language can just read a line of code out loud; it sounds quite like natural language, but probably comes closer to describing the entire solution than a corresponding read-out of the 50-line chunk of C that performs all these little steps and manipulations of temporary variables and individual array members to arrive at the same result.
For similar reasons, modern programmers tend to prefer stronger abstractions like folds, maps, or function composition to achieve some terseness (and eliminate unneeded variables and manual iteration). These generally bring the solution closer to what its description in a natural language would be.
After that, it gets a bit more controversial. How much whitespace do you need? How long are your identifiers going to be? How many assignments will you nest in a single expression?
My experience is that it gets easier to work with terse code if you get into it, and once you're into it, it actually does save you time (in reading and writing) while more verbose style becomes irritating to work with. It's like reading an article that gets sidetracked and says a lot but doesn't ever seem to get to the point, and once it's finally over, you realize you didn't get the point because it was buried in fluff.
(Fwiw, my personal style isn't quite Whitney level, but it's way more terse than what we have at work, and honestly the verbosity of work-code feels counter-productive to me.)
And so, aside from my opinion about how it is to work with terse code, there's another point: making the trees smaller makes it easier to see the forest for the trees. And if you really need to study the trees (because one of them is wrong and you have a bug?), well, you can still do it (because you learned how to work with terse code).
In other words, the point is actually to improve readability! It's the exact opposite of deliberate obfuscation, even if the end result might seem similar to the untrained eye.
Down this thread ben509 writes that > what you see with most languages is extensive commentary to help future authors understand what's going on, and that's partly because terse languages tend to be cryptic.
I'm not sure I agree. What I see with modern language developments is that they're trying to empower the programmer to make the forest smaller by eliminating the trees or making them smaller where it's feasible. We're making things higher level and more terse (but the APL family is way ahead of any mainstream language). It's the low level languages that make code seem cryptic, because you get lost in the low level details.
I still work with C day to day and the actual comments in the code bases I'm involved in tend to reflect this: C forces you to deal with lots of details (trees), and it is painfully easy to see them but not see what's actually going on at a higher level (the forest). So people write comments to explain what's going on.
Next, can you tell me the point of art?
For me, it's easier to read APL with the symbols than to read J/K. Does it get better if you work a lot with J/K?
In any case, symbols were abandoned back in the day when everyone was still using 8-bit encodings. [1][2]
Nowadays you have e.g. Dyalog APL which does use APL symbols and you have short ascii-based sequences (easier to type than latex) for inputting them:
`u`y gives you ↓↑
(You can try this online at https://tryapl.org/)[1] Remembering Ken Iverson (http://keiapl.org/rhui/)
> For the first few months, the special APL characters and the ASCII spelling co-existed in the system. It was Ken who first suggested that I should kill off the special APL characters. I myself resisted for a few weeks longer, until the situation became too confusing, for reasons described in J for the APL Programmer.
[2] J for the APL Programmer (https://www.jsoftware.com/papers/j4apl.htm)
> J uses the 7-bit ASCII alphabet. It also makes non-essential use of the box-drawing characters in the 8-bit ASCII alphabet for display. Using ASCII avoids the many problems associated with using APL symbols. It allows J to be used on a variety of machines without special hardware or software, and permits easy communication between J and other systems.
It also has a wonderful amount of books on it, most of them Creative Commons-licensed now (as all/most of Iverson's books are, I believe). It's significantly easier to learn for someone new to array languages yet who doesn't have the sort of hands-on training you can get with APL & k; you could go into the woods with nothing but the J interpreter tarball for a week and come out pretty having internalized the language, and it's even easier with some of the other books available.
It's, of course, free software.
How is that the case? I don't really see a grid system or anything like that. I much prefer delphi or wpf or something.
Erm, jqt is nice, but it ain't that nice! It isn't exactly native either.
Btw Scott, I've been meaning to ask you about what tools you're using for your work for awhile, but don't see an email address anywhere on your blog.
It seems like you've tried Lush, Clojure, Lua (Torch 7), J, R, and a bunch of other technologies. I've had the chance to try some of these for hobby purposes, but nothing for work yet, so I would trust your evaluation more than mine. Have you given up on array languages all together?
Clojure was cool, but JVM is basically worthless to me.
J is an entirely different language and I think the doc is actually pretty decent. There are free books and tutorials and books you can buy on Amazon.
I don't know how to find the documentation for the k version that is shipped with kdb+ (free non commercial download). I only find broken links (it seems they removed the k documentation from kx.com website?)
This might help a bit:
https://code.kx.com/q/basics/exposed-infrastructure/
(Look up Q docs for monadic K operators by their english name)
"abc"?"caz"
2 0 3The problem is that those languages require a completely different shift in software from the common imperative/OO style to something more declarative or in Lisp's case, just a bit different.
In a typical university course, there isn't enough time to focus on learning those things and you stay so shallow with the material that the language seems useless. It seems like 1/2 of the Prolog subreddit is about homework assignments on things like list reversal. If you can already do that in Python/Java with a single method call, Prolog just seems like a really bizarre and inefficient method. Now once you take another step and see how it can figure out how to solve Sudoku without explicit instructions...that is cool.
It's really the same thing with electrical engineering when the professor is trying to teach us microprocessors and assembly at the same time and it all seems like a waste when C is much easier. The assembly method probably teaches better, but I need more than just a few weeks to synthesize that information. Especially when you're getting slammed by other hard classes like differential equations at the same time. So perhaps you do hate J, then again, maybe you just need to learn it on your own schedule and not to answer arbitrary test questions about something you can already do in Java.
People come out of a class that used Common Lisp thinking that Common Lisp lacks any looping construct (it has several). It looks like the professor took a SICP (which uses scheme, not common lisp) based class as an undergrad, never looked at any language in the lisp family again, and then 20 years later made up a syllabus using the half-remembered information using Common Lisp because they remember SICP had something to do with Lisp and taught it without actually trying to do any of the exercises.
The point of the Incunabulum is not really what it does (it's a REPL with a handful of half-implemented J verbs and no error checking), but the style it demonstrates. I think that rewriting a similar program in another language without even attempting to reproduce the style is rather missing the point.
If you're interested in Rust, why not figure out what a semantically-compressed style looks like for it?
It's a completely apples to oranges comparison, with the rust version having some error handling, support for longer identifiers, tests, relatively idiomatic style with line breaks and indentation where you expect it, variables and type names longer than one or two characters, etcetra.
You could compress it massively before going for the preprocessor. In fact, a few of the preprocessor macros just make the C version longer (when measured in lines) than it would be without those macros if lines were allowed to be as long as in the rust version (<70 chars vs 100 chars). The printf macro (used only three times) actually makes the C code longer both in bytes and lines (or equal length in lines if you retain the 70 column limit).
There's really only one macro (DO) that expands enough to save lines, but just barely. The shorthand for return saves quite a few bytes (but not so many lines) given that it's used everywhere.
This topic in particular tends to get flooded with a few shallow clichés as soon as it arises. This is to be avoided, because they have the effect of turning a thread into the same generic argument as the last N times it came up, which is tedious.
Better to respond to the specifics of a topic. Alternatively, if you want to learn, ask curious and neutral questions. But please don't pick one of the low-hanging bombs and toss it; the results are predictably boring. Not picking on you personally—we all have this reflex, and it's actually fine in in-person conversation, but bad for internet forums.
I am fine with higher-order functions (I do that in Haskell all the time), and letters for variables, but please, please, give your functions meaningful names. Or add a comment that explain what they do.
Code is not only for the person writing it (who perfectly knows what the letters mean, at least for a short time after he or she wrote it). It's also for others who have to read it some day.
Also, for K:
https://github.com/kevinlawler/kona https://bitbucket.org/ngn/k/src/master/
1: https://github.com/jsoftware/
2: https://bitbucket.org/ngn/k/src/master/
3: https://github.com/JohnEarnest/ok