Arthur Whitney's 'B' Language
kparc.com
kparc.com
I would recommend everyone doing C to have a very careful look at this. The code is full of nice tricks and you start appreciating the terse style after a while. I understand why this style is not more commonly used, but I really like it.
I recommend you to keep the original files at hand for reference. Once you have understood it the first time, they are a better reference than my many pages document. When AW says he hates scrolling, it is not because he is lazy, it just doesn't go well with this style.
I am sure there are mistakes. If you find any, please let me know.
> That was the easy part. Now it gets complicated.
Haha lol'ed. These notes are amazing. Thank you!
Array languages are the most "obscure" of all paradigms I have looked into. Even concatenative ones have some few resources about how work.
Also, the community live in a parallel world! So when asked about stuff, redirect me to https://code.jsoftware.com/wiki/Essays/Incunabulum or https://github.com/kevinlawler/kona as if it somehow make everything so obvious ...
P.D*: Because I'm building a language that try to draw some ideas from kdb
Have you tried looking at APL books? There are dozens of books on APL you can find. There are journals and conference proceedings from ACM, IBM, Dyalog, and Vector UK. I started to get into APL more seriously last year and lack of resources is definitely not a problem in my experience, lack of time to read them all is. Also, it's more about thinking of algorithms in terms of matrix and array operations. A lot of the techniques and algorithms can be easily adapted from Matlab.
However, although I totally get your point (I've been there), I have to say there is much more to learn if you try to understand that parallel world. It can take a while, but I think it is worth it.
You got me thinking how AlphaGo reinvented the style of Go for human players. A real AI taught one of our programming languages would code in a way that would be alien to us.
Whitney’s code is fascinating. I thought that someone might use a good C IDE to rename symbols to longer human readable names, but, that would harm the terseness of the code.
Talking for myself, I would love a FOSS version of k, but I am sure as hell I would patch a couple of things, and I am sure I would write much worse code than Whitney's (in his style or in my style).
There's a lot of material on their website, though it's not well organized and some of it is out of date.
https://code.jsoftware.com/wiki/Main_Page
I find it rewarding to screw around with, though the only jobs you're gonna get are K/Kx related.
I like terse code and my style is closer to the early UNIX/K&R, yet I've noticed a lot of people already find that too terse; the "normal" style these days seems to be gradually getting more and more verbose (look at typical C# or Java, for example.)
IMHO the readability argument in favour of verboseness is a bit of a misnomer --- you can easily "understand" a file full of extremely verbosely named functions all consisting of a single line, but that doesn't help at all with seeing the big picture. Judging by the amount of file-flipping and stack-jumping when I have to debug a project in that style, it probably hinders it.
On the other hand, this is "code with a learning curve". You can't glance at it and understand immediately, but once you do take the time to read it carefully, you can understand more of the whole than if it was written in a more verbose style.
There's also a lot of code that's so tightly coupled to some business domain that it's natural to write it in the language of the business. That way you don't have to maintain separate mental models and translate between them.
https://code.jsoftware.com/wiki/Essays/Incunabulum
Once I realized it was K&R C, I thought it was pretty straightforward.
If you want to see APL expressed as C primitives, something like Nial is pretty good: https://github.com/danlm/QNial7
Whitney/kparc previously covered on HN:
Downloadable now: https://www.reddit.com/r/apljk/comments/b1l5hi/shakti_trial_...
Didn't get much attention on HN: https://news.ycombinator.com/item?id=19326007
// :+-*% ^&|<=> x64 JJ Jj o2 cc tst RET cll psh pop acdbsbsd89.. o[m[s|d]] c3 eb+1 e8+4 e9+4 [f2/66/4*][0f] 5* 7*+1 b*+4 0f8*+4 03 23 2b 3b (6b 83) 89 8b ..Not sure what was on gitlab, I assume not source?
If it is anything like k, the "shakti" binary is probably small, maybe under 260KB.
Is there a way to download just the shakti binary?
What are the restrictions on the trial version of shakti?
Is it time-limited (e.g., 30 days)?
Does it need an internet connection to be able to phone home? (I think FD/Kx started doing this with their 64bit trial.)
The backstory for all this is that Whitney sold his remaining interest in Kx around July of last year for around 53 million.
Quick US trademark search does not show any filings for a "shakti" mark that covers computer software.
bin/k is 185K
Alas, I am getting "Illegal instruction (core dumped)". Will have to try another kernel.
No phoning home etc. although license does say you're limited to 30 days (although remains to be seen if/how that is enforced)
FWIW, kx didn't _have_ a 64bit trial before they released it with the phone-home stuff, before that trial was limited to 32bit only (still available & no connection required)
If I recall correctly, the 32bit trial used to time out after some number of hours after it was launched. This prevented, e.g., running it continuously as a server. Not sure if they still do that.
Going further back, I also recall early k trials that put limits on the size of the workspace.
miniconda is just the package manager of anaconda, with no preinstalled packages.
:)
b[Ii]{h:#x;l:0;while(h>l)$[y>x[i:/l+h];l:i+1;h:i];l}
b -> function name
[Ii] -> type declaration
h:#x;l:0; -> initial definitions
And the rest of it are bog-standard binary search operations with implicit return of l. Honestly, it's more understandable than Java with annotations or other magic constructsAt first glance, it looks unsightly but when one actually starts reading it (all of it, not just looking for "clues"), how it looks makes little difference.
These threads that mention k are always entertaining. There is usually some amount of comments like yours.
There is this concept called "Tall Poppy Syndrome". I am not sure what we might call this stuff here (it should have name), but we see some of it every time k comes up in a thread on HN.
Fortunately we also see people commenting who like APL or at least are curious. Not everyone is trying to cut Arthur Whitney down for being good at what he does.
sentence(lang HumanLanguage){
if(lang!="EN"){
throw_exception(UNKNOWN_LANGUAGE_EXCEPTION);
}
with subject(combine_adjective_noun("context", "clues")){
verb[to_be::present::plural]"are"([adjective]"overestimated");
}
full_stop;
}
Do you find this more readable than a simple English sentence?AW's code is way too terse, but most code is much more verbose than I'd like it to be. Finding the sweet spot is difficult (and it may change when you are coding just for yourself and when there is a larger audience). In order to find a good compromise, it is always interesting to explore all the possibilities in both directions.
We have code with many safety measures, extensive comments and visual clues and we have code like this and, in contradiction with every theory, the terse style is doing quite well in terms of bugs and programmer efficiency. Why? Is that a coincidence? Could we have the best of both worlds?
I do not say you should use this style for your projects, you don't even have to like it a little bit, but criticizing it because of its disadvantages without trying to understand its advantages is not a very productive point of view, in my opinion.
In short: Like Whitney, I hate scrolling in code. Multiple high res monitors and with bunches of tabs and files open and then scrolling and trying to figure things out with IDE ‘jump to implementation’, ‘find references’ etc is just really not very efficient compared to this. In My Humble Opinion.
As an inveterate simplifier of code and its formatting for clarity, I have gone to only doing a single thing on a single line, even variable declarations [Note I only have personal projects now]. The primary purpose behind this is that such extreme simplification will make it easier to process algorithmically, for both meta-code (i.e. IDEish) work and to narrow the horizontal extent of its format on both paper and screen.
Now that I think about it, it seems that Whitney's perspective might flow from his not using a code-folding editor, where each chunk of code can be collapsed to its header comment that states what it does. Aren't we usually dealing with the higher-level flow of logical chunks instead of the itty-gritty details of each?
We programmers often seem to forget that we spend far, far more time reading code than writing it, therefore readability should be a (if not the) primary consideration when engineering it.
That said, I do understand where Whitney's terseness impetus comes from: I have always lamented how little of my code will fit on the screen, with the number of vertical lines being the limiting factor; that is precisely why my taskbar is on the left-hand-side. (And don't get me started on monitors designed for watching wide-screen movies -- 1600x1200 FTW!)
As such, I can't really fault his intention as much as his execution, and I feel that ultimately the solution to his (really, our) problem is that all our code (not just his) needs an IDE that can expand such terse definitions as per the current programmer's preferred style. This has been my perspective for many years, especially after wrangling various SQL dialects and the mostly awful formatting preferences of my peers.
As I see it, the ideal solution is to store all code files as its token stream and have a default format that can be customized by each programmer in the IDE. As a result, each program will be stored as its pure content (which would help with version control (so long as whitespace is ignored)) AND each programmer would get to work with their preferred perspective. Of course the problem with this methodology is that the tokenizer and formatter better be flawless or you're f*ed, not to mention the fact that there are various IDEs that people like to work with.
I seen a few bits and pieces from Geocar but never seen a release or download.
That code is extremely readable - if you assume it will take just as long to understand as the entirety of the GCC compiler: https://github.com/gcc-mirror/gcc
It is the mathematical notation, put into, as Dijkstra expressed, "technologies of the past", completed (to be Turing-complete) and made executable.
At least, I think, this is the basic idea.
I guess there should be a niche for people like that, but the APL/post-APL community seems to think that Arthur has special computer science knowledge that can be best (or only) expressed in this form. We all know that's bullshit.
What programming is about, and what's hard, is telling the next programmer what the machine does.
This makes the easy part a bit easier, and the hard part much harder, than using C. And I already disliked C to begin with, for similar reasons.
What it says is that the hard part to make the next programmer understand what the machine does is not reading code, it's code complexity. Training yourself to read terse code is O(1), understanding code is at least O(N) where N is the number of lines of code.
Before you start arguing, try it. Try writing code as simple and succinct as possible. You won't go back to anything else because it works. It does make programming simpler and more productive.
And the I had exams, and did not touch the computer for two weeks. When I came back, I forgot it all. Cx? Dzy? Those names were meaningless.
This was a very importantly lesson - programs must be maintaineable. Unless you want to spend your entire life working on a single piece of software, you want to be able to context switch, and do it quickly.
A programmer doesn't rely on variable names to understand what the machine does, or very little. The contrary is a myth. In other terms, you always check what a variable contains anyway. That being said, I make a distinction between obfuscated code and less code. The former is about presentation, which in my opinion you're free to decide while the later is about #node in the ast. Sort of.
This is 100% real, and not a myth.
APLers use extremely consistent, stereotypical naming conventions which also carry semantic information. Think "Hungarian Notation" except it's just the prefix part. With consistently applied terse names, the same idea will result in the same code. Easier to visually pattern-match.
It's also worth considering that longer names are not necessarily more meaningful. It's pretty common for everyday programmers to use single letters for something like a loop induction variable. Sometimes longer is just... longer. Consider these three semantically identical K definitions:
a:{x*x}
b:{[n]n*n}
c:{[number]number*number}
Is that third version really clearer than the first?To get to your example, the preferred signature would be:
square:{x*x} (+/*)\:
is matrix multiplication. The following computes the transitive closure of a binary relation (represented as a boolean square matrix): {x|x(|/&)\:x}
Its core has a similar "shape" as matrix multiplication: (|/&)\:
So in a sense variable names don't matter all that much! To grok this code you have to stop thinking in terms of item at a time operations and start thinking of operations over collection of items. If you know how to write unix shell pipelines you are already somewhat familiar with this paradigm except that a shell doesn't provide most of the more useful features of array languages!There is nothing inherently disadvantagous to using abbreviations as long as you have a key to the abbreviations.
Even granting that, this:
> f[iII]{m:0;$[k:-x;W(x){x-:1;j:y;N(k)y[i]:y[+i];y[k]:j;j:f[k;y;z];$[m<j;m:j;]};{N(#y)z[i]:y[i];W(j:z){m+:1;N(j){k:z[i];z[i]:z[j];z[j]:k;j-:1}}}];m}
Is not a single line of "terse code". It's a bunch of lines of code with the newlines and spaces removed.
> Try writing code as simple and succinct as possible.
I already strive to do that (I don't always succeed). But the operating word here is as possible. Not "above anything else".
Just to confirm, I'm saying that after the user, the most important thing is less code.
In more practical terms, I'd say try "above anything else" to see where the actual limit is. From my experience, it's way further from where traditional programming is.
What is that called then?
I am interested in controlling hardware, whatever that is called.
I think Torvalds once said in a comment that is why he wrote his own kernel. He wanted better control of his hardware. Something like that.
"... telling the next programmer what the machine does."
You mean what the code does?
This sounds like drudgery. I can see why one might be overly sensitive about how source code looks before even attempting to read it.
I am just a hobbyist writing programs for myself. There is no "next programmer".
I have found the best thing I can do for the "future me" is to keep programs relatively small and keep the number of files low. This is one reason I like terseness and the idea of keeping "everything" (almost) on one page/screen.
I am just not capable of understanding a large project as deeply as I would like. I thought this paragraph summarises the issue well:
"To them, once you've sufficiently studied that screen or two of code, you can understand all of it at the same time. If it's spread out over thousands of files, it's very difficult to understand all of it, which leads to bugs, unnecessary abstraction, and the need for advanced tooling just to work with your own project's code."
This is from https://github.com/tlack/b-decoded