Papers from the Lost Culture of Array Languages (2011)
prog21.dadgum.com
prog21.dadgum.com
1. It implements an algebra: Operators take arrays as input and yield arrays as output.
2. The special characters bound to those array operators.
3. It is functional.
#2 definitely gave APL programs a distinct and elegant personality, but probably didn't help with widespread adoption.
But #1 and #3 are powerful ideas, and have showed up in other forms. SETL was a set-oriented language from the late 70s/early 80s. I took compiler courses from the guy behind it, R. B. K. Dewar (of Spitbol fame). Fantastic teacher. SETL programs were concise and powerful like APL programs, exactly because of #1 and #3.
Later in my grad school career, I got into relational algebra, which also has these characteristics. My advisor was T. H. Merrett who engaged on a many-years, quixotic quest to explore relational algebra as the basis of a general-purpose programming language. I didn't buy into all of that, but as a database guy, I find relational algebra to be powerful and useful on a day-to-day basis.
Finally, I really like functional stream approaches as a great way to bridge set-at-a-time programming (for lack of a better term) with ordinary, low-level programming. Java 8 streams and lambdas is a great (if flawed) example of this approach. Even Map/Reduce is, under the vast amounts of complexity.
I like this approach so much that I built a command-line tool, osh (http://github.com/geophile/osh) based on it. The idea there is the Unix idea of composing simple commands using pipes, except in osh, the pipes carry streams of Python objects. And the commands that operate on the objects in these streams are Python. E.g., to print the pids and commands of processes whose commandline contains "emacs":
$ osh ps ^ select 'p: "emacs" in p.commandline' ^ f 'p: (p.pid, p.commandline)' $
(37268, '/usr/bin/emacs24 ')
(42107, 'emacs ')
(63758, '/usr/bin/python /usr/bin/osh ps ^ select p: "emacs" in p.commandline ^ f p: (p.pid, p.commandline) $ ')
(113605, '/usr/bin/emacs24 ')I've been searching for the 'panacea' of shells (and/or auxilary shell tools/hacks that dup() fd's 0,1,2 to enhance existing shells) that hits that same 'sweet spot' you describe. As such, over the last ~15 years I've been through a litany of setups, ranging from:
- the standard "bash/zsh/fish" approach (where you extend the shell) to
- the "scsh/ipython/eshell" approach (where you bring an inferior shell's functionality into a language) to,
- the screen/tmux approach (where you take a shell and then layer functionality over it). I.e., for directory navigation, I'd written my own f-recency+bookmark system that would hook 'cd <tab>' and generate a pane sort of like Midnight Commander to nav around
I'm not sure where I'm going with this other than, I feel your pain and I'd imagine tons of other people do/did as well. Powershell is painfully slow and RAM heavy but the ability to add custom properties(!), providers, access the registry, and manipulate all of these objects as you'd like. Your project definitely looks like an interesting take on things as well. At least we're making some progress, I suppose ;)
===
(!) This is incredibly powerful since you can take a path, like C:\users\foo\downloads\video\, take file item, and then have Powershell invoke an executable to extend functionality out. If Windows doesn't have "Length" or "Encoder" as a property on the file out-of-the-box, you can just use an auxilary tool (say, ffprobe), "mapcar" the exec to the list-of-files, grep out the Length: field, and bam, that file now has Length. ``ls|where Length -gt 15'' ends up being pretty magical.
"A lot of it is, `Let's make this all more fun.' You know, `Math is not fun, let's reduce math requirements. Algorithms are not fun, let's get rid of them. Ewww - graphic libraries, they're fun. Let's have people mess with libraries. And [forget] all this business about `command line' - we'll have people use nice visual interfaces where they can point and click and do fancy graphic stuff and have fun."
http://web.archive.org/web/20080128144630/http://itmanagemen...
The above is from an interview with Dewar about a paper he co-authored in 2008 concerning CS education and programming languages:
http://web.archive.org/web/20080103143526/http://www.stsc.hi...
I am both an APL user (k) and a spitbol user but still not a Python user. Regarding the later, I think I suffer from the Dewar perspective[1]: it's too easy and too focused on libraries. Perhaps this is misplaced, but for myself I feel k and spitbol are more "educational" than Python.
1. In truth, I have no idea if Dewar saw Python as he did Java and PHP.
That's questionable. APL doesn't treat functions as first-class values. Its higher-order functions are limited to "operators", i.e. second-order and no higher.
I meant functional in the sense that the 1st-class objects (arrays, sets, relations) are not mutable. (My memory of APL and SETL is fuzzy, so maybe there is some way to alter state, but that certainly isn't the emphasis.)
For a while there was a freely downloadable version to try out, and you can look at "q for mortals", http://code.kx.com/q4m3/, to get some flavor.
Around that time, we moved much of the mortgage models over to kdb and q, which was fun. Lots of weird dark little corners in that language. (List of dicts? Nope! It's a table now! 'sum x' drop NAs, while 0+/x was NA if any value was NA, but maybe only for ints, not floats, etc. Ugh.)
Yet, this is how we program today with words, writing english words that don't mean what they claim to be. SomethingManager, SomethingDelegator, etc. Bah! The very ideal behind APL is to notations, to represent code with symbols which are non ambiguous and always mean the same thing.
There's a lot that we have lost. It's still worth taking a look at. For those interested, J (www.jsoftware.com) is a free version with lots of great learning resources.
Further, when you have a SomethingManager, is it really representable by a general symbol? What symbol would that be?
const items = ∅; // empty set
if (myItem ∈ items) { ... } // is element of
const union = items ∪ otherSet;
assert items ≌ ∅; // items are all equalThey are rather making an egotistical (but IMHO deserved) point that their work is the 3rd revolutionary call to action in the history of relational database systems.
That said, the assertion that all elements in a set be equal doesn’t really make sense — or rather, it’d be written as “|S| < 2” by virtue of the fact that a set is defined as containing distinct elements; consequently, in a set with more than one element, all elements are not equal. If you want to use mathematical formalisms for set operations, you’d instead have to work on sets of tuples and test for a predicate: “∀x ∈ S: p(x)”.
However most of the LINQ features using the plain method call API were already exposed in Smalltalk collections.
object.Where(cond).Select(something)
would be (object select [ :element | cond ]) collect: [ :element | element ] \in => ∈
\emptyset => ∅
This allows the user to type in using an input mode they're familiar with (standard text), but get an output that looks better (IMO).The main issue is that computers don't come with a standard default layer for math and APL characters anymore so it's a huge leap of faith for someone to start to use these characters.
This depends on the implementation. In Dyalog APL for example, the special characters (glyphs) are effected using fairly simple two-key sequences. Mnemonics are used to help learning. I found it fairly straightforward, surprisingly so.
The arary languages have always struck me as the extreme end of this. The semantics are very dense: each operator does a lot. Intermediate results aren't granted names. The few people who can follow this find it pleasant as the "extraneous" information is removed. Most people find it hard to adapt to. And it doesn't readily suit a lot of business logic.
Possibly the one context where arrays-as-language-primitive has really taken off is the 3D processing world, first with OpenGL's matrix stack computation and later with shader programming. Note that shader programming is the other way round: you don't start with a screen object and apply functions to it mapped across all the pixels (fragments), you write a small program as if for one pixel. Even if the end result is executed in a highly parallel fashion.
There are many more applications than 3D, most notably APL and all its descendants applications in analysis work, logistics, stats and so on.
Though Python is making inroads here.
I would argue that most programmers are fine with this. Unix pipes are a very similar point-free way to program that we're all comfortable with. If people came to APL thinking of it as something similar to pipes as opposed to something similar to C they might have an easier time with it.
> This is a very bottom-up programming methodology, a pseudo-variant of which has recently gained the monicker “refactoring”, even though that is a ridiculous term for a number of reasons that are not worth belaboring at the moment. I also think that the formal “refactoring” stuff missed the main point, but that’s also not worth belaboring. Point being, they are sort-of related, and hopefully you will understand the similarities and differences more over the course of this article series.
Um, what? Refactoring is called refactoring because that's what it is: You're pulling out a "factor" of the program and reifying it into a symbol which then replaces the factored code. "semantic compression" is a great name for it, but it's been the done thing in Forth, for example, for many decades.
It's just like factoring polynomials.
This is a textbook example of what I call "half-smart": He's smart enough to write this article, but not smart enough not to.
(Also that site disables reader mode. Boo! Why would you do that?)
I could touch type APL on the IBM 2741 text terminals (basically Selectric typewriters with serial interface.). And could compose simpler APL codes at that speed.
The principal weakness was APLs narrow view of data type. Text was a 3rd class citizen. There was no substring searching in the early IBM implementation. As a result one had to:
Form a cross product of text with the substring, testing for equality of each character
Rotate each row of the cross product by n steps where n is the row number
(This aligns the tests along columns)
Perform AND reduction along the columns to find the full matches.
None of this was intuitive and no I didn't invent that method.I didn't learn much, sadly, but it sure piqued my interest in programming.
APL is a great language, but it shouldn't be the first one you meet.
It's extremely intuitive!
⌊/(⍳≢b)⊖a∘.=b
Let me illustrate why with q: q)a:"this is some cake i like"
q)b:"cake"
Form a cross product of text with the substring, testing for equality of each character q)a=/:b
000000000000010000000000b
000000000000001000000000b
000000000000000100000010b
000000000001000010000001b
Rotate each row of the cross product by n steps where n is the row number q)(til n:count b) rotate' a=/:b
000000000000010000000000b
000000000000010000000000b
000000000000010000001000b
000000001000010000001000b
Perform AND reduction along the columns to find the full matches. q)(&/) (til n:count b) rotate' a=/:b
000000000000010000000000b
And there we have it. A naive string search in an array language.Now some interesting advantages fall out of rediscovering this:
- It's obviously parallelisable: The "obvious" (iterative) solution in other languages isn't, and without some difficulty it isn't clear where the work can be split up.
- It uses "outer product" with compare = instead of multiplication × -- an experienced programmer might forget how deeply satisfying it is when first learning how to compose operators
- With some thought it gives a programmer ideas on how they can do string search even faster. Fast string search in an iterative language doesn't look anything like the naive solution.
b⍷a (b find-in a) is useful enough that an experienced programmer should expect a "modern" language to implement it (in q it's called "ss", other languages of course call it different things) but it is far from necessary, and we might cheat ourselves of something. After all, when we are experienced enough to see these things as obvious, other things become obvious as well.
https://en.wikipedia.org/wiki/F-Script_(programming_language...
SQL is probably not the ideal "table processing" language, but it's good enough and a strong enough standard such that competitors have yet to unseat it (although I'd like to see more competition, such as Tutorial-D/Rel, SMEQL, etc. fight it out in the market-place.)
Fortran and Python are much more general purpose languages that also happen to have array operations to speed things up.
Matlab/Octave are more math oriented than any of the above and are far more than just programming languages, they are more like interactive notebooks, the IDE for Matlab (and for Octave too now, though I haven't worked much with it) is so closely coupled with the language that the language has no stand-alone right to existence.
APL, J and K are all more closely related to each other than to any of the others, the closest of the others would be R.
Consider the lone factorial problem: Every array programmer will do some kind of the product of, one plus array of all ints up to x
apl: ×/⍳X
q: prd 1+til x
j: */1+i.x
k: */1+!x
Notice how similar they are?Now, how do you solve this problem in Fortran? In Python? In R? In Matlab? What's the first tool in your toolbox? Is it to iterate over the values?
FUNCTION FACT(N)
INTEGER N,I,FACT
FACT=1
DO 10 I=1,N
10 FACT=FACT*I
END
def fact(n):
result = 1
for i in range(1, n+1):
result *= i
return result
function b=fact(a)
b=1;
for i=1:a
b=b*i;
end
function fact = iter_fact(n)
fact = 1;
for i = 2:n
fact = fact * i;
endfor
endfunction
fact <- function(n) {
f = 1
for (i in 2:n) f <- f * i
f
}
If the native programmer's first impulse looks something like that, then it's not an array language because the programmer isn't thinking in terms of arrays.program xfact integer :: i print*,fact([(i,i=0,5)]) contains elemental function fact(n) integer, intent(in) :: n integer :: i,fact fact = product([(i,i=1,n)]) end function fact end program xfact
this version respects
line breaks while
this version does not
respect line breaks. program xfact
integer::i
print *, fact ([(i, i = 0, 5)])
contains
elemental function fact (n)
integer, intent (in)::n
integer::i
fact = product ([(i, i = 1, n)])
end function fact
end program xfact julia> fact(n) = prod(1:n)
That's 9 characters, shorter than q. (`julia> const Π=prod` brings it to 6, only apl is shorter.)I just asked someone in my office and they sent me this:
function fact(n::Integer)
n < 0 && return zero(n)
f = one(n)
for i in 2:n
f *= i
end
return f
end
They seem to have google'd the result, but I'm not judging that: That might be how Julia programmers code...I'm not sure which is faster here, often writing out the loop is a little faster but less compact.
julia> @btime fact_prod(10000)
7.833 μs (0 allocations: 0 bytes)
julia> @btime fact_loop(10000)
7.908 μs (0 allocations: 0 bytes) def product(iterable):
return reduce(operator.mul, iterable, 1)
>> prod(range(1, 5))
24
(albeit with some imports, and note range is [begin, end) so that's 4!, not 5!). It might not be the first thing a Python programmer does, but it's certainly reasonable. Haskell of course is Prelude> foldl (*) 1 [1..5]
120
(inclusive this time, hence 5!) which is only slightly more verbose than the "array" languages, but does have the advantage of clearly specifying what your base case is if the list is empty.Perhaps array languages are just getting subsumed as a special case of "functional programming", be it either the relatively weakly-guaranteed FP that gets embedded into otherwise OO/imperative languages or Haskell's stronger-guarentees FP.
The rank operator in J (and probably Dyalog APL) allows you to treat a 3D matrix as an array of 2D matrices, a 2D matrix of arrays, a single item or a bunch of unit-sized boxes. This concept generalizes to higher dimensions. I don't think this aspect of array programming has gotten as much airtime as it deserves, probably because it is complex, but this is where the semantics of array languages and conventional languages really differ.
Indeed, I'd argue that if that was your primary use case you are and always were better off with a relatively custom language (relative to programming as a whole), because the needs are so different and the wins so big using an environment set up to support those needs that you'll want that in the end. You can fuzz the line with things like NumPy, because all these lines are fuzzy, but dedicated support will be a big win in the end.
>>> import numpy as np
>>> np.multiply.reduce(np.array([[1,2,3,4], [5,6,7,8]]))
array([ 5, 12, 21, 32])
and it lets you reshape into different forms, like: >>> arr = np.array([[1,2,3,4], [5,6,7,8]])
>>> np.multiply.reduce(arr.flatten())
40320
>>> np.multiply.reduce(arr.reshape((4,2)))
array([105, 384])
>>> np.multiply.reduce(arr.reshape((2,2,2)))
array([[ 5, 12],
[21, 32]])
I believe this was influenced by the array languages. prod(1:x)fact(n::Integer) = prod(1:n)
But I wouldn't call Julia an array programming language. Not everything is an array, there are scalars too, and vectors (vs matrices) are very special arrays that exhibit the expected duality properites (since 0.6)
It's kind of crazy to think of Matlab as not an array programming language since the array is literally the only data type.
>>> from numpy import np
>>> np.multiply.reduce(np.arange(1, 10+1))
3628800
That same programmer would (hopefully) also be aware that the function will overflow for large-enough values, and might instead write it at: >>> np.multiply.reduce(np.arange(1, 1000+1, dtype=np.object))In J for example, you could factor your code out. The way you factor mathematics. Their power lies in their notation.
Outside of the notation, If you have a multidimensional array, you can operate on ranks of it. You can write loopless code, you don't need for, do, while loops. The compounding and combining of verbs allows for a very functional style of programming that you can never realize in the languages you mentioned. You have to understand APL was designed for describing algorithms before it got turned into a computer language.
No. Is like claim "C" is a OO language. Is more correct to say it have "array capabilities".
To be considered an "X" language it must that "X" be the most primitive/idiomatic/natural/default way to think and program on that. And also a primitive of the language!
You can do OO on SQL, but that is hardly the idea!
Similar, with a array language you see everything as arrays and as array manipulation and rarely deviate from that. This is the same on other languages.
Even in multi-paradigm languages like python or C# maybe you put a little functional here or linq there but that portion of the code seriously look "alien"
Mind you, I cut my teeth on R, so needing to use loops in other languages confused the hell out of me at first.