What learning APL taught me about Python
mathspp.com
mathspp.com
Like this line of JS feels so much easier to read than that line of python:
ages.filter(age => age > 17).length
Directly translating this approach to python: len(list(filter(lambda age: (age > 17), ages)))
Although a better way to write this in python I guess would be using list comprehensions: len([age for age in ages if age > 17])
which I feel is more readable (but less efficient) than the APL inspired approach. Overall, none of these python versions seem as readable to me as my JS one liner. Obviously if the function is on a hot path iterating and summing with a number is far more efficient versus filtering. In that case i'd probably still use something like reduce instead of summing booleans because the code would be more similar to other instances where you need to process a list to produce a scalar value but need to do something more complex than simply adding.The actual apl implementation: +/age>17
Apl implementation of taking the length(shape) of the filtered list: ⍴(age>17)/age
(ages > 17).sum() count(ages > 17)
and storing the ages > 17 is done with is_adult = pack(ages, ages > 17)I got curious and checked this in Rust, the generated assembly is the same!
×/ product reduce
⌊/ min reduce
⌈/ max reduce
∨/ logical OR reduce (any bool set)
∧/ logical AND reduce (all bools set)
,/ catenate reduce (join without spaces)
These all show a pattern of connection clearly where the Python names don't, that they are related operations; that suggests that you could put any function on the left or any kind of array on the right and see what happens. And they work over multidimensional arrays - and you can swap / for ⌿ to reduce down columns instead of accross the rows.Your rewritten Python and JS, by showing the operation as length instead of sum, and making a shorter filtered list, hide the connection even further instead of helping to reveal and clarify it.
> "which I feel is more readable (but less efficient) than the APL inspired approach"
I know how to read it, but just look at:
age age ages age
len age for age in ages if age
what's readable about so much repetition, what's readable bout having to spot the single character plural change in the middle of 8 short words? ([>])
that's more symbols than the APL one has, the language people reject because of the heavy use of symbols(!). Why does the [] indicate loopy-listy code but so does 'for'? In PowerShell arrays have a property .Count to use instead of Length - what's readable about counting the number of ages by indirectly looking at the length of something? Is this "it's readable because I'm familiar with it" rather than "because it's objectively readable"?julia> count(age > 17 for age in ages)
Or even shorter:
julia> count(ages .> 17)
Under the hood many of the functions are implemented using the reduce function (similar to / in APL):
julia> reduce(+,ages .> 17)
I think many languages in data science have been influenced by APL. But sometimes this influence come from an intermediate language like matlab.
myfun(list)
or list.myfun()
and this applies without exception. This means a lot of C-library code gets easier to read when used from D: Vector3 vec = (Vector3){ 1, 2, 3};
Vector3 result = Vector3Add(vec, vec2);
becomes auto vec = Vector3(1, 2, 3);
auto result = vec.Vector3Add(vec2);
I know its slightly off topic, but I want more people to use D, I really prefer it to C, C++ and Rust.On its technical merits, one of the best languages out there. It gives you insane performance but is so much easier to learn and write than C++ or Rust.
from toolz import curried, pipe
pipe(ages,
curried.filter(lambda age: age > 17),
list,
len)
This way you can follow - vertically - what happens, where, and why.You can make your own pipe function to avoid 3rd party dependencies
import functools
pipe = lambda *args: functools.reduce(lambda acc, el: el(acc), args)np.sum(age>17)
(Assumes age is a np.array)
OP here; maybe I'll add a comment to the article to make the comparison between NumPy and APL for this expression. Thanks!
length $ filter (> 17) ages
And you can also count pred = length . filter pred
count (> 17) agesselect count(*) from ages where age > 17
E: And I guess the lambda and explicit list-cast too. cf.
(count (filter #(> % 17) ages)) ages.reduce((a, c) => (a + (c > 18)), 0)
Though usually in production code you'll see this as something like ages.reduce((acc, cur) => {
return acc + c > 18 ? 1 : 0
}, 0)
Which creates some more noise but ups readabilityI think APL's ability to lift loop patterns into tensor patterns is interesting. It certainly results in a lot less syntax related to binding single values in an inner loop.
To do this he invented the APL notation.
If you find the article interesting, you might enjoy my curation of his work "Math For The Layman" [0] where he introduces several math topics using this "Iversonian" thinking.
[1] Look this up to install the J interpreter.
[0]: https://asindu.xyz/math-for-the-lay-man/
[1]: https://code.jsoftware.com/wiki/System/Installation/J9.4/Zip...
It could have been a good career choice at one point, given how many FTE's were devoted to the language at Merrill Lynch, Morgan Stanley, Lehman etc but I took the other fork
They amount to the same, but the explicit loop was much easier for me to understand (and I’m still not sure if one applies a function to a value or a value to a function, so I never remember if the function or the values go first in an apply filter call)
sum(1 for age in ages if age > 17)
with the other method you're treating a boolean as an int. Weak typing. sum(int(age > 17) for age in ages)
Every nanosecond is vital! python -m timeit 'sum(1 for age in range(100000) if age > 17)'
50 loops, best of 5: 5.08 msec per loop
python -m timeit 'sum(int(age > 17) for age in range(100000))'
50 loops, best of 5: 7.96 msec per loop
python -m timeit 'sum(age > 17 for age in range(100000))'
50 loops, best of 5: 4.78 msec per loopPlus here on each iteration `int` has to be loaded from the globals before it can be called.
[1]: https://docs.python.org/3/library/functions.html?highlight=s...
[2]: https://docs.python.org/3/library/stdtypes.html#boolean-valu...
if 23:
print('hello')
and it'll print 'hello'. But I'd prefer a strongly typed approach where this code would give an error saying, 'a bool is expected here'. Sure it's a subjective thing, and this is just my preference.“I don’t know the type hierarchy used in language X” is not the same thing as “language X is weakly typed”.
Python 3.11.4 (tags/v3.11.4:d2340ef, Jun 7 2023, 05:45:37) [MSC v.1934 64 bit
(AMD64)] on win32
Type "help", "copyright", "credits" or "license" for more information.
>>> isinstance(True, int)
True
>>> isinstance(False, int)
True
>>> issubclass(bool, int)
TrueAmen! It's of course also a C language tenet, and a great one. Life is so much simpler and more flexible when true and false are 1 and 0. It drives me crazy when I need to use a language where the logical operators only work on bools and the arithmetic only on ints, or some coercions work and others don't. When I incorporate somebody else's code into mine, first thing I get rid of is anything called "bool", a completely useless type. (as a nice side effect, that frees up the bool keyword for Boolean sets, which are quite useful)
a disappointment with unix is that process retval has this a bit backward, 0 is success, nonzero is failure (probably because errno does want for more bits than a singleton) but it's easily enough remedied with a !
I did love everything else about APL for the brief time I used it long ago (except the difficulty of entering the symbols)
Not that tightly. Which is why C and Python, which both started out with boolean values just being integers, eventually retrofitted booleans to the language. Conversions come back to bite you.
If you have suggestions for improvements for the article, let me know here! Thanks again.
How APL made me a better Python developer by Rodrigo Girão Serrão: https://www.youtube.com/watch?v=tDy-to9fgaw
If the allegedly "most popular language" (confirmed by Gartner and Netcraft) needs that much proselytizing, perhaps it is artificially popular? Or has voting rings?
len(age for age in ages if age > 17)
sum(1 for age in ages if age > 17) count_over_17 = [age > 17 for age in ages].count(True) def count(it):
return sum(1 for _ in it)
Which is basically just putting a friendly name on the approach from the article.For a safe version, you probably want to wrap it in another generator that bails with an exception at a specified size so you don’t risk an infinite loop.
def safe_count(it, limit=100):
# returns None if actual length > limit
nit = zip(range(limit+1), it)
if (l := sum(1 for _ in nit)) <= limit:
return l
Of course, you can just convert to a list and return the length, but sometimes you don’t want to build a list in memory.Not what actually happens but conceptually:
ages = [17, 13, 18, 30, 12]
sum(age > 17 for age in ages)
=> sum([False, False, True, True, False])
=> sum([0, 0, 1, 1, 0])
=> 2 # via conventional summing
Since True and False are 1 and 0 for arithmetic in Python, this is just a regular sum which also happens to produce a count.If your input was a numpy array to begin with you could skip the array comprehension, and shorten it to numpy.count_nonzero(ages > 17), since numpy automatically broadcasts the comparison operation to each element of the array.
from collections import Counter
total = Counter(range(10)).total()
assert total == 10 age >= 18
If your code is specifically about the magical age of adulthood then it ought to include that age as a literal, somewhere.It becomes more obvious when you consider replacing the inline literal with a named constant:
CHILD_UPTO = 17 # awkward
compared with: ADULT = 18 # oh the clarity
My fellow turd polishers and I would probably also add a tiny type: Age = int
ADULT: Age = 18 # mwah!
(The article was a good read, btw.)Also, Python is a wonderful functional language when used functionally.
Python having a statement heavy syntax and making complex expressions (while possible) awkward is the problem with its anonymous functions, not the fact that its anonymous functions are limited to a single expression.
f = lambda x: [
x + y
for y in range(x)
if y % 2 == 0
]
>>> f(5)
[5, 7, 9]
Lambdas which perform multiple sequential steps are fine, since we can use tuples to evaluate expressions in order; e.g. from sys import stdout
g = lambda x: (
stdout.write("Given {0}\n".format(repr(x))),
x.append(42),
stdout.write("Mutated to {0}\n".format(repr(x))),
len(x)
)[-1]
>>> my_list = [1, 2, 3]
>>> new_len = g(my_list)
Given [1, 2, 3]
Mutated to [1, 2, 3, 42]
>>> new_len
4
>>> my_list
[1, 2, 3, 42]
The problem is that many things in Python require statements, and lambdas cannot contain any; not even one. For example, all of the following are single lines: >>> throw = lambda e: raise e
File "<stdin>", line 1
throw = lambda e: raise e
^^^^^
SyntaxError: invalid syntax
>>> identity = lambda x: return x
File "<stdin>", line 1
identity = lambda x: return x
^^^^^^
SyntaxError: invalid syntax
>>> abs = lambda n: -1 * (n if n < 0 else return n)
File "<stdin>", line 1
abs = lambda n: -1 * (n if n < 0 else return n)
^^^^^^
SyntaxError: invalid syntax
>>> repeat = lambda f, n: for _ in range(n): f()
File "<stdin>", line 1
repeat = lambda f, n: for _ in range(n): f()
^^^
SyntaxError: invalid syntax
>>> set_key = lambda d, k, v: d[k] = v
File "<stdin>", line 1
set_key = lambda d, k, v: d[k] = v
^^^^^^^^^^^^^^^^^^^^
SyntaxError: cannot assign to lambda
>>> set_key = lambda d, k, v: (d[k] = v)
File "<stdin>", line 1
set_key = lambda d, k, v: (d[k] = v)
^^^^
SyntaxError: cannot assign to subscript here. Maybe you meant '==' instead of '='?Don't need return in a lambda. Assignment now has the walrus. Leaves raise and few other odds and ends.
Do believe you've cracked it! Don't think I'll use it much, but you never know... in a pinch.
You're right that the walrus makes assignment more usable; we can also call methods like .__setitem__ to get similar effects. Unfortunately the walrus seems to suffer the same broken/ambiguous scoping as assignment statements, e.g.
>>> a = 1
>>> b = lambda: (print(a), a := 2)
>>> b()
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File "<stdin>", line 1, in <lambda>
UnboundLocalError: cannot access local variable 'a' where it is not associated with a value
> Don't need return in a lambda`return` is needed for early returns. Notice that my `abs` example is trying to do that with `return n` (to skip the `-1 *` operation, since `n` is non-negative in that case).
> Leaves raise and a few other odds and ends
Off the top of my head: yield, match, def, import, class, with, for, while, assert, try, except, async, await.
b = lambda: (a := 2, print(a))
Hmm, need to put the assignment before. Can't access nonlocals unless you only read the the var and don't write to it. That's the way functions in python work as well.No, that would not have the intended behaviour: your `print(a)` is reading a fresh local variable, defined by the `:=` expression (which shadows that from the outer scope), so it will output '2'. The intended behaviour is to print '1' and reassign the outer name; but we can't do that (that's why I chose it as an example!)
> Can't access nonlocals unless you only read the the var and don't write to it. That's the way functions in python work as well.
Yes, that is precisely what I meant when I said it "suffers the same broken/ambiguous scoping as assignment statements".
Python took a haphazard, WorseIsBetter approach to assignment/declaration/scoping. The resulting behaviour is complex, confusing and involves spooky action-at-a-distance (e.g. the meaning of my `print(a)` expression was altered by the presence of an ':=' expression which hadn't been reached yet!). Since this doesn't really affect simple "top to bottom" scripts, it only became an issue once large applications and frameworks started emerging; by that time it was too late to fix, since there was too much Python code in the wild to justify such a large breaking change :(
However, I don't think complex functionality belongs in them either, so not a big loss. See the Beyonce rule mentioned elsewhere in this thread.
Crippled lambdas, no currying, "match" is a clumsy statement, weird name spaces and a rigid whitespace syntax. No real immutability.
"As such, curry is more suitably defined as an operation which, in many theoretical cases, is often applied recursively, but which is theoretically indistinguishable (when considered as an operation) from a partial application."
(while I see partial as having value, I'm struggling to see if currying would really be a useful addition to a scripting language)
That plus named tuples and a little discipline gets me 80+% of the benefits of pure functional.
The biggest thing to me wasn't in your list: lack of tail call optimization means you have to be a bit careful about where you use recursion.
I personally think long anonymous functions are an anti-pattern, naming your functionality is great for readability. I learned this from Wolfram language which makes heavy use of anonymous functions, I would often find myself returning from lunch to find my code which was perfectly clear that morning had become unintelligible. Today I try to limit anonymous functions to re-ordering arguments or "picking" functions that pull simple values out of more complicated data structures.
[0]:https://en.wikipedia.org/wiki/Miranda_(programming_language)...
If this was punch card input, or 110 baud teletypes, where program listings come back at a snails pace and use paper, then APL is great for that.
So if your typing speed hasn't gone up in proportion to the increase in Baud, I'm guessing your reading speed also hasn't gone up tens of thousands of times, and your ability to hold working state in your head hasn't gone up thousands of times, what is the advantage of increased Baud to code readability which you are talking about?
Let's say I'm not disagreeing, but I'm trying to dig into what specifically the change is which makes the difference; the computer can display more code at you per second than 1950 but humans can't read much faster than 1950 so that doesn't seem like it will help. Presumably longer books aren't inherently more readable than shorter books?
Can it be that Python is more readable because it lets you skim over and not read more of the code? Since not-reading isn't reading, it seems like 'more filler' that you don't read isn't what adds to readability.
It presumably isn't that Python is more English-y because languages which tend towards English words (SQL, Objective-C, PowerShell Cmdlets, Applescript, BASIC) are often maligned specifically for that reason, and because Python isn't English - you couldn't speak it to Shakespeare and have him understand you).
It presumably isn't because Python uses fewer symbols, or we'd all love to write Java style var1.Equals(var2) instead of == and var1.Plus(var2) instead of + and people seem to dislike that also. Why would + be preferred over .plus() but .sum() be preferred over +/ ?
Is it that Python has more visible structure to hang understanding on? Is it that it's more like walking compared to jogging compared to sprinting, that one can sustain a lower effort 'slower read' for longer?
Further proof of that: APL began as a mathematical notation on chalkboards, and only later was it decided to implement as a programming language.
As an aside, Unix also began in the era of 110 baud teletypes, which motivated brevity of many of its common commands (ls, cat, etc), and those brief names have been criticized a lot -- but people still use Unix/Linux; the brevity is a side issue, not a deciding point. As usual with most things.