Regarding scans (filter is a type of scan, howbeit easier to implement than the general case), there is in general an annoying work-span discrepancy, which bodes ill for contemporary computers with their finite parallelism. I would buffer heavily here, but fusing the scan with its input allows for the use of a work-efficient implementation, spending all latent parallelism on generating more inputs.
When I speak of loops, I am including any usage of rank (incl. implicit); in this respect, apl is chock-full of loops. Some loop optimisations may be obviated, of course, because of referential transparency, but others arise in their place. Take for instance some recent work[0] done on futhark: rather than parallelise an already in-place algorithm, they had to in-place an already parallel algorithm!
Two more points:
A problem may be bound by memory latency. I spent some time tuning j's I. recently, and it does many searches in parallel (4 for a small search space, 12 for a large one, iirc), but it is still bound by latency. If I could fetch a pivot corresponding to the first cell of y, and then generate the next cell while I wait, I would make better use of resources.
A compiler results in more transparent performance characteristics. When I target an interpreter, I must write my idioms and special combinations in exactly the way it expects, else nothing will happen; a compiler will care much less about exact phrasing. Similarly, like I mentioned, you get licm and cse for free, rather than needing to do them by hand.
0. https://futhark-lang.org/blog/2022-11-03-short-circuiting.ht...