Ruins of forgotten empires: APL Languages (2013)
scottlocklin.wordpress.com
scottlocklin.wordpress.com
The language changes the way you solve problems computationally. It changes your focus away from the mechanics programming to inmerse your mind in the problem domain. It is hard to understand this without working with APL for a year or more.
There's another element here that is often ignored: Notation.
APL uses a unique set of symbols and it is these symbols that give you expressive power not found in text-based approaches. A paper by Ken Iverson himself, titled "Notation as a tool for thought", covers this:
http://www.eecg.toronto.edu/~jzhu/csc326/readings/iverson.pd...
The value of notation has been largely forgotten today due to the fact that APL suffered a popularity gap of probably twenty years due to the language being way ahead of technoloy at the time. The language didn't fail. The hardware of the time did.
This, BTW, is the reason I think J is an absolute abomination. Yes, Ken Iverson was behind J. The motivation behind going for ASCII and losing the very symbols that made APL notation so powerful was to attempt to deal with the limitations of computers at the time. For example, on some systems you had to manually reprogram the graphic card's character ROM just to be able to display the symbols.
Iverson sought to make the APL concept more accessible by transliterating to ASCII when creating J. In doing so he utterly destroyed the very notation he understood to be critically valuable as a tool for thought.
For more than one reason it would be wrong to attempt to bring APL back today. The concepts, the approach and a form of notation for the expression of computational concepts could and should come back.
I'm reminded of this constructive review of TempleOS [1] and thinking YES I'd like to have some of those code-writing features.
I'm also reminded of Perl, a language that grows multi-character operator glyphs at a rate only slightly slower than teenagers invent emoticons. Imagine if Perl could just be programmed with proper alternative symbol support like APL.
Having proper symbols might also stop the minor endless war of deciding to use "=" as both assignment and evaluation.
When dealing with analysis workloads, his line of thinking can yield fast runtimes and doesn't bring that much of a programmer burden. However, it's killingly complex when it comes to live mutating business databases that need to scale. Even many job-based analysis workloads have a lot of concurrent mutation in trying to efficiently figure out what's been done and what needs to be done next when the actual workflow is not a straight pipeline.
I do agree that 'for' loops should be replaced by broader declarations of what needs to be done, instead of these explicit instructions on how to iterate through the data.
All this comes with the assumption that much of what people consider big data isn't really big data.
In the functional world, there are all sorts of precise constructs to operate on lists (maps, folds, unfolds, scans, etc.), and many versions of these that support concurrency. That's what I imagine when you say "declarations of what needs to be done".
Even maps, folds, etc are often defined in terms of what order they process data. This is useful for early exit, and mutable state dependency between iterations. But again, this demands serialized, specifically ordered processing of elements.
Lists themselves are ordered structures. Their generalized interface does not assume it's just a bag of items easily thrown around in sub-parts, especially if the linked list nodes are to be modified.
Of course, the above applies to impure functional languages. Purely functional of course has far more freedoms, but even the Lisps and such ultimately demand sequential processing for most of their fundamental list operations.
So, it isn't just that maps are defined in terms of car/cdr, it is that the list itself is defined in terms of car/cdr. The "aha" moment is when you can see that an array of items in c can efficiently represent a tree structure where the parent/child relationships are implied, not tagged. The performance gains from this are impressive.
Every time I see a C or C++ talk discussing value types or arrays, I get to think those guys never saw other programming languages, e.g. Algol/Wirth family.
Another example of the Lady Gaga culture, instead of learning about the world of programming languages.
Look into the Scala (and clojure?) vector sometime. Literally a 32-ary tree. Makes the algorithm somewhat pleasant to look at, but are we really shocked when something using a simple array can beat it in performance?
Now, it is a perfectly fair and valid point that optimizing development time may be preferred. We are allowed contradictary valid points in situations that are trade off defined.
If you use persistent data-structures the way they are meant to be used, they can and will be faster than a simple array.
Both have valid use-cases, but they are often distinct.
Even the one you are making is... tough. Persistent data-structures are not a sure fire win against mutable versions. Again I find myself refering to the DLX algorithm. The entire point of which is that it is easier and faster to permute through different configurations of a data structure then it is to have all instances instantiated in a persistent fashion.
Does this mean that persistent data structures are without use? Of course not! Again, you may not be optimizing speed of execution. Which is perfectly valid!!
If you're using Python or C or something else to use redis as a server, it's probably just as fast as Q/KDB, and some informal benchmarking supports this:
redis-server &
redis-benchmark -t set -n 1000000 -c 1 -P 1000000
0.00% <= 1368 milliseconds
v. q)h:neg hopen`:localhost:1234
q)\t 1000000 h/"a:69"
1013
However I don't think this necessarily a good way to build your system, because if you're going to do a million of (read something write) every second, and then do a million of (read store), then you might as well write it in KDB and just make it (read something store) and save yourself 40% on your heating bill.Another benchmark is here: http://kparc.com/q4/readme.txt
A program in KDB instead of a program in Python and Redis?
Here's an implementation of a multiuser group chat in KDB:
https://github.com/srpeck/kchat/
and here's one in Python and Redis:
http://programeveryday.com/post/create-a-simple-chat-room-wi...
I would love to see a derivate of it with proper async support and more useful error messages.
The ultra terse syntax, will have lead to an extreme learning curve. But if you replaced it with more typical function names, I doubt it would have made history different/
The software ecosystem evoloved so rapidly around APL that it didn't have time to really show off anything it was good at. All the problems that it solved were solved 100 times over by others, and they gained the market share to drive the state of the art forward. Sure Kdb might be great, but there is a reason why Oracle makes a few more orders of magnitude more money then they do, even the claim that they can replicate Hadoop and other "Big Data" solutions is nonense as even a modest Hadoop consulting firm does more buisness then the 10m that Kx brings in.
The failure is one I have seen a lot. It doesn't matter what you do, if you are living on an island as a professional of any kind your contributions will be shortly forgotten and replaced by those who had the ability to share their ideas with the larger population.
it is exactly exercise in comedy. Hive and Reddis are free. They do for regular persons and tasks what only $100K/CPU systems of yesterday were doing. It is like comparing high-professional $100K film cameras of yesterday with today's consumer 5-10M cameras. A $100K thing which i would never had access to vs. real cheap/free gigantic expansion of my capabilities. The choice is obvious to me.
I'm not sure I agree, but the assertion of the article is that much of today's pain is a refusal to acknowledge the progress of the past.
For myself, there is a lot to this. Many of my peers and the younger crowd are convinced that the problems they are solving have never been considered in the past. To the point that some of my coworkers are proud that they are "weak" in algorithms.
Why did they move a whole o/s platform but not from kdb?
KDB is a fast column oriented database with an APL-like API. Hive is a SQL translation layer for MapReduce/Tez and Redis is a key-value store with rich data structures.
And so many companies do. Others continue working with crap like Hive and Reddis where they spend more time & money on supporting the software than getting analysis done.
Fortunately we have nice counter-examples with groups such as FoundationDB doing Hadoop scale with strong consistency. Building on the right, proven principles can get one ahead.
64-bit processors and more than 4 gigs of RAM are relatively new things. Mmap is useless for datasets larger than the disk you can attach to the CPU mmapping it or the size of its virtual address space. In 1991 it would be OK to call a couple billion entries "big data" (although we were yet to invent the term) but not anymore.
Having said that, the heat dissipation of instantiating objects from their on-disk representation is probably a relevant part of the refrigeration costs of any data center. We just have unimaginable computing power available and, sometimes, we do things the easy way, not the most efficient one.
Licensing and hardware will probably be the biggest costs here.
It won't, but 512GB is nowhere near "big data". In fact, most datasets we see should not be called that. My personal rule of thumb is that nothing that can fit in less than a dozen machines can be called "big data". The whole idea begun when we started manipulating datasets that were so huge that moving the data from where it was stored to where it would be processed was no longer practical.
("Computer science worked better historically in part because humorless totalitarian nincompoopery hadn’t been invented yet. People were more concerned with solving actual problems than paying attention to idiots who feel a need to police productive people’s language for feminist ideological correctness.
You may now go fuck yourself with a carrot scraper in whatever gender-free orifice you have available. Use a for loop while you’re at it. ")
Because these systems are limited by IO latency, this way the threads respond immediatly when data is available. You dont have the 'stepping' on each other problem unless they are actually running. Further, threads on the same CPU socket, even share resources like memory access, even they can step on each other. You shouldnt be worried about threads and cpu matching if your computation takes 10ms, but data access is 150ms.