Everything You Know (about Parallel Programming) Is Wrong
splashcon.org
splashcon.org
One of the biggest challenges I see practically are that concurrency and parallelism are the hardest thing to push down under a layer of abstraction. These effects leak out all the way through your application (and even into other applications).
It strikes me that this kind of non-deterministic approach will leak everywhere. And then who can program in this way? Who among us can reason effectively about just how far from the true answer we are. If this is the only way we can really harness the many-many-core PCUs of the future then we will have the change the way we make software. The IBM approach of armies of programmers will not work. As much as I hate this approach to software it is the way most software (at least in the enterprise) is made. Will we go back to the long beards? Or will we end up in even more aggressively restrictive frameworks that work harder to prevent us from doing this wrong.
If this approach to parallelism turns out to be important I hope that the problem of reasoning about these types of program is tractible. To push well understood software development even further away from the reach of most programmers. If we all end up having to learn a bit of statistics that wouldn't be a bad thing.
If floating point math is anything to go by, most programmers won't be able to cope with this. (Then again, most languages do nothing to help programmers keep track of how large the inaccuracies get.)
The first are the cogs. These are typically either immutable data structures, or mutable components designed to be accessed only from a single thread. The second is a concurrent abstraction which sits on top of the cogs.
For example, in my project Hobo, I've got an immutable, read-only index with procedural, deterministic methods. Then sitting in front of it I've got a TThreadPoolServer which handles multithreaded access.
Single threaded: https://github.com/stucchio/Hobo/blob/master/src/org/styloot...
Library handles multithreaded part: https://github.com/stucchio/Hobo/blob/master/src/org/styloot...
Another thing that will work for data where contention is rare is software transactional memory. Most programmers are already familiar with SQL transactions, I don't see a reason they can't use it in memory as well.
This is a very good illustration of how concurrency potentially leaks out across an entire application. You can see here that your concerns for concurrency appears in the cogs. It is certainly true that you have provided a layer of abstraction but only for a certain concurrency model, and your cogs will reflect that in their implementation.
I would also note that I wouldn't claim that all concurrency problems pose these kinds of problem. Graphics pipelines being the easiest (and best abstracted) example. But in general for the day to day code we write, concurrency concerns have a habit of being very hard to tidy away.
For a really good treatment on the virtues and limitations of transactional memory there is a great conversation between Rich Hickey and Cliff Click on the clojure mailing list. I can't find the original, but here is a good cleaned up summary of the exchange.
http://blogs.azulsystems.com.sharedcopy.com/cliff/2008/05/97...
But the point I'm making is that once the API is fixed you get to program more or less normally. Look at the code I linked to - it's all standard procedural java. Some smart guys at Facebook developed the concurrency system. I did some ordinary programming subject to one additional constraint (immutability).
I imagine this is how the future will look - top programmers build APIs, and the army of programmers will build many immutable/transactional/etc subcomponents. With language support (such as what Haskell/Clojure already give you) this will be very easy to enforce - IBM's army of guys who can't handle concurrency will stick to writing pure functions.
Would appreciate it if people around here (and elsewhere) would stop acting like the word "cogs" is universally recognized, like "ram" or "dos". I have no clue what "cogs" is / are, and the search engines aren't helping to make me any wiser (or at least informed).
""" To the extent that an application remains acceptable to its end users with respect to their business needs, we aim to permit – or no longer attempt to prevent – inconsistency relating to the order in which concurrent updates and queries are received and processed """
This is something many people are already familiar with: you can giveup some correctness or timeliness, if it doesn't matter much. It's what's behind eventual consistency, yesterday's youtube architecture article, redis' nofsync settings etc.
[1] http://soft.vub.ac.be/~smarr/renaissance/ungar-kimelman-adam...
Speaking from my own experience, there's usually a way to order the results either by sorting, by accumulating them in a deterministic manner (extended fixed-point math or by ordered reduction), or by halting execution upon reaching a consistent degree of convergence.
And we already know one way to program such beasts: in a map-reduce like manner where chunks of computation are doled out to all the cores, each executed deterministically within them, but doled out in a nondeterministic order.
interview with author (about the talk): http://channel9.msdn.com/Blogs/Charles/SPLASH-2011-David-Ung...
the project homepage: http://soft.vub.ac.be/~smarr/renaissance/
It has to be true, because it makes no sense whatsoever. Web designers of the future will have a great reply to customers. "The site gives wrong answers? Well, you need to realize that asking for correct answers isn't what we do with computers. And besides that, in another universe the answers are correct. Boy, are you really happy over there!"
That was perhaps a too-long-winded snark with a point: it's not that I don't feel that we're headed in the right direction, it's that visionaries always seem to concentrate on the brain-exploding part of what they're saying instead of talking about how it all integrates together. While I think this might sell some website views, it doesn't do much for the practitioner looking for insight.
<not-snarky>It isn't clear what quantum computing has to do with this discussion.</not-snarky>
Edit: Mitigate snarkiness of final comment.
My entire point was that the way we describe future tech emphasizes the differences and exotic nature of the concepts instead of how it's all going to eventually work together anyway. Quantum computing was my second example.
Apologies if I didn't make that clear enough.
There are plenty of domains where a very fast almost-correct answer is preferable to a perfectly correct, slower one.
Take search, for an example. First off - it's very hard to determine an exact definition of the right answer. If the provided answer was good enough, it was right.
Another domain is route planning, especially with multiple stops in arbitrary order (travelling salesman). There are plenty of applications where the 99% answer in one second is vastly preferable to the 100% answer in an hour.
Ungar isn't saying that we have to give up determinism. What he's saying is that to take advantage of massively parallel systems, determinism comes at a high cost.
Clearly, we can write programs on current architectures that are deterministic (well, perceived as such, at least) and we don't have to forgo that. But there is no reason to believe it has to be the only architecture. Ungar is looking at what happens when you try to do computation in a highly networked environment with low latency (like say, a brain).
Also, if you don't see it helping for the practitioner, perhaps practitioners aren't asking good (or enough?) questions.
And no, practitioners shouldn't have to ask the question "what good it is?" Every research result begs that question.
Thinking about how the cloud model (collaborating services via RPC) maps to a single machine -- I guess that looks like a actor / message passing model where each process gets its own core or something.
Plenty of programming languages make this sort of thing straightforward -- (I'm most familiar with Go, where it is absolutely trivial).
To me, it looks like the industry is actually very well poised to handle to multi-core onslaught, without requiring some kind of fundamental rethink like nondeterminism.
First of all, even handhelds will have thousands of CPUs.
Second of all, it's not necessary to scale to this level. Rather, we will run more concurrent, independent tasks. Servers with thousands of workers will be common at home.
A server would be cheaper?