Java 7 Fork-Join Calamity
coopsoft.com
coopsoft.com
If anything, your comment is more of an ad hominem -- rather, the positive version of the same, an argument from authority -- with the implication that F/J is probably good since Doug Lea wrote it.
> The complexity criticism sounds a bit like "I don't understand this code! It must be bad!"
All complexity criticisms sound this way. Some are true and some are not, and we must attempt to separate which is which. But dismissing the criticism because of how it sounds without actually investigating the claims yourself is not good. "This sounds like X and is therefore dumb" is not a valid refutation.
It seems like you are trying to undermine some minor points in the article and claim that the whole thing is bunk as a result. What about addressing the major points instead. Can F/J scale across thousands of cores? Is work-stealing task-management actually valid in Java? Is the performance actually good? Is monitoring and stall detection available? Would a better approach use the Unsafe extension less? Etc. Otherwise your comment will not add much except for misleading weak-minded people into thinking the article is baseless.
Moving past that, there are some serious misunderstandings displayed in the article that make me nervous about accepting the rest of it at face value. For example:
> Only the Operating System (O/S) or a pseudo O/S can manage Tasks (whether they’re called tasks or processes is all the same thing.) The best an application framework can do is manage Threads.*
WTF? Tasks and processes are not the same thing in this context (yes, they are interchangeable when talking about a kernel scheduler). In the context of programming models, tasks are just reasonably independent units of work, usually thought of as within the same overall application. For examples of application frameworks that get along just fine managing Tasks, see TBB, OpenMP 3.0, Cilk and variants, Microsoft TPL, etc.
> The F/J Framework structure precludes scaling. The entanglement of client call/server processing, the spare threads necessary to help with the join() waiting problem (above), as well as the work stealing code (threads need to serially search for work among all deques) only works well on a small number of processors. This design can never scale to hundreds or thousands of processors. The overhead would kill the benefits of parallelization.
Work-stealing is designed to be more scalable than a centralized task queue by reducing contention. The idea is that, statistically, you have to serially search a very small number of neighbor queues before finding some work to steal. There are scenarios where a (logically) centralized queue wins, which are well described in the literature, but equating work-stealing with non-scalability betrays a lack of understanding of how these things actually.work.
This paragraph also echoes a theme throughout the article, which is that F/J does not fit into the author's mental model that any parallelism framework should be "enterprise-grade" and should work for a huge farm of machines. I argue that it's just fine for solutions to exist, and indeed to be present in the JDK, which prioritize simplicity and speed over logging, fault tolerance, and monitoring, and which remove the overhead of a strongly separated client-server model. Furthermore, is there evidence that the enterprise features couldn't be layered on top of the proposed framework? That would seem to be the best of both worlds.
Apart from these factual concerns, the huge degree of repetition (honestly, this article could've been 1/4 the length) and the complete lack of evidence for any of the performance and scalability claims (the table shown should not convince anyone) are red flags for me. The fact that the author happens to be highly commercially motivated to spread FUD is icing on the cake, but isn't necessary to figure out that these arguments should be taken with a huge grain of salt.
That said, there are probably some good suggestions in here for how to improve F/J. I think a lot of them are better suited for a larger piece of "parallelism middleware", like what the author sells, than to a language's standard library.
I hadn't realized that. I had assumed he was referring to others' work in this article. It's very bad form to compare your work to that of others without identifying that the work is indeed yours.
Here is a C# LINQ example:
using System.Linq;
files.AsParallel().ForAll(file => ProcessXML(file));
To achieve something similar in Java 7, one could start on perhaps with improving the java.util.Collection interface, or appropriate abstract children, by introducing the asParallel method that the Collection implementations could implement. One obvious problem I see is the lack of lambda support in Java 7, so one would have to pass an anonymous class into forXXX(..) instead which is a bit more clunky.The fact that the author is associated with a fork/join competitor and yet never acknowledges it also inflates my skepticism by a long way.