Tuple Spaces – Good Ideas Don't Always Win (2011)
software-carpentry.org
software-carpentry.org
Aside from the fairness issues that other commenters mention, tuple-spaces work at an awkward level of abstraction. You can easily implement many other distributed concurrency models with them - semaphores, message queues, producer/consumer channels, broadcasts, even MVCC and transactions - but oftentimes, in the application it's more natural to just use the more specific abstractions. For example, you could implement a message queue by putting tuples of [type tag, sequence number, ...data] in and taking out the next message of a given type - but usually you'd want guarantees that there are no name collisions among type tags, and that sequence numbers are never skipped, and a mechanism to resend missing messages if for whatever reason they aren't produced correctly. At that point you'd rather just use golang-style channels or a real message queue library rather than roll that on top of the tuple-space.
There are relatively few problem domains I can think of that map directly to a tuple-space, without building some other concurrency abstraction on top of it. Dependency graphs and workflows, perhaps, but there are other libraries to handle that specific problem which also handle things like tracing, debugging, and error correction.
JavaSpaces, TSpaces emphasized the over-the-internet aspects of it, which didn't matter to us for our project. We ended up finding and using something much simpler called LighTS. With all the nice support for concurrency in Java these days though, it wouldn't be much work to put one together from scratch.
It would simply put Matlab code and parameters in a tuple, a worker would pick it up, compute, and put the results back. Used it to distribute the heavy function evaluation in a genetic optimization.
It was very easy and trouble free..
Edit: mixed up the implementations, I used TSpaces by IBM http://www.almaden.ibm.com/cs/TSpaces/Version3/ClientProgrGu...
https://news.ycombinator.com/item?id=15327211
Jon Steinhart: "Had he done some real design work and looked at what others were doing he might have realized that at its core, X was a distributed database system in which operations on some of the databases have visual side-effects. I forget the exact number, but X includes around 20 different databases: atoms, properties, contexts, selections, keymaps, etc. each with their own set of API calls. As a result, the X API is wide and shallow like the Mac, and full of interesting race conditions to boot. The whole thing could have been done with less than a dozen API calls."
To that end, one of the weirder and cooler re-implementations of NeWS was Cogent's PIX for transputers. It was basically a NeWS-like multiprocessing PostScript interpreter for Transputers, with Linda "tuple spaces" as an interprocess communication primitive:
http://ieeexplore.ieee.org/document/301904/
The Cogent Research XTM is a desktop parallel computer based on the INMOS T800 transputer. Designed to expand from two to several hundred processors, the XTM provides a transparent distributed computing environment both within a single workstation and among a collection of workstations. Using Linda tuple spaces as the basis for interprocess communication and synchronization, a Unix-compatible, server-based OS was constructed. A graphic user interface is provided by an interactive PostScript window server called PIX. All processors see the same set of system services, and within protection limits, programs capable of using many processors can spread out over a network of workstations and resource servers, acquiring the services of unused processors.
https://en.wikipedia.org/wiki/Transputer
Contrast this approach with Erlang, there's still a lot of tuples, but you have to (somehow) know where to send them at a (sometimes high) human cost to developers, but usually low runtime cost.
I don't see how this tuple space concept can work possibly replace robust supervised process management. Also, how is tuple space different than message passing? Somehow a process needs to know what values it is supposed to consume. Something has to manage that, and before you know it, you are passing messages through tuple space and have, essentially, re-invented the wheel.
> possibly replace robust supervised process management.
No more than message passing replaces supervised process management. A supervisor might use message passing to do its work, or it might use a tuple space, or something else again. It is still a separate thing.
> how is tuple space different than message passing
Message passing is an example of a shared nothing coordination scheme. For process A to know about some information, process B needs to send a message containing a copy of that information. This is useful, but not always practical. So few systems, even that use message passing, use only this simple form of message passing (more on that below).
> Somehow a process needs to know what values it is supposed to consume.
Of course. Is is no different regardless of the communication mechanism. Processes need to know what data they need to know. If anything, this is more difficult for message passing: where the sender needs to know what the receiver might need.
> Something has to manage that
Not necessarily. The processes can just go and read the data from memory. You would be right if you think that sounds like a terrible idea. But unfortunately it was, and still is, more than common. To avoid catastrophe it requires careful use of data locks, mutexes and dances around race conditions and deadlocks.
So you are right in thinking that managed approaches are superior. Message passing is one such approach. Tuple spaces are another.
> before you know it, you are passing messages through tuple space and have, essentially, re-invented the wheel
Partially true, but the opposite is more common.
Effectively tuple spaces use a database (of tuple data) as a coordination system, along with one blocking primitive (wait for data).
You can, of course, use the data in the database as messages, create some kind of message brokering process, and reinvent message passing. But the opposite is also true. In a complicated system co-ordinated with message passing, you often end up with processes whose job it is purely to hold and be queried for data by multiple consumers/modifiers. Effectively this is coordinating via data. Or re-implementing a cut down tuple space via message passing. When you start to design a 'querying' protocol to allow different processes to access and wait for data in a consistent way, regardless of what that data is, you are very close.
This isn't unique to just message passing and tuple spaces. A sufficiently complex (and well designed) program performing parallel computation with threads and semaphores will usually have its own ad hoc reimplementation of both message passing and tuple spaces. Whatever parallel computing formalism you use (and there are others besides these three), you are essentially solving the same problem. It is no surprise that they will begin to resemble each other.
Your point seems to assume that message passing is a default. It is certainly the thing that seems to have won the battle for mindshare. But the thesis of the article is that basing a system around a tuple space is a better foundation. And by better I mean that it is easier to code at scale.
I am not totally convinced. But I have to say it kind of chimes with my experience. Very large message passing systems become unwieldy, And I have ended up coding something that looks like a tuple space to tame the maze of producers and consumers.
The interesting thoughts from the paper as far as I can see were: 1) Tuple spaces are programming language or architecture or program independent, and vastly different programs can communicate with each other 2) You don't communicate directly to other agents by address, you write to a topic and read from a topic, which is a form of decoupling producers and consumers 3) The "block when nothing to read in this topic" idea, which makes programming coordination SO easy. I guess it's a bit like unix pipes.
If tuple spaces don't seem that interesting and novel, it's probably because of the benefit of hindsight and that a lot of these ideas are so subsumed into the tools of today. I can't definitively make the claim that Linda is the cause of this, but I suspect it had some effect. I think the original author also had a lot of wacky ideas around "cyberspace" and all that, but that's another deal and I don't think it's why people find the Linda paper interesting now. The closest useful descendants of Linda to be seem to be modern Pub / Sub systems or coordination databases like RabbitMQ, Kafka, Zookeeper.
In this case, the mechanism seems too general -- if I want to read all the outstadning tuples with 2 as the second elememt, that's possible, but seems rather hard. If it's really about consuming tuples within a named channel, I want to know more about the expected or desired properties of distribution -- how is it determined which agent gets to consume a tuples when multiple request a matching tuple simultaneously, which tuple is consumed when many tuples exist that match. What guarantees are needed to ensure progress, what guarantees are hard to provide, what guarantees are useful, but not required etc. Is this the most basic abstraction for distribution, upon which other useful abstractions can be built -- or are there other underlying basic abstractions that are needed, are there useful abstractions which cannot be built upon this, etc.
How is this any different than an SQL query with 2 in the second column? Tuple spaces seem like a RDBMS with a more restricted query model.
Article is from 1998 and mentions (among other things) Tuple Spaces:
https://www.wired.com/1998/08/jini/
I wonder what ultimately decided Jini's fate.
LuaTS — A Reactive Event-Driven Tuple Space https://pdfs.semanticscholar.org/91cb/8c359920682fda35abd9c2...
https://redis.io/ (uses Lua for scripting)
https://en.wikipedia.org/wiki/Comparison_of_triplestores
Comet: An Active Key Value Store
https://vanish.cs.washington.edu/pubs/osdi2010comet.pdf
https://vanish.cs.washington.edu/pubs/osdi2010comet_presenta...
ZeroMQ http://zeromq.org/
I'd love to hear to distributed languages with first class support for tuple space type operations (not Erlang).
https://www.youtube.com/watch?v=DauJ51CTIq8
Or "Intercellular Transport in the Movable Feast Machine":
https://www.youtube.com/watch?v=6YucCpYCWpY
And especially "Robust-first Computing: Distributed City Generation":
https://www.youtube.com/watch?v=XkSXERxucPc
http://www.cs.unm.edu/~ackley/papers/paper_tsmall1_11_24.pdf
I remember playing with it, and wondering why it wasn't more popular.
Sounds like it might be overly generalized, with the developer having to implement actual mechanisms ("simple patterns") on top and the compiler/runtime having to figure out efficient implementations by presumably sophisticated analysis, all the time hoping that the two align.
I seriously considered it as a way to share monitoring events between various systems that might be interested in consuming them: logging, billing, alerting, etc.
With broken links. Nice. Try http://signallake.com/innovation/p80-gelernter.pdf
https://en.wikipedia.org/wiki/Linda_(coordination_language)#...
Now mostly just use RPCs and queues directly, depending on whether I need something now, or later.