Then I wish my c++ objects each ran in a separate (virtual) processor instance. They could have event signalling built into the language as well as the hardware.
It might force people to partition their design for more fine grained parallelism. C++ objects use function calls for interfacing with objects. Using object methods for event handling is just a ridiculous hack. Each object should have its own memory space.
I think it's very likely that this idea is well explored. CellBE, Tilera, Xeon Phi/MIC, GPUs, etc.
> Then I wish my c++ objects each ran in a separate (virtual) processor instance. They could have event signalling built into the language as well as the hardware.
You're right to identify that the challenge for this kind of design is to come up with a programming model and an associated IPC or I/O or memory tier/caching mechanism. The HPC space is a graveyard of accelerator concepts that never reached critical mass. GPUs have been the rare success. They can amortize their business across multiple industries. They're already built in to a lot of PCs, so easy to experiment with. OpenGL and later CuDA/OpenCL created an abstraction that was fast and somewhat portable (not as much in cudas case). The abstraction relieved you of the burden of having to know much about the device's internal design whole still being quite fast.
> Each object should have its own memory space.
I don't think I know what advantage this provides. Can you share more? What do you do about composition? Nested memory spaces? Sounds challenging and potentially high overhead.
GPUs rely on symmetry to simplify the hardware design. Multiple (like 64) processing elements share the same instruction decoder. They have to access adjacent registers. So they become vector processors.
CPUs devote a huge chip area to caches and instruction pipelines. GPUs took out much of that area and complexity and replaced it with raw floating point computing power. For certain applications this has proven to be a good trade off.
What I described was a similar trade off...replacing a few heavily pipelined processors with massive amounts of cache memory with smaller cheaper cores. I wonder if it might prove to be the optimal micro-architecture for certain applications and given certain languages and certain data patterns.