https://doc.rust-lang.org/beta/nomicon/concurrency.html
https://doc.rust-lang.org/book/second-edition/ch16-01-thread...
I don't think fixing C(++) can give what Rust can give, because Rust has a clean start with these strong guarantees built-in while for C(++) it would always be an addon. Defaults are powerful.
Parallelism is complicated though, and easy parallelism pretty much requires a functional style. Things like Haskell's Accelerate library (https://www.stackage.org/package/accelerate) seem the ideal way forward to me.
Apples to oranges comparison of course, since Rayon isn't [planned to become?] part of the Rust standard library.
[1] http://www.bfilipek.com/2017/08/cpp17-details-parallel.html#...
Basically, Rayon makes data parallelism really easy in a way that few if any other languages do. I’d love to have an equivalent in Python or Node, but it’s just not possible to achieve such a thing in most languages—even if you ignore the thread safety aspect.
Parallelism hasn’t seen a great deal of use until it’s urgently needed, because it’s hard to get right in most environments, and you normally need to substantially refactor code to make it happen. My hope is that with the likes of Rayon, parallelism can be a much more natural thing that people that care even a little about performance will just do, because it’s so easy to do.
This is the main reason why, initially, Windows Store APIs were all async.
Microsoft learned that when developers can choose between both models, by default most chose synch models.
And in .NET https://docs.microsoft.com/en-us/dotnet/api/system.threading...
And in Java https://blog.oio.de/2016/01/22/parallel-stream-processing-in...
Rayon is notable because Rust is a native-compiled, no-GC language.
Syntax-wise, there's OpenMP which can turn a for-loop into a parallelized for-loop (independently scheduled iterations) with just some syntactic sugar on top of the loop.
OpenMP has support for at least C++ and Fortran, and is not hard to use.
I wonder how Rayon compares to OpenMP.
Go's implementation of these things is nice, neat, and all included out of the box though which is nice.
In general you can make parallelism easy, or efficient, rarely both, at least not in a way that can solve problems generally.
edit: I should add Go does also come with a data-race detection tool which can be very useful. Not sure any other language includes that out of the box!
Still not quite as tightly integrated as Go though.
> Not sure any other language includes that out of the box!
Rust's compiler does it by default, at compile time. ;-) Go's race detector is never wrong, but it may omit things. On the other hand, Rust's compiler is also neither wrong nor does it omit things, except under one circumstance: someone, somewhere, wrote `unsafe` code and committed a bug inside that block.
ThreadSanitizer is also a thing: https://clang.llvm.org/docs/ThreadSanitizer.html
We already have these:
* Rendering, medium precision mathematics: GPU
* Low precision mathematics: TPU
* Software Defined Networking: Microsoft is deploying FPGAs, AWS has its own hardware
We could have:
* Databases: Projections, hashing, sorting in hardware.
* Dynamic runtimes: Hardware implemented memory models, HW assisted GC, code caches and user-level interrupts for the JITs. Here is the J extension RISC-V working group: [1]
etc.
Also, why not have the usual hot paths in Node.js|Spring Framework|Django directly etched into hardware? HW http header parsing surely could bring benefit to them all.
----
Of course language and programmers will have to adapt, but in a lot of cases the runtimes will take care of it automatically.
[1] https://groups.google.com/a/groups.riscv.org/forum/#!msg/hw-...
For example, WebAssembly specifically exposes SIMD primitives, which means that it may be necessary to work backwards from those SIMD primitives to make use of a true vector machine.
I think many people simply underestimate the cost of adopting a new programming model.
> Also, why not have the usual hot paths in Node.js|Spring Framework|Django directly etched into hardware? HW http header parsing surely could bring benefit to them all.
Well, in all the listed cases here, the CPU is not the bottleneck on throughput. As far as I can tell problem with HTTP is not that headers take too long to parse, it's that memory is still too slow, and context switches cost us precious time. The problem with Node is not that the hardware doesn't adequately model the semantics, it's that dynamic, weak typing makes it hard for any system (software or hardware) to understand what type things are.
Update: The J extension seems interesting, and I've read some research (not thoroughly) recently showing considerable power and time savings from hardware GC primitives. I'm excited to see what goes on in that committee.
I am too. So far, this post on the general RISC-V mailing list and a few videos online talk about it. I'd love to have some other sources of information on their progress.
Also, I've heard that the RISC-V foundation is actively seeking collaboration for Java. There is some work on having RISC-V backends in HotSpot and JikesRVM, but so far it is limited to interpretation IIRC. The fact that Oracle is not jumping at it and pouring hundred of millions into it is beyond me.
Well, JikesRVM is a proper JIT. Palmer Dabbelt from SiFive has worked on a HotSpot port before (for a different platform). I'm currently working on a V8 port. The availability of platform software and language environments is obviously of paramount importance, since it'll shape the remaining first impressions of the architecture.
> The fact that Oracle is not jumping at it and pouring hundred of millions into it is beyond me.
Well, it's a lot of work, and they have their SPARC investment to continue.
Hashing, maybe sorting, blitting and some math could be performed at the on-module DRAM controller level even, without data crossing over the slow DDR4 bus or mangling the CPU caches.
I'd love, in fact, to explore such an architecture in a simulator. What would happen to CPU performance if, say, hashes could be computed without reading the data, memory be cleared without zeroes hitting the bus or some SIMD operations be conducted on the memory.
Edit: clarify the processing could be done on the module side of the memory bus.
In some specific cases, such as when the Linux kernel maps memory into your process, this is exactly what happens. When you write the page, it faults and clears it on demand; but I don't think there would be a considerable benefit to doing this at a finer granularity.
This article mostly mentions a UC Berkeley effort: https://en.wikipedia.org/wiki/Computational_RAM
This page mentions others: http://www.ai.mit.edu/projects/aries/course/notes/pim.html
https://research.google.com/pubs/pub46518.html The Case for Learned Index Structures - Research at Google
People have been warning us about this for about 10 years now, but I still don't see those 64-core CPUs I was promised anywhere.
If we had the amount of parallelism we were told we were going to get, we could give every app its own core. OSes could even consider disabling context switches altogether for the majority of apps. Instead, we're left complaining about Electron apps like it matters.
That said, I'm not sure what stagnation you refer to. There's a reason languages like Rust, Elixir/Erlang and Go are getting popular. My PHP app could handle hundreds of concurrent connections on a single machine, my Elixir app handles hundreds of thousands. Yet, processors didn't get 1000x faster (and Elixir isn't even a particularly fast language). This is the opposite of stagnation, it's progress.
Threadripper and EPYC exist now though. With 32 and 64 logical cores respectively.
The people telling us we had to go change our code to use parallel processing fast predicted a significantly faster increase in amount-of-cores on commodity hardware. Instead, CPUs stopped getting faster and hardly gained more parallellism.
I'm still waiting for the massively parallel NUMA machine in a chip, but there are many manufacturing problems keeping those away.
In CPU terms, a 1080 is 40 cores, each with a 64 way vector unit.
The people promising that were crackpots and no one really called them out on it, so that meme got repeated everywhere despite being wrong. Processor vendors can't release a new processor that runs existing apps slower because no one would buy it (not counting monopolistic tactics). And since many existing apps are single-threaded, that means new processors have to at least maintain the same single-threaded performance which means keeping brainiac cores which means you can only afford 6-8 of them. (And arguably there are 64 weak cores in your CPU; they're just in the IGP and you have to program them with OpenCL.)
Go/Rust/Elixir are not so much progress IMO as undoing the negative progress of writing large-scale software in scripting languages.
In case you dont know, Golang goroutines are a marvel of parallelism. They are coroutines which are dispatched on a few OS threads. So you can use 100% of a multi-cores CPU and yet, spawn, say, 10K of those light threads without worrying about context switches PLUS have them all run concurrently. I've found that golang is one of those rare language, like Lisp, that actually change the way you think about programming. Makes you feel really more powerful.
If you dont know that language, I suggest to run the following and watch your cpu activity and memory (or any metric):
import "time"
func main() {
for i := 0; i < 10000; i++ {
go func() {
for {
time.Sleep(time.Second)
}
}()
}
}Once you execute truly on multiple CPU cores (by increasing GOMAXPROCS), you'll be having the same kind of race conditions in Go as in any other imperative language (inb4 Rust Evangelism Strike Force saying "except Rust").
GOMAXPROCS defaults to number of cores.
Prove it by replacing Sleep in your example with some number crunching, and show how it scales with the number of cores in your CPU.
Running the following code.
https://play.golang.org/p/k_rRxNAyb0i
I can assure you that go runs across all processors by default.
https://docs.google.com/document/d/1At2Ls5_fhJQ59kDK2DFVhFu3...
Wrong. GOMAXPROCS defaults to the number of logical CPUs, IIRC since version 1.5. For example, I have four cores with 2 threads each, so goroutines will be executed on up to 8 threads unless I set GOMAXPROCS to something else or the application explicitly changes it using the runtime package.
And sure, you'll really have the same problems, but IMO channels and goroutines minimize the friction of implementing thread safe programs using CSP. GP seems a bit optimistic, agreed, but I think that there is at least some substance to the idea that go makes it easier to correctly utilize multiple cores.
Additionally, combining them with the power of channels makes quick work of many tasks. Channels of course can be implemented in C++ too, but having the compiler take care of it for you with additional tools such as the race detector is very handy.
For a large set of problems, they are very nice to work with.
Then there is the ongoing work to add async/await patterns into C++20.
Notice how you can always have that answer when it comes to programming languages: you can do the same in. The point is that, the way it's made in go, is awfully handy.
go func () { }
and Task.Run(() => { })
With the benefit that on the latter example, the runtime allows me to customise how scheduling is done.These aspects are key to goroutines.
It seems like both languages are equally capable here, with C++ having more power and foot guns when required as usual.
C++ with PPL on Windows, would be
task<T> handle = create_task([] { /* ... */ });
And with standard C++ future<T> handle = std::async(std::launch::async, [] { /* ... */ });
In both cases, with C++20 it will be possible to co_await handle, which you can already play with on clang and VC++.This is probably exactly the same behavior as goroutines.
Goroutines and .net tasks allow a nice mix of CPU bound and IO bound code. While a goroutine/task is waiting for IO or something else to complete (timer in this example), the runtime will immediately use the hardware thread for some other task, without OS involved.
I'm unsure if the intent of not including a map function was due to this, however with a for loop such as
for x := range ch {
slice = append(slice, x)
}
, it is immediately obvious there are allocations happening in the background.The fact that map access, and slice access is not thread safe means there is no trickery going on in the background. The fact that it is not threadsafe means I don't need to worry about a lock if I only write to a map when it's created. Sure the compiler could take care of this - they have the race detector after all, but the compile speed is one of the design goals of go. I really like being able to compile in less than 1 second.
`sync.Map` however calls a function which implies there is more going on in the background.
If you follow the go mantra of share memory by communicate, don't communicate by sharing memory - handling data becomes a whole lot easier. It does allow you to share memory in case you do need the extra speed.
Like most things, all designs are a matter of trade offs. Sacrifice one thing for another. There are other languages that provide the functionality you desire - but I understand the frustration when one thing is 90% what you want.
Go obviously has room for improvement, and perhaps a native threadsafe map is one of those areas.
That it can't do something effectively does not mean that it shouldn't or didn't mean to. (And users can't really fix it, because hey, no generics. Sigh.)
I'm pretty sure I disagree with that largely because Go is low-level enough so that the GC is easily handled.
And users can't really fix it, because hey, no generics. Sigh.
Users can fix it by writing languages on top of Go (particularly dynamic ones).
What's really hard is to break down a problem into parallelizable chunks, figure out as much independent work as possible to reduce the touchpoints, and coordinate all those tasks such that they keep the CPU as busy as possible and as a whole finish as early as possible.
Beyond this "parallel breakdown design", it's the little touchpoints with shared data structures and synchronization that creates the difficulty of implementation, and I haven't seen any language or system that does magic there.
Common mistake number 1. Goroutines are running in parallel AND concurrently - they are coroutines (common mistake number 2 is to think they are only coroutines). I suggest not to underestimate that, it's the big deal. Additionnaly, it's important to note that goroutines yields on sleeps (any kind of sleeps / waits, like disk reads, network requests, channel writes/reads, etc), and while it may sound like a detail, it's a wonder of cpu control. Due to that, there is also a rare elegance to the way Go solve sharing data via channels.
What's really hard is to break down a problem into parallelizable chunks, figure out as much independent work as possible to reduce the touchpoints, and coordinate all those tasks such that they keep the CPU as busy as possible and as a whole finish as early as possible.
That's exactly what Go is a wonder for due to the combination of parallelism, concurrency, yield-on-sleep and channels.
If Go's m:n is blowing your mind, you should check out Erlang. Spawning a million "processes" is not a big deal.
Not sure how rare, .NET does that as well. Below is your sample translated to C#. It requires C# 7.1 because async main, but the rest of the stuff is available for many years, since 2010.
static async Task routine()
{
await Task.Delay(TimeSpan.FromSeconds(1));
}
static Task Main()
{
return Task.WhenAll(Enumerable.Range(0, 1000).Select(i => routine()));
}I don’t have hands-on experience with golang but based on what I know about it yes, they are.
> there seems to be a call to yield but its unclear what it really does
You mean “await”?
It waits for the completion of whatever is on the right side of “await”. If the result is already available, it just continues. If the result is not yet available (e.g. Task.Delay creates a task that will complete in some moment in the future), the control goes away from the async method to the scheduler. The scheduler can then run some other task on the same OS thread. When the result of that operation becomes available (in the sample code, when 1 second delay passes), the scheduler resumes execution of that async method, on the statement after the await.
> if they yield on any sleeps
No, not any sleep. You can call Thread.Sleep() which will put the whole OS thread to sleep. It’s up to programmers to avoid calling blocking APIs from their async methods, i.e. use await Task.Delay() instead of Thread.Sleep(), await stream.ReadAsync() instead of stream.Read(), and so on.
The problem with parallelism is that C-like language don't fit well, only functional languages do. If you want to use multi-threading you have to forget about state and only work with input/output paradigms. For OSes it might mean a deep re-design, but I don't really know.
A possible design would be a small but very fast CPU that only takes care or scheduling and task control, and another chip with many cores that deal with payloads and user software.
AMD had some kind of hybrid chip that planned to do both graphics and task, but it was thrown away.
Going parallel would require to change both hardware and software, and by software I mean stateless.
Logic programming languages like Prolog and Mercury are also much more amenable to parallelization than C-like languages.
In fact, different Prolog clauses could in principle be executed in parallel without changing the declarative meaning of the program, at least as long as you stay in the so-called pure subset of the language which imposes certain restrictions on the code.
.. eh? What are you referring to here? Mainstream CPUs don't have error correction either, unless you're talking about ECC on the higher-end ones.
https://en.wikipedia.org/wiki/Machine-check_exception
Realise that if you only protected RAM with ECC then you're leaving a lot of data vulnerable in caches and registers, so those need parity bits too (as well as lots of other error checks on CPU operations). And anyway, CPU errors are common due to bad power supplies and overclockers. And you don't want to add a whole lot of design effort to create a marginally cheaper-to-produce version of the CPU which doesn't do any error checking.
I've seen a lot of MCEs on non-ECC CPUs :(
IANAEE
There has. GPU programming is exactly that. CPU-heavy tasks (games, bitcoin mining, machine learning) have already migrated.
https://en.wikipedia.org/wiki/Occam_(programming_language)
Even though claims of multi-tasking etc persist, the truth is good parallel programmers are a rare thing.
Many ordinary programmers already get into Hot-Water when they use two threads and access data where a Semaphore might be needed.
In addition many algorithms are sequential, so parallelizing them is tricky or gives you no true reward due to cross-thread communication. Add to that the OO-Software Structures that subtle encourage using sequential programming.
I think the future in parallel-programming is actually hiding the parallel programming completely - accept the fact that most humans are not made for it, allow for experts to unlock the ability to override that behavior- let compilers go as far as they can and live with the results.
It will suffer the same fate as functional programming. Really useful, but never dominant, due to limitations in the applying humans.
Now, with helpful systems like no-side-effect functional languages and reactive stream frameworks, a lot of gory detail can be abstracted away. I think this has recently lead to more parallel-by-default software development.
You still have to decide upon the unit of work that is going to be sent to a different thread/core/processor/NUMA node/whatever. The different units of work that are distributed should not share state; one really doesn't want to be sharing a lot of state between different processors, because synchronizing the processors memory caches in NUMA is a extremely slow.
I guess it is really hard to break up both the program and data and decide upon the optimal granularity of the work units, it is not something that can be easily done behind the scenes - human intervention is still required.
The right kind of language looks at serial program formulations and based on flow-analysis automatically identifies parallelizable fragments that are large enough to benefit from multicore, then schedules these fragments e.g. by using work-stealing in a system of green threads, i.e., mapping green threads to OS cores as efficiently as possible. Something like that.
In a good parallel language there need to be many immutable constructs by default, exception handling is tricky, and ordinary flow control needs to be compatible by default with parallel evaluation. The languages I've seen such as Parasail are not yet production ready.
Making the programmer control parallelism can be okay, like in Go and Ada, but in the end it should be automatic.
Edit: The problem is also that finding a neat way of solving the problem academically does not readily translate into an efficient implementation, so much that I wonder whether green threads are actually worth it over OS threads. In most languages/VMs they aren't but Go seems to be an exception.
CUDA. OpenCL. Vulkan Compute Shaders. DirectCompute. C++ AMP. AMD's ROCm. Intel's SPMD. Khronos SYCL.
One of the best examples was StarLisp for the Connection Machine.
the really* cool language from Hillis and Steele was cmlisp, but I don't know how far they got, they never released anything.
There has. It's called a GPU. Things like OpenCL and CUDA are the new languages.
It's partly what made the PS3 so hard to write for, the SPUs only have 256kb of directly addressable memory, everything else is DMA'd. That said when you had your code+data fitting in 256kb it screamed everywhere else as well since you fit in L1+L2 cache neatly.
With the current fragmented and sometimes proprietary forest of programming platforms, an application needs to be quite specialized to warrant investment in GPU compute outside the original niche of graphics acceleration. There are other giant problems too, after you get over rewriting your application for numerous different platforms - atrocious quality of GPU drivers causing OS crashes for users, lack of any common way to debug GPU code, the colourful quality of compilers, the wildly different performance characteristics of different platforms necessitating per-platform algorithm changes, etc...
Consider what a minority of applications bother to even put in the work to exploit large amounts of CPU parallelism, which is vastly easier. There is after all >10x parallelism available on a typical PC CPU, after you count cores, threads and SIMD lanes.
There will, but it takes time. The world of scalar hardware in the 1970's was no less fragmented. Honestly most of the incompatibilities in the SIMD world at this point are bugs and not fundamental problems. The vector world has settled on a broad architecture at this point for most things.
On the server side, when we deploy a Node.js web service on AWS we start one instance per logical core, for 4-64 processes all independently serving connections.
It seems the process has become the new thread, the smallest unit you should design for. So today's workloads actually make pretty good use of all those cores. Unless you're doing high performance computing and need to squeeze every last drop of performance, processes are a straightforward way to parallelize.
I think we should, really, start thinking about such things. Maybe prefixing instructions with the execution unit that should handle them (and overflow back to the first one in a circle if we have more EU's in software than the actual hardware provides), separating dependencies within code flow in a more explicit way and, at the same time, not bothering with creating threads.
[0]: https://en.wikipedia.org/wiki/Chapel_(programming_language)
[1]: https://en.wikipedia.org/wiki/Fortress_(programming_language...
[2]: https://en.wikipedia.org/wiki/List_of_concurrent_and_paralle...
Scaling upwards, your opinion on that changes when a single engineer’s service is running on 10k machines.
At the other end of the spectrum, if you’re developing high-performance applications for small systems (desktop, laptops, mobile), Your workload isn’t going to look like tens of thousands of concurrent independent requests, So the approach of getting parallelism by deploying multiple copies of the application no longer works
I often experienced that this backfired. Single machines are still constrained in their power and while it's easy to spin up additional VMs in the cloud, scaling a program properly to run on dozens of machines takes a lot of work. It can be faster to develop a program that is really efficient and can solve the problem on one machine than to develop faster only to then spend the time scaling it to a large fleet of servers.
used to yield orders of magnitude more performance, but optimizers have evolved and today there's not much difference.
If this were true, you'd expect see a lot more native python and the like.From reading the open literature and advertisements by foundry companies, I think you could make a 6502 equivalent processor with Indium Phosphide with 64kb of static RAM that clocks at 30 GHz. With a more refined process you might push 90 GHz and a much more complex processor.
Yes, InP is more expensive than Silicon but part of that is the low volume that InP parts are made in. Advances in Silicon are getting much more expensive, and one InP microprocessor could do the work of ten Silicon-based cores so you can save on die area without the "race to the bottom" in size.
The main issue with high clocks is fast access to memory, probably you would need an optical interface to off-chip RAM, also I don't know what the InP equivalent of DRAM is. (Something like Optane?)
2. The problem today in CPUs is not really clock speed but much more the memory access latency, optane is much slower* than DRAM and has much lower endurance.
*Even though silicon HKMG transistors use high k gate dielectrics now, they still use a silicon dioxide interfacial layer.
there has and it's called labview, although by hardware you may have meant processors. labview has many quirks, but it surprisingly gets many things right, even in futuristic ways. when i move back to text-based languages it's always a jolt primarily due to the serial nature of them, even those that support asynchronous computation. it's really hard to recalibrate to having to assign things to temporary variables and the like. and the lower dimensions of a text file compared to a higher dimensional canvas is something that sticks out as a limiting factor in supporting parallel by default.
> I find the apparent stagnation extremely depressing.
agreed.
Obviously this problem domain is easily parallelized but it’s nice to see parallelism be the refs to standard when possible and reasonable to do so.
If your request spins in a for-loop doing lots of work without function calls, other goroutines on the same thread won't get a chance to run, and you'll be limited to GOMAXPROCS simultaneous requests. In practice this never really happens though.
While IDE tooling can still be improved, the parallel debugging tools in .NET and Java eco-systems are already quite good.
On VS, I can have at any given moment a graphical snapshot on how all threads and tasks are interacting with each other, or just execute some of the threads.
It doesn't solve everything, but it makes it easier than a typical gdb session.
https://www.realworldtech.com/forum/?threadid=146066&curpost...
Prominent examples being Go and Perl 6.
Perl 6 especially, given how audacious the project is. There are performance issues with it currently though, from what I hear they are working to fix it soon.
There are things which are slower, but since it is a higher level language it may be easier to try multiple algorithms one of which may be significantly faster. It also has many useful features included, which can be optimized in ways that aren't recommended for user code. (writing the algorithm in NQP) There is also a code specilizer (spesh) and a JIT.
Basically for many things it can be fast enough. Also if you profile your code and find something that is egregiously slow you should report it. Many times such things get optimized quickly.
For example, for web programming, you can throw several machines (or processes) of your serial program to have it run in parallel for all practical purpose.
Of course this makes some other things a bit more difficult.
The unit tests are what convinced me this will be the next big thing. Beautifully clear syntax, succinct tests, and most importantly: Parallel out of the box. Hundreds of unit tests run instantaneously. Ruby TDD setups run tests that changed with maybe 1-2s lag... Elixir runs all the tests so fast that, at the beginning, I wasn't sure the tests were running.
[1] Plus a couple of background processes, like breathing, that execute in parallel.
You could look at society as a whole and see the zeitgeist as a singular "train of thought", but you'd probably still recognize humans as individual agents. I think we have a bias towards thinking of ourselves as the ultimate individuals, neither recognizing the processes within us (like those of our cells) or the processes beyond us (like those of a group of people, animals, plants etc.) as having a similar nature. This is probably a genetically advantageous trait.
I disagree. Conscious thinking is a crucial process that's inherently single-threaded, even though it runs on highly parallel hardware (the billions of neurons).
However, I must admit that my point doesn't necessarily contradict pjc's point, and I don't really agree with digi_owl's claim that parallelism is a "wild goose chase".
I don't even know what to make of that. What do you mean by "single-threaded" if at the same time you recognize that it "runs on highly parallel hardware"? If you actually mean that our consciousness emerges from a purely sequential process that happens to go on in a highly parallel system, no, that's clearly wrong. Experience, the fundamental basis of consciousness, actuates many parts of the brain at the same time. They process this information largely independently in different ways, and sometimes those processes result in a clear "train of thought" but most of the time they do not. You can not reason about the inherent nature of our consciousness in terms of trains of thought if you recognize any subjectivity to our experience that exists without reasoning or language. That's a matter of definition, of course, and without agreeing on a precise definition it's probably no use talking about what is inherent about it.
If that's not what you mean, CPU execution models are probably not a very helpful metaphor to explain your idea. The clearly defined layer of abstraction that separates a fully pipelined CPU design built with simultaneously operating logic gates from "single threaded" programs being executed on it doesn't exist in brains. I guess that's what bugs me most about this type of discussion on HN. It seems developers are very fond of taking their (admittedly versatile) hammers and hammer away at anything they can think of, for better or for worse.
I'm not talking about consciousness (as in qualia or subjective experience), but about conscious thinking as in train of thought or intentionally thinking about something. For me, this process itself feels very sequential.
That's roughly what Intel tried with Itanium. I don't know if whatever barriers they hit are still barriers today.
Good SE answer here: https://softwareengineering.stackexchange.com/questions/2793... especially the ones focusing on cache misses.
That sounds a lot like "parallel-by-default C++" kind of language + hardware system to exploit it"
Just swapping "compiler" for language. They didn't succeed, but they did try
Edit: helping me understand where I'm off might be more helpful than a downvote. Does swapping "compiler" for "language" not respresent what Intel was trying to do?