Why has CPU frequency ceased to grow? (2014)
software.intel.com
software.intel.com
One thing that's changed as transistors have gotten smaller is that leakage has gotten to be more of a problem. You used to just worry about active switching power but now you have to balance using higher voltage and lower thresholds to switch your transistors quickly with the leakage power that that will generate.
And finally velocity saturation is more of a problem on shorter channels making current go up more linearly with the gate voltage than quadratically.
Here are some free relevant courses. You might have to go back and take the pre-reqs.
https://ocw.mit.edu/courses/electrical-engineering-and-compu...
https://ocw.mit.edu/courses/electrical-engineering-and-compu...
I tinker with electronics and make some remote controlled robots for fun (internet controlled, live video with multi user input, sort of crowd controlled). I am now trying to self teach myself about kalman filters and control theory and want to build more autonomous robots.
But any info on getting into robotics for a day job would be nice.
* Okay, I know of some people, but their design is different.
But the clock setup, hold times also gets shorter and shorter when the clock frequency goes up. The clock signal will have jitter. The end result is that less and less time of the clock edge is usable to sample the signal into the register.
And this in turn put a strain on how well balanced the logic between the registers are. To allow all signals traverse the logic paths through the gates and stabilize in time to be sampled.
To add to the complexity, as we move down the geometries, the difference in performance of different transistors becomes relatively larger. One rason for this is that oxide layers consist of (in average) fewer and fewer molecules. When the layer was made up of 100 molecules, 101 or 102 didn't really make much of a difference. But when the average is 4 molecule one more or less will have a huge impact on the performance.
So controlling variance (clock tree balance, jitter in clock generatiom, imbalances between paths and variance in chip production) becomes ever more problematic and important.
In case you dont know, Golang goroutines are a marvel of parallelism. They are coroutines which are dispatched on a few OS threads. So you can use 100% of a multi-cores CPU and yet, spawn, say, 10K of those light threads without worrying about context switches PLUS have them all run concurrently. I've found that golang is one of those rare language, like Lisp, that actually change the way you think about programming. Makes you feel really more powerful.
If you dont know that language, I suggest to run the following and watch your cpu activity and memory (or any metric):
import "time"
func main() {
for i := 0; i < 10000; i++ {
go func() {
for {
time.Sleep(time.Second)
}
}()
}
}Once you execute truly on multiple CPU cores (by increasing GOMAXPROCS), you'll be having the same kind of race conditions in Go as in any other imperative language (inb4 Rust Evangelism Strike Force saying "except Rust").
GOMAXPROCS defaults to number of cores.
Prove it by replacing Sleep in your example with some number crunching, and show how it scales with the number of cores in your CPU.
Running the following code.
https://play.golang.org/p/k_rRxNAyb0i
I can assure you that go runs across all processors by default.
https://docs.google.com/document/d/1At2Ls5_fhJQ59kDK2DFVhFu3...
Wrong. GOMAXPROCS defaults to the number of logical CPUs, IIRC since version 1.5. For example, I have four cores with 2 threads each, so goroutines will be executed on up to 8 threads unless I set GOMAXPROCS to something else or the application explicitly changes it using the runtime package.
And sure, you'll really have the same problems, but IMO channels and goroutines minimize the friction of implementing thread safe programs using CSP. GP seems a bit optimistic, agreed, but I think that there is at least some substance to the idea that go makes it easier to correctly utilize multiple cores.
Additionally, combining them with the power of channels makes quick work of many tasks. Channels of course can be implemented in C++ too, but having the compiler take care of it for you with additional tools such as the race detector is very handy.
For a large set of problems, they are very nice to work with.
Then there is the ongoing work to add async/await patterns into C++20.
Notice how you can always have that answer when it comes to programming languages: you can do the same in. The point is that, the way it's made in go, is awfully handy.
go func () { }
and Task.Run(() => { })
With the benefit that on the latter example, the runtime allows me to customise how scheduling is done.These aspects are key to goroutines.
It seems like both languages are equally capable here, with C++ having more power and foot guns when required as usual.
C++ with PPL on Windows, would be
task<T> handle = create_task([] { /* ... */ });
And with standard C++ future<T> handle = std::async(std::launch::async, [] { /* ... */ });
In both cases, with C++20 it will be possible to co_await handle, which you can already play with on clang and VC++.This is probably exactly the same behavior as goroutines.
Goroutines and .net tasks allow a nice mix of CPU bound and IO bound code. While a goroutine/task is waiting for IO or something else to complete (timer in this example), the runtime will immediately use the hardware thread for some other task, without OS involved.
I'm unsure if the intent of not including a map function was due to this, however with a for loop such as
for x := range ch {
slice = append(slice, x)
}
, it is immediately obvious there are allocations happening in the background.The fact that map access, and slice access is not thread safe means there is no trickery going on in the background. The fact that it is not threadsafe means I don't need to worry about a lock if I only write to a map when it's created. Sure the compiler could take care of this - they have the race detector after all, but the compile speed is one of the design goals of go. I really like being able to compile in less than 1 second.
`sync.Map` however calls a function which implies there is more going on in the background.
If you follow the go mantra of share memory by communicate, don't communicate by sharing memory - handling data becomes a whole lot easier. It does allow you to share memory in case you do need the extra speed.
Like most things, all designs are a matter of trade offs. Sacrifice one thing for another. There are other languages that provide the functionality you desire - but I understand the frustration when one thing is 90% what you want.
Go obviously has room for improvement, and perhaps a native threadsafe map is one of those areas.
That it can't do something effectively does not mean that it shouldn't or didn't mean to. (And users can't really fix it, because hey, no generics. Sigh.)
I'm pretty sure I disagree with that largely because Go is low-level enough so that the GC is easily handled.
And users can't really fix it, because hey, no generics. Sigh.
Users can fix it by writing languages on top of Go (particularly dynamic ones).
What's really hard is to break down a problem into parallelizable chunks, figure out as much independent work as possible to reduce the touchpoints, and coordinate all those tasks such that they keep the CPU as busy as possible and as a whole finish as early as possible.
Beyond this "parallel breakdown design", it's the little touchpoints with shared data structures and synchronization that creates the difficulty of implementation, and I haven't seen any language or system that does magic there.
Common mistake number 1. Goroutines are running in parallel AND concurrently - they are coroutines (common mistake number 2 is to think they are only coroutines). I suggest not to underestimate that, it's the big deal. Additionnaly, it's important to note that goroutines yields on sleeps (any kind of sleeps / waits, like disk reads, network requests, channel writes/reads, etc), and while it may sound like a detail, it's a wonder of cpu control. Due to that, there is also a rare elegance to the way Go solve sharing data via channels.
What's really hard is to break down a problem into parallelizable chunks, figure out as much independent work as possible to reduce the touchpoints, and coordinate all those tasks such that they keep the CPU as busy as possible and as a whole finish as early as possible.
That's exactly what Go is a wonder for due to the combination of parallelism, concurrency, yield-on-sleep and channels.
If Go's m:n is blowing your mind, you should check out Erlang. Spawning a million "processes" is not a big deal.
Not sure how rare, .NET does that as well. Below is your sample translated to C#. It requires C# 7.1 because async main, but the rest of the stuff is available for many years, since 2010.
static async Task routine()
{
await Task.Delay(TimeSpan.FromSeconds(1));
}
static Task Main()
{
return Task.WhenAll(Enumerable.Range(0, 1000).Select(i => routine()));
}I don’t have hands-on experience with golang but based on what I know about it yes, they are.
> there seems to be a call to yield but its unclear what it really does
You mean “await”?
It waits for the completion of whatever is on the right side of “await”. If the result is already available, it just continues. If the result is not yet available (e.g. Task.Delay creates a task that will complete in some moment in the future), the control goes away from the async method to the scheduler. The scheduler can then run some other task on the same OS thread. When the result of that operation becomes available (in the sample code, when 1 second delay passes), the scheduler resumes execution of that async method, on the statement after the await.
> if they yield on any sleeps
No, not any sleep. You can call Thread.Sleep() which will put the whole OS thread to sleep. It’s up to programmers to avoid calling blocking APIs from their async methods, i.e. use await Task.Delay() instead of Thread.Sleep(), await stream.ReadAsync() instead of stream.Read(), and so on.
One of the best examples was StarLisp for the Connection Machine.
the really* cool language from Hillis and Steele was cmlisp, but I don't know how far they got, they never released anything.
https://doc.rust-lang.org/beta/nomicon/concurrency.html
https://doc.rust-lang.org/book/second-edition/ch16-01-thread...
I don't think fixing C(++) can give what Rust can give, because Rust has a clean start with these strong guarantees built-in while for C(++) it would always be an addon. Defaults are powerful.
Parallelism is complicated though, and easy parallelism pretty much requires a functional style. Things like Haskell's Accelerate library (https://www.stackage.org/package/accelerate) seem the ideal way forward to me.
Apples to oranges comparison of course, since Rayon isn't [planned to become?] part of the Rust standard library.
[1] http://www.bfilipek.com/2017/08/cpp17-details-parallel.html#...
Basically, Rayon makes data parallelism really easy in a way that few if any other languages do. I’d love to have an equivalent in Python or Node, but it’s just not possible to achieve such a thing in most languages—even if you ignore the thread safety aspect.
Parallelism hasn’t seen a great deal of use until it’s urgently needed, because it’s hard to get right in most environments, and you normally need to substantially refactor code to make it happen. My hope is that with the likes of Rayon, parallelism can be a much more natural thing that people that care even a little about performance will just do, because it’s so easy to do.
This is the main reason why, initially, Windows Store APIs were all async.
Microsoft learned that when developers can choose between both models, by default most chose synch models.
And in .NET https://docs.microsoft.com/en-us/dotnet/api/system.threading...
And in Java https://blog.oio.de/2016/01/22/parallel-stream-processing-in...
Rayon is notable because Rust is a native-compiled, no-GC language.
Syntax-wise, there's OpenMP which can turn a for-loop into a parallelized for-loop (independently scheduled iterations) with just some syntactic sugar on top of the loop.
OpenMP has support for at least C++ and Fortran, and is not hard to use.
I wonder how Rayon compares to OpenMP.
Go's implementation of these things is nice, neat, and all included out of the box though which is nice.
In general you can make parallelism easy, or efficient, rarely both, at least not in a way that can solve problems generally.
edit: I should add Go does also come with a data-race detection tool which can be very useful. Not sure any other language includes that out of the box!
Still not quite as tightly integrated as Go though.
> Not sure any other language includes that out of the box!
Rust's compiler does it by default, at compile time. ;-) Go's race detector is never wrong, but it may omit things. On the other hand, Rust's compiler is also neither wrong nor does it omit things, except under one circumstance: someone, somewhere, wrote `unsafe` code and committed a bug inside that block.
ThreadSanitizer is also a thing: https://clang.llvm.org/docs/ThreadSanitizer.html
The problem with parallelism is that C-like language don't fit well, only functional languages do. If you want to use multi-threading you have to forget about state and only work with input/output paradigms. For OSes it might mean a deep re-design, but I don't really know.
A possible design would be a small but very fast CPU that only takes care or scheduling and task control, and another chip with many cores that deal with payloads and user software.
AMD had some kind of hybrid chip that planned to do both graphics and task, but it was thrown away.
Going parallel would require to change both hardware and software, and by software I mean stateless.
Logic programming languages like Prolog and Mercury are also much more amenable to parallelization than C-like languages.
In fact, different Prolog clauses could in principle be executed in parallel without changing the declarative meaning of the program, at least as long as you stay in the so-called pure subset of the language which imposes certain restrictions on the code.
.. eh? What are you referring to here? Mainstream CPUs don't have error correction either, unless you're talking about ECC on the higher-end ones.
https://en.wikipedia.org/wiki/Machine-check_exception
Realise that if you only protected RAM with ECC then you're leaving a lot of data vulnerable in caches and registers, so those need parity bits too (as well as lots of other error checks on CPU operations). And anyway, CPU errors are common due to bad power supplies and overclockers. And you don't want to add a whole lot of design effort to create a marginally cheaper-to-produce version of the CPU which doesn't do any error checking.
I've seen a lot of MCEs on non-ECC CPUs :(
IANAEE
Now, with helpful systems like no-side-effect functional languages and reactive stream frameworks, a lot of gory detail can be abstracted away. I think this has recently lead to more parallel-by-default software development.
While IDE tooling can still be improved, the parallel debugging tools in .NET and Java eco-systems are already quite good.
On VS, I can have at any given moment a graphical snapshot on how all threads and tasks are interacting with each other, or just execute some of the threads.
It doesn't solve everything, but it makes it easier than a typical gdb session.
[0]: https://en.wikipedia.org/wiki/Chapel_(programming_language)
[1]: https://en.wikipedia.org/wiki/Fortress_(programming_language...
[2]: https://en.wikipedia.org/wiki/List_of_concurrent_and_paralle...
That's roughly what Intel tried with Itanium. I don't know if whatever barriers they hit are still barriers today.
Good SE answer here: https://softwareengineering.stackexchange.com/questions/2793... especially the ones focusing on cache misses.
That sounds a lot like "parallel-by-default C++" kind of language + hardware system to exploit it"
Just swapping "compiler" for language. They didn't succeed, but they did try
Edit: helping me understand where I'm off might be more helpful than a downvote. Does swapping "compiler" for "language" not respresent what Intel was trying to do?
Even though claims of multi-tasking etc persist, the truth is good parallel programmers are a rare thing.
Many ordinary programmers already get into Hot-Water when they use two threads and access data where a Semaphore might be needed.
In addition many algorithms are sequential, so parallelizing them is tricky or gives you no true reward due to cross-thread communication. Add to that the OO-Software Structures that subtle encourage using sequential programming.
I think the future in parallel-programming is actually hiding the parallel programming completely - accept the fact that most humans are not made for it, allow for experts to unlock the ability to override that behavior- let compilers go as far as they can and live with the results.
It will suffer the same fate as functional programming. Really useful, but never dominant, due to limitations in the applying humans.
You still have to decide upon the unit of work that is going to be sent to a different thread/core/processor/NUMA node/whatever. The different units of work that are distributed should not share state; one really doesn't want to be sharing a lot of state between different processors, because synchronizing the processors memory caches in NUMA is a extremely slow.
I guess it is really hard to break up both the program and data and decide upon the optimal granularity of the work units, it is not something that can be easily done behind the scenes - human intervention is still required.
Scaling upwards, your opinion on that changes when a single engineer’s service is running on 10k machines.
At the other end of the spectrum, if you’re developing high-performance applications for small systems (desktop, laptops, mobile), Your workload isn’t going to look like tens of thousands of concurrent independent requests, So the approach of getting parallelism by deploying multiple copies of the application no longer works
I often experienced that this backfired. Single machines are still constrained in their power and while it's easy to spin up additional VMs in the cloud, scaling a program properly to run on dozens of machines takes a lot of work. It can be faster to develop a program that is really efficient and can solve the problem on one machine than to develop faster only to then spend the time scaling it to a large fleet of servers.
used to yield orders of magnitude more performance, but optimizers have evolved and today there's not much difference.
If this were true, you'd expect see a lot more native python and the like.https://en.wikipedia.org/wiki/Occam_(programming_language)
[1] Plus a couple of background processes, like breathing, that execute in parallel.
You could look at society as a whole and see the zeitgeist as a singular "train of thought", but you'd probably still recognize humans as individual agents. I think we have a bias towards thinking of ourselves as the ultimate individuals, neither recognizing the processes within us (like those of our cells) or the processes beyond us (like those of a group of people, animals, plants etc.) as having a similar nature. This is probably a genetically advantageous trait.
I disagree. Conscious thinking is a crucial process that's inherently single-threaded, even though it runs on highly parallel hardware (the billions of neurons).
However, I must admit that my point doesn't necessarily contradict pjc's point, and I don't really agree with digi_owl's claim that parallelism is a "wild goose chase".
I don't even know what to make of that. What do you mean by "single-threaded" if at the same time you recognize that it "runs on highly parallel hardware"? If you actually mean that our consciousness emerges from a purely sequential process that happens to go on in a highly parallel system, no, that's clearly wrong. Experience, the fundamental basis of consciousness, actuates many parts of the brain at the same time. They process this information largely independently in different ways, and sometimes those processes result in a clear "train of thought" but most of the time they do not. You can not reason about the inherent nature of our consciousness in terms of trains of thought if you recognize any subjectivity to our experience that exists without reasoning or language. That's a matter of definition, of course, and without agreeing on a precise definition it's probably no use talking about what is inherent about it.
If that's not what you mean, CPU execution models are probably not a very helpful metaphor to explain your idea. The clearly defined layer of abstraction that separates a fully pipelined CPU design built with simultaneously operating logic gates from "single threaded" programs being executed on it doesn't exist in brains. I guess that's what bugs me most about this type of discussion on HN. It seems developers are very fond of taking their (admittedly versatile) hammers and hammer away at anything they can think of, for better or for worse.
I'm not talking about consciousness (as in qualia or subjective experience), but about conscious thinking as in train of thought or intentionally thinking about something. For me, this process itself feels very sequential.
People have been warning us about this for about 10 years now, but I still don't see those 64-core CPUs I was promised anywhere.
If we had the amount of parallelism we were told we were going to get, we could give every app its own core. OSes could even consider disabling context switches altogether for the majority of apps. Instead, we're left complaining about Electron apps like it matters.
That said, I'm not sure what stagnation you refer to. There's a reason languages like Rust, Elixir/Erlang and Go are getting popular. My PHP app could handle hundreds of concurrent connections on a single machine, my Elixir app handles hundreds of thousands. Yet, processors didn't get 1000x faster (and Elixir isn't even a particularly fast language). This is the opposite of stagnation, it's progress.
Threadripper and EPYC exist now though. With 32 and 64 logical cores respectively.
The people telling us we had to go change our code to use parallel processing fast predicted a significantly faster increase in amount-of-cores on commodity hardware. Instead, CPUs stopped getting faster and hardly gained more parallellism.
I'm still waiting for the massively parallel NUMA machine in a chip, but there are many manufacturing problems keeping those away.
In CPU terms, a 1080 is 40 cores, each with a 64 way vector unit.
The people promising that were crackpots and no one really called them out on it, so that meme got repeated everywhere despite being wrong. Processor vendors can't release a new processor that runs existing apps slower because no one would buy it (not counting monopolistic tactics). And since many existing apps are single-threaded, that means new processors have to at least maintain the same single-threaded performance which means keeping brainiac cores which means you can only afford 6-8 of them. (And arguably there are 64 weak cores in your CPU; they're just in the IGP and you have to program them with OpenCL.)
Go/Rust/Elixir are not so much progress IMO as undoing the negative progress of writing large-scale software in scripting languages.
There has. GPU programming is exactly that. CPU-heavy tasks (games, bitcoin mining, machine learning) have already migrated.
The right kind of language looks at serial program formulations and based on flow-analysis automatically identifies parallelizable fragments that are large enough to benefit from multicore, then schedules these fragments e.g. by using work-stealing in a system of green threads, i.e., mapping green threads to OS cores as efficiently as possible. Something like that.
In a good parallel language there need to be many immutable constructs by default, exception handling is tricky, and ordinary flow control needs to be compatible by default with parallel evaluation. The languages I've seen such as Parasail are not yet production ready.
Making the programmer control parallelism can be okay, like in Go and Ada, but in the end it should be automatic.
Edit: The problem is also that finding a neat way of solving the problem academically does not readily translate into an efficient implementation, so much that I wonder whether green threads are actually worth it over OS threads. In most languages/VMs they aren't but Go seems to be an exception.
We already have these:
* Rendering, medium precision mathematics: GPU
* Low precision mathematics: TPU
* Software Defined Networking: Microsoft is deploying FPGAs, AWS has its own hardware
We could have:
* Databases: Projections, hashing, sorting in hardware.
* Dynamic runtimes: Hardware implemented memory models, HW assisted GC, code caches and user-level interrupts for the JITs. Here is the J extension RISC-V working group: [1]
etc.
Also, why not have the usual hot paths in Node.js|Spring Framework|Django directly etched into hardware? HW http header parsing surely could bring benefit to them all.
----
Of course language and programmers will have to adapt, but in a lot of cases the runtimes will take care of it automatically.
[1] https://groups.google.com/a/groups.riscv.org/forum/#!msg/hw-...
For example, WebAssembly specifically exposes SIMD primitives, which means that it may be necessary to work backwards from those SIMD primitives to make use of a true vector machine.
I think many people simply underestimate the cost of adopting a new programming model.
> Also, why not have the usual hot paths in Node.js|Spring Framework|Django directly etched into hardware? HW http header parsing surely could bring benefit to them all.
Well, in all the listed cases here, the CPU is not the bottleneck on throughput. As far as I can tell problem with HTTP is not that headers take too long to parse, it's that memory is still too slow, and context switches cost us precious time. The problem with Node is not that the hardware doesn't adequately model the semantics, it's that dynamic, weak typing makes it hard for any system (software or hardware) to understand what type things are.
Update: The J extension seems interesting, and I've read some research (not thoroughly) recently showing considerable power and time savings from hardware GC primitives. I'm excited to see what goes on in that committee.
I am too. So far, this post on the general RISC-V mailing list and a few videos online talk about it. I'd love to have some other sources of information on their progress.
Also, I've heard that the RISC-V foundation is actively seeking collaboration for Java. There is some work on having RISC-V backends in HotSpot and JikesRVM, but so far it is limited to interpretation IIRC. The fact that Oracle is not jumping at it and pouring hundred of millions into it is beyond me.
Well, JikesRVM is a proper JIT. Palmer Dabbelt from SiFive has worked on a HotSpot port before (for a different platform). I'm currently working on a V8 port. The availability of platform software and language environments is obviously of paramount importance, since it'll shape the remaining first impressions of the architecture.
> The fact that Oracle is not jumping at it and pouring hundred of millions into it is beyond me.
Well, it's a lot of work, and they have their SPARC investment to continue.
Hashing, maybe sorting, blitting and some math could be performed at the on-module DRAM controller level even, without data crossing over the slow DDR4 bus or mangling the CPU caches.
I'd love, in fact, to explore such an architecture in a simulator. What would happen to CPU performance if, say, hashes could be computed without reading the data, memory be cleared without zeroes hitting the bus or some SIMD operations be conducted on the memory.
Edit: clarify the processing could be done on the module side of the memory bus.
In some specific cases, such as when the Linux kernel maps memory into your process, this is exactly what happens. When you write the page, it faults and clears it on demand; but I don't think there would be a considerable benefit to doing this at a finer granularity.
This article mostly mentions a UC Berkeley effort: https://en.wikipedia.org/wiki/Computational_RAM
This page mentions others: http://www.ai.mit.edu/projects/aries/course/notes/pim.html
https://research.google.com/pubs/pub46518.html The Case for Learned Index Structures - Research at Google
On the server side, when we deploy a Node.js web service on AWS we start one instance per logical core, for 4-64 processes all independently serving connections.
It seems the process has become the new thread, the smallest unit you should design for. So today's workloads actually make pretty good use of all those cores. Unless you're doing high performance computing and need to squeeze every last drop of performance, processes are a straightforward way to parallelize.
I think we should, really, start thinking about such things. Maybe prefixing instructions with the execution unit that should handle them (and overflow back to the first one in a circle if we have more EU's in software than the actual hardware provides), separating dependencies within code flow in a more explicit way and, at the same time, not bothering with creating threads.
From reading the open literature and advertisements by foundry companies, I think you could make a 6502 equivalent processor with Indium Phosphide with 64kb of static RAM that clocks at 30 GHz. With a more refined process you might push 90 GHz and a much more complex processor.
Yes, InP is more expensive than Silicon but part of that is the low volume that InP parts are made in. Advances in Silicon are getting much more expensive, and one InP microprocessor could do the work of ten Silicon-based cores so you can save on die area without the "race to the bottom" in size.
The main issue with high clocks is fast access to memory, probably you would need an optical interface to off-chip RAM, also I don't know what the InP equivalent of DRAM is. (Something like Optane?)
2. The problem today in CPUs is not really clock speed but much more the memory access latency, optane is much slower* than DRAM and has much lower endurance.
*Even though silicon HKMG transistors use high k gate dielectrics now, they still use a silicon dioxide interfacial layer.
Of course this makes some other things a bit more difficult.
Prominent examples being Go and Perl 6.
Perl 6 especially, given how audacious the project is. There are performance issues with it currently though, from what I hear they are working to fix it soon.
There are things which are slower, but since it is a higher level language it may be easier to try multiple algorithms one of which may be significantly faster. It also has many useful features included, which can be optimized in ways that aren't recommended for user code. (writing the algorithm in NQP) There is also a code specilizer (spesh) and a JIT.
Basically for many things it can be fast enough. Also if you profile your code and find something that is egregiously slow you should report it. Many times such things get optimized quickly.
The unit tests are what convinced me this will be the next big thing. Beautifully clear syntax, succinct tests, and most importantly: Parallel out of the box. Hundreds of unit tests run instantaneously. Ruby TDD setups run tests that changed with maybe 1-2s lag... Elixir runs all the tests so fast that, at the beginning, I wasn't sure the tests were running.
there has and it's called labview, although by hardware you may have meant processors. labview has many quirks, but it surprisingly gets many things right, even in futuristic ways. when i move back to text-based languages it's always a jolt primarily due to the serial nature of them, even those that support asynchronous computation. it's really hard to recalibrate to having to assign things to temporary variables and the like. and the lower dimensions of a text file compared to a higher dimensional canvas is something that sticks out as a limiting factor in supporting parallel by default.
> I find the apparent stagnation extremely depressing.
agreed.
For example, for web programming, you can throw several machines (or processes) of your serial program to have it run in parallel for all practical purpose.
CUDA. OpenCL. Vulkan Compute Shaders. DirectCompute. C++ AMP. AMD's ROCm. Intel's SPMD. Khronos SYCL.
Obviously this problem domain is easily parallelized but it’s nice to see parallelism be the refs to standard when possible and reasonable to do so.
If your request spins in a for-loop doing lots of work without function calls, other goroutines on the same thread won't get a chance to run, and you'll be limited to GOMAXPROCS simultaneous requests. In practice this never really happens though.
There has. It's called a GPU. Things like OpenCL and CUDA are the new languages.
It's partly what made the PS3 so hard to write for, the SPUs only have 256kb of directly addressable memory, everything else is DMA'd. That said when you had your code+data fitting in 256kb it screamed everywhere else as well since you fit in L1+L2 cache neatly.
With the current fragmented and sometimes proprietary forest of programming platforms, an application needs to be quite specialized to warrant investment in GPU compute outside the original niche of graphics acceleration. There are other giant problems too, after you get over rewriting your application for numerous different platforms - atrocious quality of GPU drivers causing OS crashes for users, lack of any common way to debug GPU code, the colourful quality of compilers, the wildly different performance characteristics of different platforms necessitating per-platform algorithm changes, etc...
Consider what a minority of applications bother to even put in the work to exploit large amounts of CPU parallelism, which is vastly easier. There is after all >10x parallelism available on a typical PC CPU, after you count cores, threads and SIMD lanes.
There will, but it takes time. The world of scalar hardware in the 1970's was no less fragmented. Honestly most of the incompatibilities in the SIMD world at this point are bugs and not fundamental problems. The vector world has settled on a broad architecture at this point for most things.
https://www.realworldtech.com/forum/?threadid=146066&curpost...
If there ever were a return to exponential scaling, we would very soon run into the Launder limit.
I hadn't heard of this before, but it sounds like we are a long way off from reaching the limit; as the article states, modern computers use millions of times more energy than what Landauer's Principle implies is the lowest possible amount.
No we wouldn't. We're around ~10,000X off and would run into thermal danger zones long before.
Factoring in the redundancy requirement we are likely only off by somewhere between 100x and 1000x. If there were ever a return to exponential technological improvement, we would run out of road after a few years.
Preface, I'm not a programmer, I'm a hardware guy.
It's all well and good to make sure your programs and future programs are able to be run in a parallel fashion but there is a big hole to that and it's the operating systems methods of handling cores and threads.
Let's use folding@home as an example. Very multithreaded. Now let's use, at first, Ryzen 1800x as the hardware we'll run it on. We have 8 physical cores. We also have two separate dies. Each die has four cores. Each die module has their own level 3 cache. As you use your system and you are also folding, even in the newest Linux kernel, data and instructions might get evicted and bounced around and take latency hits and thus performance hits. Nothing really locks the work to the cores or threads taking into account locality. You can adjust this with HTOP and set each thread of folding manually.
Beyond AMD, even Intel has similar issues still with the 8700k. Hell, in general just efficient multithreading seems like a tough compromise for OS development. "Users" want things to be smooth upon interaction, so you have preemption. Work wants to get done but it also wants to be a good citizen to the rest of the system.
Developers are going to have to learn about, and keep up to date with, much more then a fancy new language. You're going to have to learn each new CPU inside and out and how each OS treats it.
And it doesn’t make sense to apply the term ”risc” to internal CPU design anyway.
There was nothing magic in BeOS, simply multithreading that seemed novel at the time in consumer-level hardware.
Hate to say it, but there was no magic, just engineering that everyone now has.
Applications scale, not operating systems.
The problem is now in the application/algorithm level, not the OS level.
https://www.karlrupp.net/2018/02/42-years-of-microprocessor-...
That's usually a good trade-off until someone wants to make your CPU look terrible.
Interesting chart! However that looks like moving goalposts. If you allow functions to still be exponentials when the rate repeatedly diminishes, any monotonic function can be an exponential: f(x) = x is an "exponential" with diminishing growth rate log(x). That curve looks approximately logistic to me, as most technologies usually do.
Actually I believe speculative execution theoretically allows increasing linearly single-threaded performance for exponentially more multi-threaded performance -- so a transistor price Moore's law (if not power/density) should allow a continued linear single-threaded growth. The problem is that without power (inverse) scaling, this would also cost exponentially more power. It seems the slight increase in power visible in that chart could account for a fraction of increased single-threaded performance (other factors would be improved efficiency and architectural gains); expect those factors to also stall soon (which would fulfill the "logistic prophecy").
While they've never quite taken off (the extra gates decrease speed and our harder to manufacture), with the recent side-channel attacks on processor pipelines, I've been hopeful that I would see something pop up. Imagine a world where our processors run without a clock!
It turns out that the advantages are not so great as might appear, and the situation gets worse as DRAM delay dominates. Also logic designers are a conservative bunch and getting everyone to replace the industry standard design toolchain is a big ask.
He was one of my lecturers at uni, shame not much became of Armulet tho.
DARPA manufactured a THz transistor made of InP back in 2014.
Silicon isn't the only semiconductor in nature, and others are actively being researched.
Also, "when you increase the frequency you increase the power" (which is their argument) doesn't explain why they can't increase the frequencies. That was always the case even back in 1960s.
What they actually need to explain is why they can't make the silicon more power-efficient anymore; all toy-physics arguments (such approximations/linearizations work only for a very limited range of frequencies, if they do at all, anyway meaning their scaling-relations aren't universal like they're trying to portray and those coefficients they ignore aren't constant across voltage, frequency, materials, ... either; you almost never get such simple and universal answers in condensed matter physics even for much simpler problems) mentioned there could have been made 50 years ago as well, but silicon CPU frequencies did go up.
But if your switching frequency is slow, it doesn't matter if you use a superconductor for wires. It is the switching frequency that truly determines the limits for gate times, which in turn determines how fast your CPU is.
For the record, SiGe is also very promising in terms of switching speeds. There were experiments which shows near THz frequencies.
No it's not. What is this magical material with ultra low resistance? And how do you plan to reduce capacitance?
Btw, manufacturing terahertz speed transistors is very difficult. There are Mott FETs which will switch at 10 terahertz, but they're incredibly hard to manufacture and very power hungry.
Not to mention the manufacturing challenge of integrating superconductors into chips (I think InP would be the easiest candidate, and that's saying something...)
Do you have any references for FETs that switch at 10THz? I've never heard of it and I'm interested in the physics of it.
Here is a great paper on comparing to silicon using far higher clock rates.
https://arxiv.org/ftp/arxiv/papers/1704/1704.04760.pdf
Now what will be interesting is this new architecture can be used for more traditional CS functions.
I love this paper from Jeff Dean on using the TPUs instead of a CPU to replace a b-tree for example.
https://research.google.com/pubs/pub46518.html
This also solves our multi-thread issue. Basically it done in a multi-thread manner from the ground up.
We get a round peg for a round hole.
https://www.youtube.com/playlist?list=PL5Q2soXY2Zi9OhoVQBXYF...
>One could object to this and note that due to shorter clock ticks, the small steps will be executed faster, so the average speed will be greater. However, the following diagram shows that this is not the case.
Said diagram shows that the two-clock-tick step locks the pipeline, that is you can't execute the first clock tick for the next instruction if you're still running the 2nd part of the previous one. When would this be the case? Isn't entire point of pipelining to divide a function into smaller steps that can be run in parallel? If you can split "step 3" across two clock cycles, couldn't you effectively subdivide it into two steps that could run in parallel?
I suppose that eventually you run across the issue that adding additional pipelining stages increases the logic size which in turn causes it to run slower or something like that. I wish the document was a little more specific, after all it doesn't hesitate to throw the physical formulas for power dissipation in the 2nd part so clearly it's not afraid to dig into technical details.
There is a difference between splitting "step 3" across two clock cycles and splitting "step 3" into two separate steps. The underlying assumption here is that "step 3" is indivisible. E.g. say "step 3" was memory access and the latency for that is 500 picoseconds, it's not like you can just split it into two steps and make it load faster.
Don't date robots!
Multi-core design seems now to compensate for slower clock rates, but it also has its trade-offs. It makes software more complex. In case of the CISC arch Intel established it's a huge trade-off, since CISC is supposed to make its processors easier to program, as opposed to RISC. I don't think that CISC is a good choice when it comes to massive parallelism.
But, since chip design is so expensive and is considered state of the art high-tech, we'll need to deal with everything that chip makers throw at us. Or do we?
It still matters at the small end, which is why Cortex-M exists.
> Or do we?
A startup can design its own chips, but good luck getting anyone to use it.
On the contrary, I think the increased code density (reduced fetch bandwidth --- very important for multiple cores) and greater semantic information of CISC instructions is crucial for parallelism. Large operations can be broken up into individually scheduled uops inside the core, and those uops can then be parallelised, without the equivalent of fetching all those uops from memory as might occur in a classic RISC.
In fact, even modern ARMs use this uop-based "instruction splitting" in their microarchitecture.
1) most actual RISC ISA have compact modes, typically using 16 bits instructions mixed with the regular 32 bits ones. That's Thumb2, Micro MIPS, RISC-V Compact mode, and others for embedded CPUs (ARC, Andes, ...). Their code density are competitive (and sometimes better) then x86. So with practical RISC implementation the code density is not a factor in RISC vs CISC;
2) there's a big outlier if I remember correctly: ARM in 64 bits mode dropped Thumb2 support. They certainly have to know how to keep a compact mode, and they decided not to bother. So I guess the I-cache limitation is maybe not a such a problem in real life? I don't have the data but I trust ARM to take benchmarking seriously, particularly for an ISA that also target server chips.
Certainly, Moore's law is just an observation and cannot go on forever. Would it fair to say that we've simply reached the point where we can no longer "keep up" with Moore's observation because the technology is getting harder and not because we've actually reached any limit of physics?
And the main method of making an instruction faster is by splitting it, but all instructions have now already been split as much as is possible, while still having them operate correctly.
Or put it another way, let’s say phase 3 contains several important instructions that cannot be reduced to a length of less than 1.7 clock ticks. If the pipeline stalls you have other pipelines that won’t.
Or you get crazy and put in 2 copies of the slow paths of phase 3 and one takes the even ticks and the other the odd ones.
"Nice LEDs bro" "Nah that's my CPU"
Also the promise of Pony is a garbage collection that is concurrent with program execution and since I want to write low latency server code this feature sounds very appealing.
Using a proper thermal interface material in their CPUs would be a start...
When you can decrease the temps of Intel CPUs by 20°C with delidding, the heat argument seems quite constructed.
Nicely written. Seems like the intended headline was "why it's bad to overclock", though!
Or, better yet, wave pipelining...
No one in the office understands why i'm laughing....
Edit: Oh, 2014. The sweet irony.
But there are limits to my hubris - this is on intel.com, so I'm going to go with "I'm the one missing something". Is the speed of light and number of transistors you can put in that path (due to die size) just not a practical constraint? Neither is mentioned.
(source: worked on this for a chip design software company. The delay approximation was based entirely around R/L/C modelling and had no terms for the speed of light per se. If I remember rightly it was calculated in integer pico-meters; I definitely remember it emitting an error message if you had more than 2cm of wire in any one net!)
I recall reading clear back in the 1970s that IBM mainframes were trying to do a dual processor setup. This wasn't multiple cores on one die, this was separate physical boxes. And they were having trouble because they wanted them to operate in sync (in the sense of presenting one image to the OS and applications), but they were more than a foot apart, and they were operating at sub-nanosecond frequencies. For them, the speed of light was definitely a constraint. Even if they got around all the electrical stuff, the speed of light still put a limit on how "in sync" those two CPUs could be.