“256 cores by 2013”?
herbsutter.com
herbsutter.com
Each Xeon Phi core has a 32K L1 data cache and a 512K L2 data cache. A Xeon Phi core can issue 16 SIMD operations/cycle and it needs 4 threads to saturate the instruction pipeline. The 60 cores in a 5110P have to share 320 GB/s of main memory bandwidth or 5.33 GB/s per core.
And all of this means that code customized for one processor (OpenCL) will run like crap on the other one. The Xeon Phi needs that 30 MB of total individual L2 cache to avoid slamming into memory bus contention while the K20 needs to operate entirely inside its L1 cache and register file to hit peak performance. Main memory fetches are less fatal to the K20 both because optimized code will be running 40+ warps to bury the latency and because of the significantly higher bandwidth.
What strikes me at this point is absolutely paucity of compelling Xeon Phi benchmarks. All we see are SGEMM, DGEMM, and a bunch of synthetic tests. They've had 6 years to get this right so why didn't they go after all the jewels in NVIDIA's many-core crown from the get-go?
Finally, languages like OpenCL and CUDA subsume SIMD, multithreading, multi-core, and cache optimization into the programming model, all but implicitly forcing the programmer into optimizing many-core performance. In contrast, Intel continues to expect programmers to use processor-specific intrinsics to hit peak performance that change with vendor and processor generation. Sure, it's easier to write serial code for a serial Intel core. But I thorughly disagree that it's easier to write many-core applications by adding processor intrinsics and a threading library to fundamentally serial code.
Edit: Worth mentioning that there has been lots of interesting things happening in Haskell to support concurrency: http://stackoverflow.com/questions/3063652/whats-the-status-...
Already in the middle 90's there where some HPC research with Lisp as the main language.
Functional languages have the ability to work similar to how SQL does querys, by abstracting how the DB engine works.
Of course we are seeing this type of language constructs come to the mainstream languages as well, because you really need to have some kind of inteligent runtime to fully explore parallelism.
It is a bit like assembly programming vs compilers. Sure clever programmers can beat most compilers when targeting simple processors, but when you level up with out-of-order executing, multiple instruction pairing, multiple level caches, NUMA, and so on, the compiler optimizer usually wins hands down.
I disagree. Real HP software have their computation kernels hand written (or generated) in assembly. As for relying on the smartness of the compiler to take into account all those factors I invite you to recall or read the story of the itanium architecture.
That seems pretty real HP to me.
If you are working in a research context, it's not really worth writing assembly-- just getting it to work is more important. If you have to buy more hardware, then just do it. Traditional HPC is not really known for being very cost-sensitive.
When you're working in a commerical context, performance starts to matter more. It's the same reason why a one-off hand-soldered electronic device isn't built to the same standards as an iPod. If you're only building one, don't waste time on polish.
Why aren't you using compiler vector operations for such case?
On the project I used to work on, we were building the infrastructure to perform real time data analysis for the data coming straight out of the accelerator.
We have written custom memory allocators, our own network stack and protocols, measured and optimized every operation in the years before the accelerator went live.
The code is massively parallel in core and cluster distributed.
No extra need for Assembly.
The problem with the "sufficiently smart compiler" argument, as always, is that the compiler may have a lot of optimizations, but it's still not an artificial intelligence. It can't tell that what you are trying to do with your set of operations is actually perform CRC, and there is hardware support for that.
Check it out at: http://svn.apache.org/viewvc/hadoop/common/trunk/hadoop-comm...
We have written custom memory allocators, our own network stack and protocols, measured and optimized every operation in the years before the accelerator went live.
I would actually argue that you don't need to do these things in most cases. SCTP and DCCP are alternatives to TCP/IP that have been in the Linux kernel for a while now. If you want to trash TCP/IP and go full custom, you can do so without writing a line of protocol code. However, a better approach is to use TCP/IP with tweaks like Fast Open (what Google uses internally.) You can then continue to buy commodity hardware.
Similarly, memory allocators have been done to death. Just use tcmalloc or jemalloc rather than rolling yet another malloc(). The exception is if you want to create something like a slab allocator where you just hand out lots of mostly identically-sized buffers, or a database-like application where you manage your own writeback to disk.
This might be a special case, as you're taking advantage of a specific use case instruction, which will fail in portable code anyway.
I have seen Assembly programmers put to shame in modern processors by C and C++ developers, just by making use of better data structures, algorithms and a special mix of compiler flags.
Anyway thanks for the follow up, very interesting.
http://cdsweb.cern.ch/record/616089
A more up to date link is available here, http://atlas-proj-hltdaqdcs-tdr.web.cern.ch/atlas-proj-hltda...
ROBins are the special purpose network cards.
Personally, I like the Cilk[0, 1] approach, and it is simple to implement in most cases.
[0] https://en.wikipedia.org/wiki/Cilk [1] https://en.wikipedia.org/wiki/Intel_Cilk_Plus
When you have lots of cores, I think your bottleneck would be cache and to a lesser degree, memory bandwidth. You can't have multiple cores manipulate the same memory address simultaneously.
Am I right?
Of course, many of us could have told them that that wasn't going to work, but they didn't ask us now did they? :-)
I'm not trying to be dismissive, I'm just really interested to hear about the tradeoffs involved.
I think the whole focus on "cores" misses lots of issues. Memory infrastructure is only one. Instruction-level parallelism (superscalar & out-of-order) execution is another - even single-core processors like ye olde Pentium can execute multiple instructions per cycle. It's very easy and tempting to look at the number of cores and use that as a rough estimate of system performance. But this approach will land you WAY off of real-world figures. It's akin to using the number of cylinders in a car's engine to determine how fast it is - sure, to the first order, cylinder count is correlated with engine output and hence car speed, but it's only a very rough correlation.
Is there solid empirical data on the threads vs cores thing?
A lot of modern app architectures are effectively 3 threads (UI thread, compute thread, and compositor thread), which seem like good candidates for SMT. Although I don't think many/any mobile architectures use SMT -- I don't know how it effect power efficiency.
If a smaller cluster of 256-core servers had better bang/buck than a bigger cluster of 16-core servers then that would still be a win for many applications (not everything CPU bound is also I/O bound, e.g. web application servers written in hilariously inefficient languages, virtualization).
Take a game like guild wars 2 and their pvp where in a screen's line of sight, you have somewhere of 300 players all using some form of animated actions with projectiles and particles and other fluff. CPU is indeed in need there, through the bottle neck could well be their utilization of the cpu cores. Games are always pushing for more, and cpu usages has not exactly slowed down.
On the web-services side, if you don't go cloudy with clouds, you still want to have a large cpu resource for scaling.
The use of old P54C core, while understandable from hardware point (design already existed, no need to validate/test) this decision has big consequences. The core can only support a single outstanding memory operation, crippling memory performance. Doubly so because it does not support out-of-order execution, meaning the core is stalled until memory requests finish.
Furthermore, the lack of L2 cache consistency lets you experiment nicely with NUMA and the combination of reprogrammable lookup tables and more physical memory than address space (34 bit vs 32 bit) lets you do nice tricks to share/communicate data between cores with zero copying needed. Unfortunately (due to time/budget constraints) Intel didn't implement any programmable expiration/invalidation/flushing of the L2 cache. This means that any application wanting to use these tricks had no solution beyond reading in a full 256KB (with the proper striding to make sure you actually cache miss on all of those 256KB!) to flush the cache. Such a flush takes approximately 500k CPU cycles, almost guaranteed killing any performance gained from zero copy tricks.
The fact that cores were 32bit is also extremely, as you quickly consume your entire address space when you start playing around with the above mentioned sort of memory shenanigans.
There were some more fairly minor issues, but these are the ones that initially pop in my head.
This is up there with The Next Web always vainly touting what they "told us" in the last article.
Say something interesting. Don't tell me how great you are. It immediately diminishes my opinion of people and companies because it makes it clear they are more interested in recognition and taking credit than in being interesting.
Iirc my gtx 670s have 1344 cores each, compared to 4 in my 3770k.
In OpenCL terms it's called a PE (Processing Element).
Also what AMD calls a "GPU core" is something smaller, than what NVIDA calls "GPU core".
Talking about specialty hardware like the Xeon Phi HPC accelerator (which most developers have never heard of, let alone programmed for), just doesn't make sense. Similarly, GPUs may be ubiquitous, but how many programmers have actually written for them? Very few.
The reality is that CPU architects have done everything they possibly can to keep down the amount of parallelism. When you get transistors, you can use them for things like L1 cache, instead of adding more execution units.
The bottom line is that it's hard to predict the future. We can read stuff from the past, like mailing list posts about how Linux is irrelevant because "we'll all be using Sparcs in a few years," but the lesson that it's hard to make predictions never really seems to sink in.
I hope that we'll see more manycore chips in the future for developers to play with. The realities of physics seem to be dragging CPU architects kicking and screaming into the multicore world. But it may not be as fast as we once thought.
It wasn't a prediction about the state of tools for programmers. The multi-threaded hardware is ubiquitous, but it's still hard to program for.
Now, this was pretty much the low end of his original prediction: 'If we stick with "just more of the same" as in Figure 2's extrapolation, we'd expect aggressive early hardware adopters to be running 16-core machines (possibly double that if they're aggressive enough to run dual-CPU workstations with two sockets), and we'd likely expect most general mainstream users to have 4-, 8- or maybe a smattering of 16-core machines (accounting for the time for new chips to be adopted in the marketplace).'
It is true that new, mainstream systems offer 4 to 8 way parallelism, and that you see up to 24 way parallelism on the high end. So, his "more of the same" prediction was perfectly correct.
We did not see the large jump that he hypothesized might happen. But if you notice, he said "But the gating factor is software that can use them effectively; specifically, the availability of scalable parallel mainstream killer applications. The only thing I can foresee that could prevent the widespread adoption of manycore mainstream systems in the next decade would be a complete failure to find and build some key parallel killer apps, ones that large numbers of people want and that work better with lots of cores."
As you say, most programmers don't develop for these big many-core monstrosities. So he was right; software that takes advantage of multi-core machines is still a gating factor. The chips are available, but because that many cores does not help most software that hasn't been specially crafted for it, they are not mainstream.
I think the one major fault is in his optimism that it's possible for most software to take good advantage of that many cores. Parallelism is hard. Correct software is more important than fast software. Heck, a lot of software these days is written in languages like JavaScript, Python, and Ruby, which are not known for their speed, but rather their ease of development. Telling people to get with the multi-core bandwagon, when really just writing large, correct software is in most cases more important that squeezing out the last ounce of performance, is not exactly productive.
As GPUs have shown us, the best way to take advantage of that extra parallelism is to write special purpose libraries for the computationally intensive parts, and then just call out to that from simpler, less parallel code. So I'm not sure that the average programmer is going to spend that much time adapting their code for the parallel world, other than using existing libraries and frameworks for taking advantage of that.
A database backed web app is a great example of this kind of parallelism. The hard parts of parallelism are generally handled by the database itself. The web app is usually written completely single threaded, but you can spin up multiple processes all talking to the same database to take advantage of your multiple cores.
I suspect that is how we'll see parallelism play out in general. Most software written single threaded, with parallelism only by running multiple processes, talking to specially tailored databases, message queues, rendering libraries, numerical computing libraries, and the like.
The Nvidia Tesla K20 has 2496 cores (I know that it is a GPU and not a CPU)
As others have pointed out, with AMD Opterons you can stuff 64 cores on a server motherboard today.