Intel shows 8 core 528 thread processor with silicon photonics
servethehome.com
servethehome.com
https://citeseerx.ist.psu.edu/document?repid=rep1&type=pdf&d...
(I’d put up money that the DARPA project that funded this work is from the same lineage of TLA interest that got Tera enough money to buy Cray!)
Bill Dally was working on the networking side (maybe equiv to the photonics interconnect here) and almost got it up and running at BBN (search for "monarch") http://franki66.free.fr/Principles%20and%20Practices%20of%20...
Here's some refs for the chip https://slideplayer.com/slide/7739558/ https://viterbischool.usc.edu/news/2007/03/monarch-system-on...
6 main RISC processors with 96 ALUs for the lightweight IO/compute processes
You can see the latter in your code, you can’t—on a modern superscalar—see the former. Is the proposed architecture any different?
Also, I'm not even sure it'd be too painful to program (in assembly, at least). It'd be perhaps inconvenient to derive those ops from C code, but Rust, with its explicit ownership, may have a better hand here.
The primary weakness of barrel processors is human; only a handful of people grok how to design codes that really exploit their potential. They look deceptively familiar at a code level, because normal code will run okay, but won’t perform well unless you do things that look extremely odd to someone that has only written code for CPUs. It is a weird type of architecture to design data structures and algorithms for and there isn’t a lot of literature on algorithm design for barrel processors.
I love barrel processors and have designed codes for a few different such architectures, starting with the old Tera systems, and became quite good at it. In the hands of someone that knows what they are doing I believe they can be more computationally efficient than just about any other architecture given a similar silicon budget for general purpose computing. However, the reality is that writing efficient code for a barrel processor requires carrying a much more complex model in your head than the equivalent code on a CPU; the economics favors architectures like CPUs where an average engineer can deliver adequate efficiency. At this point, I’ve given up on the idea that I’ll ever see a mainstream barrel processor, despite their strengths from a pure computational efficiency standpoint.
This was a bet that completely failed for Itanium, but maybe this time...
But it still can't really design things. It's probably about as far away from being able to design a complex CPU architecture as GPT-4 is from Eliza.
https://techmonitor.ai/technology/ai-and-automation/nvidia-a...
Google has been working on AI's to optimize code:
https://blog.research.google/2022/07/mlgo-machine-learning-f...
> But it still can't really design things.
What do you really mean by "really" in that sentence?
In any case, the claim you responeded to was not that the chips would be designed from the ground up by AI, only that AI will enable us to run code on top of chips that are even more complex than current chips.
The Google research you linked is using AI to make better decisions about when to apply optimisations. For example when do you inline a function? Which register do you spill? Again this is a very simple (conceptually) low level task. Nothing like writing a compiler for example.
> What do you really mean by "really" in that sentence?
I mean successfully create complex new designs from scratch.
This would work only if ChatGPT 6 is AGI, but then it doesn't make much sense to label it as ChatGPT.
If the actual requirement for this is AGI, then it's clearer to just say AGI instead of ChatGPT which might or might not ever become AGI.
And then, it'd require a lot of software rewriting, because we are not used to write for hundreds of threads for context switching is a very expensive operation on modern CPUs. On this CPU a context switch is fast and happens on any operation that makes the CPU wait for memory access, therefore, thinking in terms of hundreds of threads pays off. But, again, this will need some clever OS designs to hide everything under an API existing programs can recognize.
It may even be that nothing comes out of it except another lesson on how not to build a computer.
Existing programs don't need to see these things at all, they can just act like microservices, and we could load them with an image and call them as black boxes.
I would think a lot more work would be done by various accelerator cards by now.
Not too psyched to use some proprietary black box compilerthough.
From skimming Wikipedia, it looks like a big challenge is cache pollution. Is it possible that the hit to cache locality is what inhibits uptake? After all, most threads in the OS are sitting idle doing nothing, which means you’re penalized for any “hot code” that’s largely serial (ie typically you have a small number of “hot” applications unless your problem is embarrassingly parallel)
The challenge is mostly that you have to create enough fine-grained parallelism and that per thread performance is relatively low. Amdahl's law is in full effect here, a sequential part is going to bite you hard. That's why each die on this chip has two sequential-performance cores.
The graph problems this processor is designed to handle have plenty of parallelism and most of the time those threads will be waiting for (uncached) 8B DRAM accesses.
> From skimming Wikipedia, it looks like a big challenge is cache pollution
This processor has tiny caches and the programmer decides which accesses are cached. In practice, you cache the thread's stack and do all the large graph accesses un-cached, letting the barrel processor hide the latency. There are very fast scratchpads on this thing for when you do need to exploit locality.
A barrel processor would make branching more efficient than on a GPU, at the cost of throughput. The set of problems that are economically interesting and that would strongly profit from that is rather small, hence these processors remain niche. Conversely, the incentive to have problems that map well to a GPU is higher, because they are cheap and ubiquitous.
edit: s/eight/four
Nowadays, apps that run straight on baremetal servers are the exception instead of the norm. Some cloud applications favour tech stacks based on high level languages designed to run single-threaded processes on interpreters that abstract all basic data structures, and their bottleneck is IO throughput instead of CPU.
Even if this processor is not ideal for all applications, it might just be the ideal tool for cloud applications that need to handle tons of connections and stay idling while waiting for IO operations to go through.
Though CPU/GPU combinations are slowly moving in that direction anyway. Incidentally, if this sort of thing really interests you: when I was playing around with machine learning my trick to see how efficiently my code was using the GPU was really simple: first I ran a graphics benchmark that maxed out the GPU and measured power consumption. Then I did the same for just the CPU. Afterwards while running my own code I'd compare the ratio between my own code's power draw with the maximum obtained during the benchmarks, this gave a pretty good indication of whether or not I had made some large mistake and showed a nice and steady increase with every optimization step.
Imagine what kind of performance you could get out of hardware that is task specific. That's why for instance crypto mining went through a very rapid set of iterations: CPU->GPU->ASIC in a matter of a few years with an extremely brief blip of programmable hardware somewhere in there as well (FPGA based miners, approximately 2013).
Any loss of generality will result in efficiency and vice versa, the question is whether or not it is economically feasible and there are different points on that line that have resulted in marketable (and profitable) products. But there are also plenty of wrecks.
There are some of the complications related to barrel processing
Which is what "average engineer" would use.
This sort of architecture. I wouldn't be surprised if current GPU were doing something similar.
If you think about executing a shader program. You typically are running that same code over a bunch of data. You can map that to multiple threads.
https://en.wikipedia.org/wiki/Thread_block_(CUDA_programming...
https://yosefk.com/blog/simd-simt-smt-parallelism-in-nvidia-...
It does share the latency-hiding-by-parallelism design, but GPUs do that scheduling on a pretty coarse granularity (viz. warp). The barrel processors on this thing round-robin through each instruction.
GPUs are designed for dense compute: lots of predictable data accesses and control flow, high arithmetic intensity FLOPS.
In contrast, this is designed for lots of data-dependent unpredictable accesses at the 4-8B granularity with little to no FLOPS.
This part of each core is very similar to the existing GPUs.
What is different in this experimental Intel CPU and unlike in any previous GPU or CPU, is that each core, besides the GPU-like part, also includes 2 very fast threads, with out-of-order execution and a much higher clock frequency than the slow threads. Each of the 2 fast threads has its own non-shared execution units.
Separately, the 2 fast threads and the 64 slow threads are very similar with older CPUs or GPUs, but their combination into a single core with shared scratchpad memory and cache memories is novel.
It could be that eg 97% of your threads are looking things up in big hashtables (eg computing a big join for a database query) or binary-searching big arrays, rather than ‘some I/O task’
the 'plus' side here is that that condition gets handled gracefully, but yes, certainly you can end up in a situation where memory transactions per second is the bottleneck.
its likely more advtangeous to have a lot of memory controllers and ddr interfaces here than a lot of banks on the same bus. but that's a real cost and pin issue.
the mta 'solved' this by fully dissociating the memory from the cpu with a fabric
maybe you could do the same with cxl today
The protocols for HPC are so amorphous that they bubbled up into the lowest common denominator, completely software defined async global workspace
Networking is a huge part of cloud applications, and network connections take orders of magnitude longer to go through than disk access.
There are components of any cloud architecture which are dedicated exclusively to handling networking. Reverse proxies, ingress controllers, API gateways, message broker handlers, etc etc etc. Even function-as-a-service tasks heavily favour listening and reacting to network calls.
I dare say that pure horsepower servers are no longer driving demand for servers. The ability to shove as many processes and threads on a single CPU is by far the thing that cloud providers and on prem companies seek.
The PPE of the Cell was a rather weak CPU, meant for control functions, not for computational tasks.
Here the 2 fast threads are clearly meant to execute all the tasks that cannot be parallelized, so they are very fast, according to Intel they are eight time faster than the slow threads, so the 2 fast threads concentrate 20% of the processing capability of a core, with only 80% provided by the other 64 threads.
It can be assumed that the power consumption of the 2 fast threads is much higher than that of the slow threads. It is likely that the 2 fast threads consume alone about the same power as all the other 64 threads, so they will be used at full-speed only for non-parallelizable tasks.
The second big difference was that in the Cell the communication between the PPE and the many SPEs was awkward, while here it is trivial, as all the threads of a core share the cache memories and the scratchpad memory.
Getting some Cell[1] vibes from that, except in reverse I guess.
https://www.servethehome.com/marvell-thunderx3-arm-server-cp...
IIRC it was targeted at database workloads, and the pitch was the same: if the core is usually twiddling its thumbs waiting on RAM, it might as well work on another thread in the meantime.
And I guess IBM and Zen4C kinda fufill this demand, but more SMT16 cloud instances explicity targeted at these low IPC loads would be neat.
Interesting thank you!
Maybe I misread it. Maybe the optical component is for something else like die stacking so you get grids of chips on a super carrier, with optical interconnect.
Yikes, Intel. Has to be a pretty low moment as a chipmaker to have to use your competitor's fabs for something like this.
But yes, the 10nm node has been a disaster for Intel and set the company back 5/10 years on the chip making business.
Yes 10nm+++++ was a big problem, too.
Apple was also busy going independent and I think their path is to merge iOS with MacOX someday here so it makes sense to dump x86 in favor of higher control and profit margins.
Intel downsides for Apple:
1. No reliable control of schedule, specs, CPU, GPU, DPU core counts, high/low power core ratios, energy envelopes.
2. No ability to embed special Apple designed blocks (Secure Enclave, Video processing, whatever, ...)
3. Intel still hasn't moved to on-chip RAM, shared across all core-types. (As far as I know?)
4. The need to negotiate Intel chip supplies, complicated by Intel's plans for other partner's needs.
5. An inability to differentiate Mac's basic computing capabilities from every other PC that continues to use Intel.
6. Intel requiring Apple to support a second instruction architecture, and a more complex stack of software development tools.
Apple solved 1000 problems when they ditched Intel.
Apple doesnt have on chip RAM either. They do the exact same thing PC manufacturers do: use standard off the shelf DDR.
I can't seem to find any Intel chips that are packaged with unified RAM like that.
I doubt it. Apple loves doing full vertical stack as much as possible.
Their strategy for desktop and mobile processors has been different since the 90s and they only consolidated because it made sense to ditch their partners in the desktop space.
Not this heavily. They bought an entire CPU design and implementation team (PA Semi).
The PA Semi purchase and redirection of their team from PowerPC to ARM was completely different and obviously signaled they were all in on ARM, like their earlier ARM/Newton stuff did not.
I think a more parsimonious explanation is the accepted one: Intel was floundering for ages, Apple’s phone CPUs were booming, and a company which had suffered a lot due to supplier issues in the PowerPC era decided that they couldn’t afford to let another company have that much control over their product line. It wasn’t just things like the CPUs failing further behind but also the various chipset restrictions and inability to customize things. Apple puts a ton of hardware in to support things like security or various popular tasks (image & video processing, ML, etc.) and now that’s an internal conversation, and the net result is cheaper, cooler, and a unique selling point for them.
That and they are not paying for Intel's profit margins either. Apple is the quintessential vertical integration - they own their entire stack.
Under load, my M1 laptop can pull similar wattage to my old Intel MacBook Pro while staying virtually silent. Meanwhile the old Intel MacBook Pro sounds like a jet engine.
I couldn't find good data on the older mbpros, but the m1 max mbpro used 1/3 the power vs an 11th gen Intel laptop to get almost identical scores in cinebench r23.
https://www.anandtech.com/show/17024/apple-m1-max-performanc...
Intel for the last 10 years has been saying “if your CPU isn't 100c then theres performance on the table”.
They also drastically underplayed TDP compared to, say, AMD, by taking the average TDP with frequency scaling taken into consideration.
I can easily see Intel marketing to Apple that their CPUs would be fine with 10w of cooling with Intel knowing that that they wont perform as well, and Apple thinking that there will be a generational improvement on thermal efficiency.
But that was my entire point (root thread comment.)
It's not that Apple was taking existing Intel CPUs and designing bad thermal solutions around them. It's that Apple was designing hardware first, three years in advance of production; showing that hardware design and its thermal envelope to Intel; and then asking Intel to align their own mobile CPU roadmap, to produce mobile chips for Apple that would work well within said thermal envelope.
And then Intel was coming back 2.5 years later, at hardware integration time, with... basically their desktop chips but with more sleep states. No efficiency cores, no lower base-clocks, no power-draw-lowering IP cores (e.g. acceleration of video-codecs), no anything that we today would expect "a good mobile CPU" to be based around. Not even in the Atom.
Apple already knew exactly what they wanted in a mobile CPU — they built them themselves, for their phones. They likely tried to tell Intel at various points exactly what features of their iPhone SoCs they wanted Intel to "borrow" into the mobile chips they were making. But Intel just couldn't do it — at least, not at the time. (It took Intel until 2022 to put out a CPU with E-cores.)
On a 15/16" Intel MBP, the CPU alone can draw up to 100w. No Apple Silicon except an M Ultra can draw that much power.
There is no chance your M1 laptop can draw even close to it. M1 maxes out at around 10w. M1 Max maxes out at around 40w.
Intel doesn't publish anything except TDP.
Being generous and saying TDP is actually the consumption; most Intel Mac's actually shipping with "configurable power down" specced chips ranging from 23W (like the i5 5257U) to 47W (like the i7 4870HQ); (NOTE: newer chips like the i9 9980HK actually have a lower TDP at 45w)
of course TDP isn't actually a measure of power consumption, but M2 Max has a TDP of 79W which is considerably more than the "high end" Intel CPU's; at least in terms of what Intel markets.
Keep in mind that Intel might ship a 23w chip but laptop makers can choose to boost it to whatever it wants. For example, a 23w Intel chip is often boosted to 35w+ because laptop makers want to win benchmarks. In addition, Intel's TDP is quite useless because they added PL1 and PL2 boosts.
The major pains for Apple was when the thermal situation was so bad that CPUs were performing below base clock. -- at that point i7's were outperforming i9's because they were underclocking themselves due to thermal exhaustion; which feels too weird to be true.
My 2019 macbook pro 15 with the i9-9880H can maintain the stock 2.3GHz clock on all cores indefinitely, even with the iGPU active.
And you can compare both of those and Intels newer chips to Apples ARM offerings.
I think Apple really wanted to unify the Mac and iOS platforms and it would have happened regardless.
They're all using ASML lithography machines anyway, so who's feeding the wafers into the machine is kind of inconsequential.
If that was true Intel wouldn't be years behind.
Source requested, I follow this news closely and have not seen anything that AMD is using Intel's fab.
Nvidia praised their test chip and that indicates they might use Intel's fab. They have not definitively announced that either (https://www.tomshardware.com/news/nvidia-ceo-intel-test-chip...)
Amazon was using Intel Foundry for packaging, not fabrication. And Qualcomm was considering 20A (https://www.pcgamer.com/intel-announces-first-foundry-custom...) in 2021, but there was a rumor earlier this year that Qualcomm might not use it after all (https://www.notebookcheck.net/Qualcomm-reportedly-ditches-In...)
TSMC is in a unique position in the market and its integration with ASML is one of the chapters in this novel.
If you search Google with the time filter between 2001 and 2010 you'll find news on it.
There must be some reason HN users don't know very much about semiconductors. Probably principally a software audience? Probably the highest confidence:commentary_quality ratio on this site.
I anticipate that photonics will be introduced to general-purpose computing in the years to come, even if only to get a handle on the rising excess heat problems. Most notable is the 10 -> 7 nm process.
Perhaps HN user bcantrill will give us the inside scoop :-)
Maybe not a big factor, but it also was bothersome for sysadmins because much of the work we had to do was serial, single core, etc. Meaning it showed it's worst side to the group that usually signed the vendor checks.
Here is an interesting one. Intel has a 66-thread-per-core processor with 8 cores in a socket (528 threads?) The cache apparently is not well used due to the workload. This is a RISC ISA not x86.
The diagram shows 8 gray boxes (in 2 columns of 4). Each gray box is a core.
Above that is a detailed view of what each core looks like. There's a middle section labeled "Crossbar", and above that are 3 boxes labeled STP, MTP, and MTP. Below that are another 3 identical boxes. So each core has 4x MTPs and 2x STP.
What are those? The text of the same slide says MTP is "Multi-Threaded Pipelines" with "16 threads per pipeline" and STP is "Single-Threaded Pipelines" with (self-evidently) 1 thread per pipeline.
And 4 * 16 + 2 * 1 = 66, so 66 threads per core.
where 2 is number of slow threads, like current CPU,
and 64 are more like GPU threads.
This architecture could be next step in GPU integration. Hope they manage to write efficient implementation for standard math libs. It could be faster than separate CPU+GPU in one package.
there were no interrupts, just a thread waiting for someone to wake it up.
There are some other explainers out there that go into more detail, but Nvidia is also doing this with Grace Hopper (I think?), and Apple with their M-series chips.
Update: It looks like it's just DDR5 ("custom DDR5-4400 DRAM") in this case.
If you are only accessing a single 4B or 8B element of that cache line >80% of the time, as shown on the slides, you are wasting 7/8th of all memory bandwidth with irrelevant data. If you were to use that 64B memory bus to access 8x different 8B memory locations, you get a large boost to your effective bandwidth.
[1b] https://www.servethehome.com/intel-xeon-e-cores-for-next-gen...
[2]https://www.servethehome.com/google-details-tpuv4-and-its-cr...
Very interesting to see the adoption of fiber optic in chips (PIC, photonics in chip/photonic integrated circuit) from so many big players and newcomers.
The YouTube channel for those who like me are too impatient to ever read.
I'm far too impatient for YouTube videos. Gimme an article I can scan.
He's far too impatient for articles. Give him a YouTube video he can listen to while doing something else.
20 years and we haven't changed each others mind :->
Can it be further exploited by not requiring it to be 100% precise, ie. when close approximation is good enough?
What would be a good use case for a high-thread / low-core chip?
Parallelism, like Erlang systems?
Sadly it never took off.
Normal multithreading is not a solution because it is not granular enough, you can not context switch in the middle of each memory access to run a different thread while waiting for the memory access to complete. You could manually unroll loops and instead of only following one path through the graph you could follow several paths at the same time by interleaving their operations. While this will be able to hide the waiting times because you can continue working on other paths, it will make your code more complex and increase register pressure as you still have the same number of registers but are operating on multiple independent paths. Letting the hardware interleave many parallel paths will keep your code simpler and as each thread has its own registers also not increase register pressure.
Downside: Your business logic of what to do with every node needs to be written in SIMD intrinsics. But that has to be cheaper to do than asking Intel to design a new CPU for you.
But with hundreds of hardware threads I could just continue executing some other thread that already got his pointer dereferenced. SIMD is great when you can issue loads and they all get satisfied with the same cache line, but if you have to gather data, then values will arrive one after another and you can not do any dependent work until the last one arrived. I guess the whole point is that all the threads can be at slightly different execution points, while some are waiting for reads to complete others can continue working after their read completed without having to wait for an entire group of reads to complete.
This is also why hyperthreading yields a benefit for sparse/graph workloads. I was getting good results on the KNL Xeon Phi with its four hyperthreads. But you easily hit the 'outstanding loads' limit if your hyperthreads all do gather instructions every 4-8 instructions. The architecture needs to balanced around it.
If you use SIMD gather instructions to fetch all the children of a node, that is still going to be very inefficient on e.g. a Xeon. Each of those gather loads are likely to hit different cache lines, which are accessed only once and waste 7/8th of the memory bandwidth pulling them in. SIMD is also not going to help if you are memory bandwidth or latency bound, we're talking ~0.1 ops/byte here.
"14 Popular Air-Gapped Data Exfiltration Techniques Used to Steal the Data" - https://thesecmaster.com/14-popular-air-gapped-data-exfiltra...
"A Survey on Air-Gap Attacks: Fundamentals, Transport Means, Attack Scenarios and Challenges" - https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10054827/
Also none of the above are about software level security.
You'll probably not doing those anyway, even if you have all the security fixes for cpu timing stuff (as was the context) enabled.
If you only care about security, then I think OpenBSDs approach is currently the best, but it also seems like they got lucky a few times, like with Zenbleed, where they for unknown reason never really adopted the AVX to the same extend as Linux or Windows.
Security concerns are typically solved in software by them as long as they can get access to the hardware in a free (as in freedom) way.
I actually applaud them for their hard stance and am happy that we have this end of the spectrum as well as the Linux end (pragmatic, just get devices to work somehow, willing to accept some non-freedom). It's certainly not the easiest path to follow.
Your comment reads as someone who doesn’t really interact with the OpenBSD ecosystem very much.
I’m pretty sure the commenter you’re replying to was referring to the fact that hyperthreading is disable by default on OpenBSD systems out of caution: https://news.ycombinator.com/item?id=17350278
And as for their attitude toward firmware blobs, while they ideally prefer them to be free, they only require them to be redistributable; this is a less hard stance than GNU. Plenty of OpenBSD drivers require proprietary blobs to function.
Now, when they disable Zen SMT, then there might be a decent talking point.