JEDEC publishes GDDR7 graphics memory standard
jedec.org
jedec.org
This more complex encoding scheme is just the next level in that process, indeed moving it closer to techniques used in RF engineering.
Increasing the symbol complexity of each channel does more than just move the bottleneck around, because it allows fewer chip to chip interconnects to carry more data.
I don't work in this regime, but as a layman I'm not convinced using full QAM for on-board chip to chip interconnects makes sense. One major advantage you natively have over the RF case is you can be easily coherent (shared clock). Throwing this away to do carrier recovery introduces a lot of complexity and potentially reduces the available bandwidth. Assuming you transmit without a carrier, can you have "baseband" QAM without a separate I and a Q signal? If you transmit an I and Q signal separately, does that not just become the same thing as two PAM-32 signals?
Did you mean higher?
- one needs 2 signals instead of one (2x total bandwidth) - requires each channel bandwidth to extend to to DC, which had many other challenges
If one modulates the signal to shift it away from DC, the “negative/mirror” frequencies also shift, which means now bandwidth has doubled.
A QAM signal still has double the bandwidth of an equivalent PAM one but pays for it by encoding two PAM signals.
Of course, Discrete Multitone Modulation puts QAM to shame for non-flat channels as it can adapt near-perfectly to such. Not likely to happen for high speed interconnects in our lifetime. I suspect photonics will happen first.
Most of the fancier schemes take advantage of the fact that traditional binary signalling has excess noise margin, ie, they were throwing away energy to start with, and they are encoding extra bits in the energy budget. But to maintain noise margins as you cram more bits in, you have to up the amount of energy per bit.
The other half is that the physical layer implementation that does the encoding and decoding consumes more energy because it's doing more computation. This also figures into the energy per bit metric if you're being honest about your comparisons (and because it is not always clear if this is included in a metric you find papers where people cherry pick numbers to make their case). This number can become quite big because the baseline of a binary tx/rx is so low compared to doing effectively a DAC/ADC and phase recovery system.
What you find is that QAM or more schemes are certainly possible, but they can consume more power than the CPU just to keep the link idle and trained. The real art is picking the implementation and developing new circuit tricks that we hadn't thought of before to wring a little more bandwidth without killing the power budget.
These complex modulations have a huge drawback though: latency. You need to pack, ramdomize, encode and modulate at the transceiver, and undo all these + equalization at the receiver. Especially feed-forward equalization is a huge latency source.
> Cable latency is about 5ns/m, so 2.6 µs is equivalent to latency of a 520m cable. This latency might become a big issue for the HPC-based applications.
PCB manufacturers start offering FR-4 with strands of single-mode fiber embedded in it?
How does the receiving technology get built? Surely at least someone will have to go there the first time, and they will have to take the long way. It will still be quite a problem to get to a system 10k light years away.
So we ship off these receivers to circumvent that limitation. Instead of travelling ourselves, we can send off our consciousness to inhabit a human-life analog to explore.
What that does to your psyche, and your body in limbo, are probably good material for a story, if it hasn't already been written.
At best, you'll throw a bunch of nanoprobes everywhere to get new entropy into the system.
The tricky thing is that hacking is usually an iterative process, and these iterations are going to be an extreme exercise in patience.
Actually, another tricky thing: how do you know that the other end is actually cooperating? If the aliens are dicks they could give you the thumbs up while having zero intention to reconstitute your consciousness. If you wanted to round-trip some brave soul as a means of verifying everything works, they could just send one of their own minds back instead, just for the fun of wreaking havoc.
No kidding! On the first try you accidentally end up causing a revolution because the targets/specimens ended up learning about the scientific method, gunpowder, and other dangerous things instead of just getting a proper advanced consciousness installed. So now all you can do is try to shape said species technological progress towards building the correct technology that you can hijack for your own purposes when ready.
"Just be patient"
I presume humans are the result of such a hack a few billion years ago.
Interstellar travel require patience, at least to get beyond the initial latency.
Besides, the serialisation process is a form of quantum measurement. Depending on how coarse-grained it is, there might be no way to take a snapshot without modifying you (maybe the measurement process turns the original brain matter into soup).
Cloning a hard drive can produce the same data, but without any networking, there's no reason for the original machine to know anything from the perspective of the new one
What if there were 100 billion of us in total?
Rip yesterday me.
Sounds like a copy, not a transfer. If you didn't physically transport the atoms, the matter, you would end up with two duplicates living at different places and time, and with different ways of thinking after the copy, as the living experiences will diverge from that moment.
This unless you exterminate the original with each copy. Also should be considered each copy may lose information, degrade (signal integrity through distance, number of travels, and so on).
The problem of synchronization is gonna be particularly nasty in this case.
And most certainly after sleep.
It's like walking into a room and forgetting why you walked in there.
On the other hand, there exists two spots in the universe that are so far separated from each other that they observers in both spots will never be able to see, affect, or pass information between each other, simply because they're too far away and the speed of light is too slow. That's silly.
Time is relative to your frame of reference. If you travel to Proxima Centauri at 99.99% c, it will take you 22 days from your point of reference sitting in the spaceship, which is quite acceptable. On Earth, 4.24 years have passed. So, your family and friends grow old quite fast when you do interstellar travel without them; hence, it's better to take them with you.
From what i see reading https://en.wikipedia.org/wiki/Multi-channel_memory_architect... different channels could, in theory, be used "autonomously of each other".
Memory channels are independent, however generally all cores use all channels. There are two common designs. Intel has 4 dies, 2 memory channels per core, so 2 channels are closer/lower latency than the other 6. AMD has multiple chiplets, but a single memory controller with 12 channels. So All cores have the same latency to all channels.
Generally Intel has lower latencies to 25% of the channels, but AMD has more throughput (bandwidth or random IOPs).
One thing that surprised me is that for maximum throughput you want at cache misses queued to the memory controller, at least twice the number of memory channels. These days missing in L1/L2/L3 is often approximately half the total memory latency. So on an Intel Xeon at least 16 misses (per socket), AMD at least 24 misses (per socket.).
So on Intel you could tune things (and the NUMA support helps) to prefer the local channels. Most OSs help, and C calls like numa_alloc_local() allows local control.
For memory intensive codes I have found the best scaling when there's 2 cores per channel. Of course most codes are pretty friendly.
Oh but the speed of the signal does depend quite a lot on the transmission medium. In Cat-6 signals travel 2/3c. Can't find a quick reference for on-die or motherboard kinds of interconnects. If you had optical interconnects traveling through vacuum in a silicon chip, that's a full 50% faster (as in lower travel time for one bit over a distance) than Most copper ethernet.
And this (actually, phase velocity) is what makes refraction a thing.
I knew that we could slow down light to subsonic speeds, but TIL we can put it at a complete standstill! Amazing.
I highly doubt that. With on-die ECC and the ridiculously complicated PAM3 encoding/decoding, I would bet that latency is going to increase over GDDR6.
I assume by "the ridiculously complicated PAM3 encoding/decoding" you are referring to section 2.9.3,
"The total burst transfer payload per channel is encoded using 23 x 11b7S and 1 x 3b2S for the data, 6 x 3b2S for the CRC and 1 x 2b1S for the SEV/PSN, it adds up to the 176 PAM3 symbols that can be allocated for a 16 burst over 11 data lines."
that does seem complicated using 3 (11b7s, 3b2s, 2b1s) different modulation schemes in 1 burst.
Yes, but the transactions are still 16 WCK half cycles (beats) just like in GDDR6. The designers opted for a narrower bus (per channel, and more of them) rather than shorter transactions. So that doesn't save any time. I didn't find anything on the WCK rates, but it looks like they're pretty similar to GDDR6 based on all of the examples I was able to find. So I'm not convinced of much of a time savings there either.
Now, latency numbers are measured in units of tCK, not WCK, and with GDDR6 those were pretty long relative to the time it took to actually send the transaction (2 tCK.) I'm not too familiar with the internals of the DRAM, but I assume that the process of loading the data into and out of the DRAM cells is a bit involved if it takes that much time. If that were to be sped up, then we could see improvements in latency, but I'm not holding my breath.
https://www.techpowerup.com/gpu-specs/radeon-r9-fury-x.c2677
> AMD has paired 4 GB HBM memory with the Radeon R9 FURY X, which are connected using a 4096-bit memory interface.
Who cares about consumer inference? The money is in training because that's where utterly insane amounts of compute capacity are needed, and as long as no competition comes even close to CUDA, NVIDIA has a cash cow to milk.
As for potential competition, Apple doesn't sell their silicon to anyone else, and AMD lacks the available manufacture capacity (keep in mind, they also supply the game console market), the driver/tooling quality and most importantly developer trust/ecosystem quality. Everyone is using CUDA.
Obviously, the big news is PAM3 signaling, and on die ECC. These aren't all that new, as NVIDIA's GDDR6X used PAM4 signaling (at lower frequencies than traditional GDDR6) and an unnamed DRAM vendor had GDDR6 DRAMs with on die ECC, though at the cost of having annoyingly high read and write latencies. Thankfully, only the DQ (data) pins use PAM3 signaling.
I'm going to do my best to explain how this works in GDDR7, since I need to understand this for my work:
If you don't know what PAM3 is, it stands for "Pulse Amplitude Modulation, 3 levels" Traditional communication could be thought of as PAM2, since there's a level for 0 and 1. But we usually call it NRZ for "Non-return to Zero" There's a bit of nuance, since not all binary data communication is NRZ. But that's a different discussion. Now you might ask, why PAM3 and not PAM4? With 4 levels you can transfer 2 bits, and that seems much easier to work with. Well, it's because we hate ourselves, and we hate you. That's why.
For context, a GDDR6 channel uses 16 DQ (data) pins + 2 EDC (Error Detection/Correction) pins and 16 transfers for 256 data bits and 32 CRC bits transaction. GDDR6X (from my understanding) does the same thing, but since it's PAM4, it sends 2 bits per transfer with (I think) half the number of transfers.
GDDR7 on the other hand has 11 DQ pins of PAM3 signaling. Since PAM3 is 3 state {-1, 0, 1} these are defined as "symbols" or "trits" (I really hate the word "trit".) These symbols have are 3 separate encoding methodologies (yikes) that are used in this protocol: 11b7S, 3b2S, 2b1S. These essentially determine how many bits of data you can encode in a given number of symbols. 11b7S means 11 bits of data encoded in 7 symbols. 11 bits of data has 2048 unique combinations, 7 symbols have 2187 unique combinations (3^7). 3b2S is 3 bits (8 combinations) encoded in 2 symbols (9 combinations). And 2b1S is a misnomer but it's application specific to the Poison and Severity flag bits and the Severity flag takes precedent, so the unrepresented combination is invalid.
Like GDDR6, there are still 16 transfers, but the DQ bus width is reduced to 11 DQ pins, and a parity error pin. With this, we get a total of 176 symbols per transaction. When decoded, this allows us to send 256 bits of data, 2 bits for Poison/Severity flags (this is the 2b1S symbol), and 18 bits of CRC. Encoded though is a bit of a doozy: 163 symbols for the data, 1 symbol (2b1S) for the Poison/Severity flags, and 12 symbols (3b2s) for the CRC. If that's not confusing enough, The 163 data symbols don't even all use the same encoding. The first 161 symbols encode 253 bits in 23 sets of 7 symbol to 11 bit (11b7S) sets, for 23 sets of 11 bits. The remaining 3 bits a are encoded in the remaining 2 symbols as a 3b2S set.
How this is mapped to each pin is outlined in the spec. I'll take a look at that part later, as I have had enough unfriendly math for the day.
On a different note, there were some other notable changes that I noticed:
They also shrunk the CA (Command Address) bus width from 10 pins down to 5 and run it twice as fast relative to the CK and WCK clocks. Instead of 2 cycles of 10 bits, it's 4 cycles of 5 bits. The 5 bits are split into "Row" [0:2] and "Column" [3:4] bits. It was a bit weird at first glance, but it's actually kind of nice. The commands make a lot more sense now. Interestingly enough, it looks like the CABI (Command Address Bus Inversion) and CAPAR (Command Address Parity) are part of the 20 bit CA command now. Though I'm not sure how useful a CABI is if it's once every 4 cycles. Not sure how much power you're really saving.
There are 64 mode registers instead of 16. Still 12 bit wide though. Wait WTF? This (kind of) made sense in GDDR6, since the address and the data fit nicely into 16 bits (4 address, 12 data) and you could use the other 4 bits as the command ID for MRS. Now we have 12 bit wide registers, 6 bit wide addresses, and the MRS command is a double length command because it doesn't use the column bits. This is really weird. Registers 0-31 are defined by the spec, 32-47 are reserved for future use, and 48-63 are Vendor Specific.
Found this funny - There's a feature for addressing clock drift between CA and WCK clocks due to varying Voltage and Temperature. It's called the "Command Address Oscillator" I wonder why they picked that name... "There is a CAOSC associated with each channel and operates fully independent of any channel’s operating frequency or state" Someone had fun with that one!
Other Quick blurbs:
- No mention of Quad Data Rate (QDR)
- CTLE looks like it's supported from the DRAM side.
- They added an RCK (read clock) pair, and is generated on the DRAM side.
- They added a static data scrambler! You don't know how happy I am to see this!
> For context, a GDDR6 channel uses 16 DQ (data) pins + 2 EDC (Error Detection/Correction) pins and 16 transfers for 256 data bits and 32 CRC bits transaction. GDDR6X (from my understanding) does the same thing, but since it's PAM4, it sends 2 bits per transfer with (I think) half the number of transfers.
GDDR6X did not transfer 2 bits at a time despite being PAM4, because they disallowed transitions between the lowest voltage and the highest voltage.
Because of that each group of 4 symbols has 139 easily-stackable sequences, so they went with 7b4S.
https://research.nvidia.com/publication/2022-04_saving-pam4-...
I couldn't tell you how each transaction is laid out.
I would think NVIDIA in particular and other chip makers/integrators like apple make up their own standards now. It also seems less relevant because memory is rarely interchangeable anymore
Sad that the vast majority of x86-64 laptops and desktops have the same bus width of decades ago, while the core counts are ever increasing.
Apple doesn't have any 1024-bit offerings. The Ultras are 2x 512-bit which is different from a true 1024-bit since it's a NUMA configuration (two different memory controllers).
> Sad that the vast majority of x86-64 laptops and desktops have the same bus width of decades ago, while the core counts are ever increasing.
The consistent trend is CPU performance on consumer workloads is still much more sensitive to memory latency than it is bandwidth. This is probably also why those 512-bit M1/2 Max's don't even bother to let the CPU make use of it, they top out at "just" 200GB/s memory bandwidth even though there's 400GB/s available to the SoC as a whole.
Being able to pair a lot of DRAM with the GPU of an M* Max/Ultra is definitely a currently unique perk, but the bandwidth numbers available to those GPUs are not actually all that special. Desktop GPUs passed 400GB/s mark way back in 2016, and are currently pushing 1TB/s in modern flagship offerings.
So the solution is to pick (or create) a separate, notionally independent body staffed and supported by representatives of all the relevant stakeholders, and have them write "standards" that everyone agrees to adhere to. The body doesn't invent the technology, that happens at the individual chip companies. They then present their proposals to JEDEC[1] and everyone argues and agrees on what will go into GDDR19 or whatnot. And JEDEC then publishes the standards for all to see.
[1] Or whoever, JEDEC does DRAM, but there's a USB Consortium, Bluetooth SIG, WiFi is under IEEE, etc...
Ah, a FOSSy dev insistent that his bits are gibi- and mebi- and kibi-bytes.