639 karma · joined August 6, 2015
Current side project: https://github.com/Lramseyer/vaporview
There are people working on PPA optimization and trying to shake up how things are done, just not with LLMs.
Surfer is fantastic, and the developers of Surfer are pretty great people too! It has been on my to-do list to learn Spade.
The real danger here is how over-leveraged Open AI is. No other AI player is as exposed. Their massive spending commitments are all precariously balanced on the other end by their user base, and if that evaporates, the whole thing will fall apart and that could crash the stocks of other players ...and by crash, I mean bring them down to a realistic value. But the economy is counting on this to work, which is why I believe that Open AI's strategy here really is to make the market exposed to Open AI's risks.
Show HN: My Project - A description for my vibe coded project [3 weeks]
A lot of the good stuff I see on Show HN are projects that have been worked on for a long time. While I understand that vibe coding is newer trend, I also know that vibe coded projects are less likely to stand the test of time. With this, we don't have to worry about whether a project is AI assisted or not, nor do we ban it. Instead just incentivize longer term projects. If the developer lies about how long they worked on the project, they will get reported and downvoted into oblivion.
I'm a little biased though since I work in chip design and I maintain an open source EDA project.
I agree with their take for the most part, but it's really nothing insightful or different than what people have been saying for a while now.
OK, what about those of us who aren't writing libraries?
As a personal anecdote, the amount of opportunities that have been opened up to me as a result of my open source project are worth way more than any $1 per mention or user.
I realize the irony here that Zed is fast because it's not web based, but I stand by my claim that being able to optionally display web UIs would be a really cool feature to have. It would open the door to a lot of extensions.
Give me Yelp for date spots and take a cut of the ad revenue. That way, there's at least an incentive to get people to not ghost each other long enough to actually meet up for a date. Hopefully that will do some level of incentivizing human connection.
I have a love/hate relationship with the VScode webview panels, but the message handler is not my favorite implementation in the world. I would love a way to send binary data, and get semantic token colors.
The only issue is that when you have a custom build of VScode, you have to manage a fork of VScode, and potentially pull in updates as VScode updates. How do you manage that?
The challenge of a HDL over a regular sequential programming (software) language is that a software language is programmed in time, whereas a HDL is programmed in both space and time. As one HDL theory expert once told me "Too many high level HDLs try to abstract out time, when what they really need to do is expose time."
Also, the latency with GDDR7 is pretty terrible. It uses PAM3 signaling with a cursed packet encoding scheme. At least they were nice enough to add in a static data scrambler this time around! The lack of RLL was kind of a pain in GDDR6.
Also, I should clarify that "brilliance" and "stupidity" in this theory are not raw intelligence, but the application of said intelligence.
I should also mention that I am a VScode extension developer and I'm one of the weirdos that actually takes the time to read about API updates. They are putting in a lot of effort in developing language model APIs. So it's not like they're outright blocking others from their marketplace.
The 1064nm exciter laser is pumped by an 808nm pump laser, and based on what I know about how inefficient lasers are, I can guarantee that those beams are way more powerful than the output beam! If those leak because the manufacturer cheaped out on filters, those lasers mat not visible, but they are still dangerous!
I would bet money that most of the "certification" process for technicians (aside from high voltage electrical safety, which is common to ALL EVs) is an NDA that grants access to unnecessarily proprietary diagnostic software. All this fancy talk about certification is just an excuse to charge consumers more for maintenance by adding an illusion of complexity.
This sort of crap is exactly why we need stronger Right to Repair protection. If a company [that sells consumer products] goes under, the specs, schematics, drawings, and source code, etc should go into the public domain. It can be the owners' job to parse through it all and figure out what to do about making/procuring replacement parts and finding mechanics to service the technology, but that data should at least be available.
Encouraging modern practices and enabling developers to migrate to newer development and development adjacent tools will be the huge value add with a product like this. At my company and some of the other companies I have worked for, we primarily use Synopsys for our tooling, but in reality, we use Cadence and Siemens tools occasionally. Being able to be more vendor agnostic, and tool agnostic, would be extremely useful. I noticed that you're using ventilator, but are there plans in the future to support other vendor tools?
Promoting the use of natively running apps (even if they're thin client web apps) is a huge win in my book too! VNC and VDIs are a terrible way to work. I really hate having to deal with the 40ms latency for every key and mouse event and font scaling that never works properly.
Another question I have, is about cost - I'm not a cloud billing guy, so I don't know the numbers off hand. But from my understanding, hardware development typically sees pretty high compute resource utilization, which is why I had always assumed that in-housing compute infrastructure made financial sense. Since it's on the cloud, how does it compare to on-prem computing from a cost perspective?
Yes, but the transactions are still 16 WCK half cycles (beats) just like in GDDR6. The designers opted for a narrower bus (per channel, and more of them) rather than shorter transactions. So that doesn't save any time. I didn't find anything on the WCK rates, but it looks like they're pretty similar to GDDR6 based on all of the examples I was able to find. So I'm not convinced of much of a time savings there either.
Now, latency numbers are measured in units of tCK, not WCK, and with GDDR6 those were pretty long relative to the time it took to actually send the transaction (2 tCK.) I'm not too familiar with the internals of the DRAM, but I assume that the process of loading the data into and out of the DRAM cells is a bit involved if it takes that much time. If that were to be sped up, then we could see improvements in latency, but I'm not holding my breath.
I highly doubt that. With on-die ECC and the ridiculously complicated PAM3 encoding/decoding, I would bet that latency is going to increase over GDDR6.
Obviously, the big news is PAM3 signaling, and on die ECC. These aren't all that new, as NVIDIA's GDDR6X used PAM4 signaling (at lower frequencies than traditional GDDR6) and an unnamed DRAM vendor had GDDR6 DRAMs with on die ECC, though at the cost of having annoyingly high read and write latencies. Thankfully, only the DQ (data) pins use PAM3 signaling.
I'm going to do my best to explain how this works in GDDR7, since I need to understand this for my work:
If you don't know what PAM3 is, it stands for "Pulse Amplitude Modulation, 3 levels" Traditional communication could be thought of as PAM2, since there's a level for 0 and 1. But we usually call it NRZ for "Non-return to Zero" There's a bit of nuance, since not all binary data communication is NRZ. But that's a different discussion. Now you might ask, why PAM3 and not PAM4? With 4 levels you can transfer 2 bits, and that seems much easier to work with. Well, it's because we hate ourselves, and we hate you. That's why.
For context, a GDDR6 channel uses 16 DQ (data) pins + 2 EDC (Error Detection/Correction) pins and 16 transfers for 256 data bits and 32 CRC bits transaction. GDDR6X (from my understanding) does the same thing, but since it's PAM4, it sends 2 bits per transfer with (I think) half the number of transfers.
GDDR7 on the other hand has 11 DQ pins of PAM3 signaling. Since PAM3 is 3 state {-1, 0, 1} these are defined as "symbols" or "trits" (I really hate the word "trit".) These symbols have are 3 separate encoding methodologies (yikes) that are used in this protocol: 11b7S, 3b2S, 2b1S. These essentially determine how many bits of data you can encode in a given number of symbols. 11b7S means 11 bits of data encoded in 7 symbols. 11 bits of data has 2048 unique combinations, 7 symbols have 2187 unique combinations (3^7). 3b2S is 3 bits (8 combinations) encoded in 2 symbols (9 combinations). And 2b1S is a misnomer but it's application specific to the Poison and Severity flag bits and the Severity flag takes precedent, so the unrepresented combination is invalid.
Like GDDR6, there are still 16 transfers, but the DQ bus width is reduced to 11 DQ pins, and a parity error pin. With this, we get a total of 176 symbols per transaction. When decoded, this allows us to send 256 bits of data, 2 bits for Poison/Severity flags (this is the 2b1S symbol), and 18 bits of CRC. Encoded though is a bit of a doozy: 163 symbols for the data, 1 symbol (2b1S) for the Poison/Severity flags, and 12 symbols (3b2s) for the CRC. If that's not confusing enough, The 163 data symbols don't even all use the same encoding. The first 161 symbols encode 253 bits in 23 sets of 7 symbol to 11 bit (11b7S) sets, for 23 sets of 11 bits. The remaining 3 bits a are encoded in the remaining 2 symbols as a 3b2S set.
How this is mapped to each pin is outlined in the spec. I'll take a look at that part later, as I have had enough unfriendly math for the day.
On a different note, there were some other notable changes that I noticed:
They also shrunk the CA (Command Address) bus width from 10 pins down to 5 and run it twice as fast relative to the CK and WCK clocks. Instead of 2 cycles of 10 bits, it's 4 cycles of 5 bits. The 5 bits are split into "Row" [0:2] and "Column" [3:4] bits. It was a bit weird at first glance, but it's actually kind of nice. The commands make a lot more sense now. Interestingly enough, it looks like the CABI (Command Address Bus Inversion) and CAPAR (Command Address Parity) are part of the 20 bit CA command now. Though I'm not sure how useful a CABI is if it's once every 4 cycles. Not sure how much power you're really saving.
There are 64 mode registers instead of 16. Still 12 bit wide though. Wait WTF? This (kind of) made sense in GDDR6, since the address and the data fit nicely into 16 bits (4 address, 12 data) and you could use the other 4 bits as the command ID for MRS. Now we have 12 bit wide registers, 6 bit wide addresses, and the MRS command is a double length command because it doesn't use the column bits. This is really weird. Registers 0-31 are defined by the spec, 32-47 are reserved for future use, and 48-63 are Vendor Specific.
Found this funny - There's a feature for addressing clock drift between CA and WCK clocks due to varying Voltage and Temperature. It's called the "Command Address Oscillator" I wonder why they picked that name... "There is a CAOSC associated with each channel and operates fully independent of any channel’s operating frequency or state" Someone had fun with that one!
Other Quick blurbs:
- No mention of Quad Data Rate (QDR)
- CTLE looks like it's supported from the DRAM side.
- They added an RCK (read clock) pair, and is generated on the DRAM side.
- They added a static data scrambler! You don't know how happy I am to see this!
This has always felt like a gaping security hole waiting to be explored.
Modern, high end FPGAs have a feature known as Raw SerDes, which in essence allows you to bypass a PCIe or Ethernet controller and use those lanes (yes, PCIe lanes) to your heart's desire ...provided you can design a working communication protocol. Difficult, but not impossible by any means.
So if you wanted to, you could design your own PCIe controller and give it whatever device ID, vendor ID, memory space, or capability space you want! Normally these things are not writable on a PCIe controller. But if you designed your own, you could write them to whatever you want and spoof device types, memory spaces, or driver bindings, and probably get yourself access to memory you shouldn't be touching. While I don't know how the linux kernel would handle these potentially out of spec conditions, it never sat right with me from a security standpoint.