Sam Zeloof is known from his YouTube channel[1] where he shows how to build microprocessors in detail.
I hope that this step is the first of many to build an efficient processor with fresh ideas and innovation.
Sam Zeloof is known from his YouTube channel[1] where he shows how to build microprocessors in detail.
I hope that this step is the first of many to build an efficient processor with fresh ideas and innovation.
Long story short if I had money that would be relevant in this context I would invest really hard into Tenstorrent.
Investing time to listen to the guy seems to be not worse.
1a. https://www.youtube.com/watch?v=Nb2tebYAaOA
1b. https://www.youtube.com/watch?v=G4hL5Om4IJ4
Some former AMD employees started it.
From Jim himself:
> IC: A few people consider you 'The Father of Zen', do you think you’d scribe to that position? Or should that go to somebody else?
> JK: Perhaps one of the uncles. There were a lot of really great people on Zen. There was a methodology team that was worldwide, the SoC team was partly in Austin and partly in India, the floating-point cache was done in Colorado, the core execution front end was in Austin, the Arm front end was in Sunnyvale, and we had good technical leaders. I was in daily communication for a while with Suzanne Plummer and Steve Hale, who kind of built the front end of the Zen core, and the Colorado team. It was really good people. Mike Clark's a great architect, so we had a lot of fun, and success. Success has a lot of authors - failure has one. So that was a success. Then some teams stepped up - we moved Excavator to the Boston team, where they took over finishing the design and the physical stuff, Harry Fair and his guys did a great job on that. So there were some fairly stressful organizational changes that we did, going through that. The team all came together, so I think there was a lot of camaraderie in it. So I won't claim to be the ‘father’ - I was brought in, you know, as the instigator and the chief nudge, but part architect part transformational leader. That was fun.
https://www.anandtech.com/show/16762/an-anandtech-interview-...
If he says it was a team effort, especially if he’s speaking clearly about the different aspects and who was responsible for them, I wouldn’t disagree with him. He would know after all.
Its human nature to pick out heros, say they did everything by themselves and idolize them. That doesn't mean it's real. It's just a story you're telling yourself.
I've never heard about the people mentioned here, and I don't know anything about semiconductor manufacturing. Except that it's one of the absolute most complicated things humans have ever attempted and there's no way a single person could be the creator of an entire modern processor architecture. So I've been reading through this thread kind of surprised to see people saying this one person did so. But your comment made me realize it's just typical hero worship.
Note: this doesn't belittle the work of anyone. Schumacher is an amazing driver, one of the best in the world. But saying he is responsible for winning formula one races belittles the work of the other people involved, engineers at the top of their game just as much as he is, and yet faceless to most people. He could be exactly the driver he is, or even a hundred times better, and without equally talented engineers behind him he'd still finish in last place. So did he win the races, or did they? Obviously, the answer is that they won together.
Just because they arent celebrities it doesnt mean that they do not exist
https://en.wikipedia.org/wiki/AMD_Platform_Security_Processo...
His greatest hits: DEC Alpha, AMD K7 (Thunderbird/Athlon), AMD K8 (Athlon 64), Apple A5, AMD Zen... probably others I'm forgetting.
Like god damn can someone else in the industry have some ideas of their own please, lol /s
And again not that other people aren't involved either, in particular Mike Clark was really the guy who executed Zen development, but Keller was there at least as an advisor for a lot of the early development. He's one of Zen's uncles, Clark is the father. https://www.youtube.com/watch?v=3vyNzgOP5yw
I think it also says a lot that he completely fucking bailed from Intel after only being there a couple months... I think he saw they just weren't ready/willing to execute well and his time would be wasted there. He's the wandering silicon samurai who drops some golden nuggets of advice and wanders off into the sunset, not a babysitter while you purge middle-managers playing office politics.
(He left due to a "family situation" and while I have no doubt that it was real... I also think he probably might have stayed if Intel wasn't a complete dumpster fire too.)
I was disappointed when Apple switched to x86 instead, and a few years later, Apple acquired P.A. Semi, which I believe became the bulk of Apple's mobile processor team.
If rumor is to be believed, it really is too bad that Ken Olsen refused to drop margins and increase volume on the Alpha in order to create a scaled down version for Apple back when Apple was looking to leave the m68k architecture.
even if he's not the determining factor of success - as he clearly said it's a team effort - he's very likely a safe bet, and a good canary, a good advisor (not a yes-man).
https://www.youtube.com/watch?v=3vyNzgOP5yw
Probably the most interesting nugget is that Zen2 was actually considered a tweak internally and not a full uarch revision, while Zen3 is actually the clean-sheet revision. But it's just a good summary of the general tone and tempo of CPU development I think.
How would one go about doing this?
I would happily accept a bet (e.g. $100 bucks or a nice bottle of whiskey) from you (or anyone) that atomicsemi is successful. I will just bet against it because I am skeptical that a new company is successful in a capital intensive segment.
It is probably slightly insulting, foolish, and arrogant to bet against the famous Jim Killer on a website dedicated to startups. This should not be a dig against Sam and Jim Killer. I guess that will try something truly innovative. But if I look a the past, the odds seem to be stacked against them.
And Tenstorrent I still see as a risky bet, but in my mind it's very +EV.
Everything is ML nowadays and we still use very naive approach driven by what hardware was available at hand.
The Alpha made all of that moot because in one fell stroke it increased the amount of RAM that could be addressed directly to the point where the whole thing could happen in memory without any cluster communications overhead. It was still an expensive machine but it cost a fraction of the setup that it replaced, and performed really very well. A nice example of how vertical scaling can be a very viable option. The 64 bit file system also allowed for much larger files, which helped the project in different ways.
One downside was that spare hardware was difficult to obtain but the system was built like a tank and ran for many years until there were many other suppliers of 64 bit systems.
It was way ahead of the anything else in the 'affordable' range of computers. Though it still cost as much as a nice car fully decked out, especially the RAM was quite expensive.
Azul later realized some of these things on Java. Building a virtual machine and even language to take advantage of that from the ground up would have been cool.
The processor's firmware (PALCode) was essentially a single-tenant hypervisor, and the OS kernel made upcalls to the firmware in order to perform any privileged instructions. Had the architecture survived longer, this would have been handy for virtualization. Modern OS kernels have special cases for upcalls when running on top of hypervisors in order to avoid some of the overhead of the trap-and-emulate code in the hypervisor.
The designers were brutal in only including instructions that could show a performance improvement in simulations. The first versions of the processor didn't have single byte loads or stores, presuming that the standard library string functions would load and store 64-bit words at a time and perform any necessary bit manipulations in registers. They later relented and included an instruction set extension for single-byte operations.
They were also famously brutal in their memory model, leaving as much leeway as possible for hardware to re-order operations. As long as you're correctly using mutexes to protect shared state, the mutex acquisition and releasing code will properly synchronize all of your memory operations. However, if you're implementing lockfree data structures, the Alpha is particularly liberal in its read ordering, and you need read fences on the reader side of lockfree structures, which is unusual. Experience has shown that for most code, the potential performance improvements aren't very significant, especially considering the increased potential for concurrency bugs.
I'm pretty sure that if you had dropped that from the 10th floor of a random office building you'd be fined for damage to the pavement but that machine would have still worked ;) It also took two people to lift it.
And now your average phone has more CPU power and more storage...
Times, they are a changing...
I not a architect and don't know enough about the topic, but I thought that might be something interesting for RISC-V. Love to read about the advantages and disadvantages of that.
In the absence of a read fence, you can speculate the address of a second load, execute this second load in parallel with a first load, and ignore a cache invalidation (or just not wait until you can rule out an asynchronously transmitted one) hitting the second load so long as the address (computed from the first load's data) was correctly predicted. It hits even harder when the predicted address was a cache hit and the first load experienced a cache miss, because now you can speculate execution using the second load's data (delivered from the cache) and retain/confirm/retire the results of the speculated computation as soon as the first load's data returns and the computation of the second load's address confirms the speculated one.
An example is read-only access to data structures with pointer chasing while a different core performs copying garbage collection. Because the data is (semantically) read-only, the old copy and the new copy are both equally valid, and as long as you don't accidentally read the new space before the copy was written into it, you can pointer-chase freely through these structures reading the next e.g. linked-list entry from either the old or the new place (if you speculate correctly).
Critically, this could get by with invalidating only cached data for the copy target range, ensuring readers don't get the uninitialized data, without invalidating their cache of the copy source range. Of course that would require sufficiently targeted invalidation.
Other cases like e.g. typical union-find / disjoint-set datastructures work just fine with standard fence-free Alpha memory accesses, at the slight cost of `union` operations not coherently affecting outcomes of `find` operations. That's often not a problem, though, as parallel applications already have to cope with the `union` racing the "subsequent" `find` operations (and ending up with the `union` happening last).
Different spaces, different people, but I'm going to sit back and observe I think. I'm just kind of done being excited by people announcing the partnerships, once the partnership bears fruit I will have a look and then decide if I'm excited.