Jim Keller moves to AI chip startup
reuters.com
reuters.com
"Intel’s press release today states that Jim Keller is leaving the position on June 11th ( 2020 ) due to personal reasons. However, he will remain with the company as a consultant for six months in order to assist with the transition." [1]
Exactly six months later he took a new job. Some may want to look back at their comment on the subject. [2] [3]
Still waiting for the Story between Jim and Gerard.[4]
[1] https://www.anandtech.com/show/15846/jim-keller-resigns-from...
[2] https://news.ycombinator.com/item?id=23496083
I loved also how he explained why books are good, some takes 20 years of his experience and writes it in 200 pages...
Thanks for the link!
Absolutely fascinating interview.
Lex has a PhD in machine learning but doesn't seem to be familiar with branch prediction, apparently.
I mean, I have a PhD and had no idea what branch prediction was until I listened to that podcast.
My provisional answer is that I'd host a conversation between experts.
For example, I'm now very curious about Apple Silicon M1 wrt to Java's Memory Model and Project Loom (structured concurrency). But I don't know nearly enough to even ask smart questions, much less understand the answers.
So my dream future perfect interview would have Ron Pressler, Doug Lea, and one or two people really smart about M1 (the only name I know is Dan Luu) sit around and chat it up.
I'd ask them open ended questions, like "What's new and different?" "What happens next?" "What are you excited about?"
The conversations would likely happen over multiple sessions and different mediums. Because the experts would share and ask each other stuff which would prompt followups.
As podcast host, I'd try to be catalyst, try to remove myself from the convo as much as possible. I can't think of any examples, role models. While I'm a huge fan of Ezra Klein and Adam Gordon Bell (Corecursive), I'm not confident I could lean in like they do.
One tactic both Lex and AGB do really well is prompt their guests to explicitly define jargon. I suspect that some of the perceptions of Lex's ignorance are him trying to make topics more accessible. eg Working close to the metal with AI, I'm quite confident Lex knows about branch prediction.
Even if you know about branch prediction, then asking the guest to explain it, maybe even pretending not to know about it, is a great way to have concepts introduced and make things more approachable.
Lex wouldn't be as popular as he was if he didn't have a good sense for the level of knowledge his ideal listener has about the subject.
If his successors applied his management style, mobile phones would have an Intel processor in them.
PS: Nice nickname. Read the book from Th. Mann - over a period of 7 years I think.
There's no arguing that Keller is a smart guy, but he doesn't design an entire CPU architecture himself. If you're desperate and say "Okay, we are hiring the smartest guy we can find to build our new CPU, and give everything he needs to make it happen", then perhaps you get the AMD64, Zen or the A4 and A5. If you try to just dump a smart guy into a team as just another engineer, maybe you get nothing, like Intel.
Perhaps AMD, who already knew him, just gave Keller everything he need to build a team that can deliver on a new architecture, even when he's no longer there. Same with Apple. Intel on the other hand may have been unwilling to grant Keller the same level of autonomy and control. Then it also makes sense that he would leave Intel, for personal reasons, those being: "I can't work here, they won't let me do my job".
If the company is growing, there are new X-of-Y positions to move up to.
If the company is stable or shrinking, people start watching out for their own careers with knives out.
AMD possibly avoided this because of size & realization of what needed to be done. Intel's too big & old: I would be very surprised if they weren't much more internally resistant to that sort of change.
And one can only deal with your colleagues throwing up brick wall after brick wall on every bit of minutiae for so long, as least if you're talented enough to have other options.
I think we are already starting to see fruit of his work. Intel doesn't need Jim Keller for CPU uArch design. Intel has had their uArch roadmap ready, and they were the best in the Industry if it wasn't for the 10nm delay. They also have work in the pipeline all being held back by their process node.
Jim described it in one of his interview ( Sorry I spent 10 min but couldn't find the source, so I may have remembered it wrong ) about not having process node held back your chip design, where he has experience in doing so in AMD and Apple. Being flexible enough to back port your design should anything happen as Plan B. Where previously Intel was just keep waiting for the process guys to fix it. That in itself is a huge workflow changes. It is hard to imagine the amount of work required to push this through especially with all the internal politics at Intel.
And Intel is at least looking at TSMC / alternative paths for some of their product lineup now ( Gaming Focused Large- Die Size GPU ) . Whether that is decided or not is unclear. But at least we have Rocket Lake launching soon which is sort of a half baked Willow Cove ( Used in Tiger Lake ) ported back to 14nm on Desktop. And we have Sapphire Rapid as well as other product roadmap hinting at multiple node ( Shown in Investor meetings notes ). That is at least showing Intel has changed their Internal design to be flexible enough in case of another 10nm like fiasco. And I think Jim Keller has some credit in this transition.
That is of course, having flexible design still doesn't fix their problem if TSMC is 2 years ahead of Intel in leading edge node design, volume and cost. And as I have repeatedly stated, Intel's problem is not design, but their business Model. And It would not surprise me if TSMC have shipped more 5nm wafers last year ( 2020 ) than Intel's entire 10nm production history since 2017.
Just let that sink in for a bit.
You're probably going to see a whole lot more of this sort of thing given the limits to process scaling. Keeping things simple and backwardly compatible made sense when you could just throw more transistors at the problem. Now you're seeing more and more specialized circuitry that software people are just going to have to deal with.
Compare that to the non-scalable SIMD instructions that mean you have to rewrite your code to take advantage of them and resultingly people don't bother to use them at all.
AMX allocates a huge die area to GEMM functionality that gets used a lot less in real numerics than you'd gather from reading a linear algebra textbook.
There are other approaches to the problems the industry faces other than 'fill up the die with registers that will never be used', nvidia and apple are going that way and that is why they are succeeding and Intel is failing.
These statements almost seem contradictory. What if instead of "not being able to right that ship", it is instead an example to the contrary?
Intel started out with memory. They never left and are on by far the cutting edge with memory tech. In particular Optane NVM DIMMs are so fast they basically define a new layer in the performance/cache hierarchy. Intel might see a shift in their focus over time away from CPUs to chalcogenide based persistent memory, where it seems they have held the lead for some time now.
Why? Because I never ever see any articles talking about any other interesting employees. Every single time it's Jim Keller.
I'm sure he's good, his interview with Lex Fridman shows that he's knowledgeable and creative, but there's no way he's as exclusive a major force as the media portrays him.
For example, one could round up as many scientists as they could find in 1900, but there is no number that would guarantee the progress made in theoretical physics by someone like Einstein alone.
The person can be a key motivator to bring the team forward, or reach decisions they wouldn't otherwise have taken.
Was he exclusively working on "that ship" ... wait, I thought you said "chip".
Keller can move the entire market if he’s given enough space and resources. 10,000x is closer to him than 10x.
A 10x engineer might not get 10x done, but their work will be 10x better in a combination of ways: quality, maintainability, speed, portability, extendability, etc. Hopefully the ways in which their work is better fits the priorities of the organization.
That’s the only way you can really call someone an N-xer compared to a productive individual contributor.
Someone like Jim Keller is a big multiplier at a higher level. People today understand the value an executive like Steve Jobs brings, but usually there’s debate on the value before it becomes clear a few years later.
[1] https://www.youtube.com/watch?v=Nb2tebYAaOA
[2] https://www.youtube.com/watch?v=Nb2tebYAaOA&t=3973s
I think you meant to refer to [2], and that part of the interview is Lex Fridman interjecting him when he was trying to make a deep point about how to think about things.
Love the TechTechPotato breakdown of this BTW :)
I can't wait to stop writing machine code.
Mind you, that same thing is now done to software devs, that is, product manager tries to explain what they want, designers and devs interpret in a certain way.
There are a few aspects of it that I'm really enjoying:
- I can now actually understand the disassembled code that I see during debugging. This includes recognizing some of the assembly patterns that appear because of ABI requirements and/or common programming idioms.
- I'm becoming comfortable with a programming idiom that I've never really used in the past: registers, flags, various kinds of memory addressing.
- It helps my understanding of compilers' lower levels / backends, and the related problems: register allocation, instruction selection, etc.
- It provides a clear path for my first attempt at writing JIT code (using Xbyak[0]).
So as Richard Feynman might have said, it's great fun!
I don't notice anything in particular that stands out vs. the many other AI chips people are making, at first glance. But I'm far from an expert. There are several other technical videos on their YouTube channel as well: https://www.youtube.com/channel/UC7041p6DlAh0r4_Fnlk10pQ
I’m pretty surprised no-one has actually exposed the actor model for parallelising neural networks, it seems it would work quite well and allow you to have a layer per node (or actually many split configurations). Maybe data locality would be an issue with actor based approaches. They seem to be solving this at a lower level but with less knowledge of the actual parallelism in software.
I wonder what they are referring to. Are they accelerating what SHAP's GradientExplainer [1] does? (namely: crafting inputs at a specific layer, propagating forward to see the influence on class prediction, and sort of backpropagating to pixels) Or is it about something more related to Judea Pearl's work on causality?
[1] https://github.com/slundberg/shap#deep-learning-example-with...
I think the same goes with this guy, Jim Keller quit Intel and could join another company sooner or later, and that will not be a former one, likely a startup. We are Genius.
Could it be possible he's "famous for being famous"?
E.g. Some might say Jeff Dean (Google) fits this mode a bit, whereas Sanjay Ghemawat (Google) has contributed arguably just as much if not more - but is mentioned radically less than Jeff.