AMD CEO: The Next Challenge Is Energy Efficiency
spectrum.ieee.org
spectrum.ieee.org
But wow there's a bunch of power burned on interconnect between CCDs and CCX. And now AMD's new southbridge, Promontory 21, made by Asmedia, is another pretty significant power hog, and the flagship X670 tier is powered by two of these.
There's absolutely a challenge to bring power down. I'm incredibly super impressed by AMD's showing, & they've done very well. But they've been making trade-offs that have pretty large net impacts, especially if we measure at idle power.
Lots of fast PCIe lanes eat a lot of energy. A server motherboard contains tons of more PCI devices when compared to a consumer desktop systems, and they are not the cards, but the small units enabling the advanced features in servers, which are embedded on the motherboards themselves.
Thunderbolt is just a PCIe encapsulator of some sort, which can also do plethora of other things.
This explains Intel’s squandering of PCIe lanes for consumer desktops versus AMD’s generosity.
So overall, AMD has more and faster IO available directly from the CPU, but less lanes from the chipset, and with a weaker connection to the chipset. If PCIe 5.0 drives become available and the transfer speed to storage is important, I'd say AMD is better, otherwise I'd say Intel has more IO.
Where "top" means "most expensive". The Z790 board I recently purchased for around $300 was pretty barebones and lackluster (no TB, meager IO from ports and headers, wattage constrained VRM relative to 13th gen TDP, etc), but it was the least costly way to work with an Intel proprietary technology.
It'll last another four or five years, but it was my first Intel build since the slocket days, and likely my last.
You do get to use the AMD boards for more than two CPU generations though.
I wish modern chips did a better job of breaking down where power went. It'd be so interesting to know how much power is going to usb controllers, how much is going to PCIe. I'd also hope that they could do things like shut down parts of the chip, if there's no USB or PCIe devices plugged in. But these chips seem to have a pretty high starting place of power consumption. Although maybe it's in part because the first example were flagship motherboards with a whole bunch of extra things peppered across the board - fancy NIC chips, supplementary thunderbolt controllers, sound cards, wifi - so maybe there was just an unusual lot of extra stuff going on. But it has been shocking seeing idle power raise so much on the modern platforms. It feels like there's a lot of room for improvement in power-down.
[0] https://www.macrotrends.net/stocks/charts/TSM/taiwan-semicon...
CCD stands for Core Complex Die (and neither terms refer to the IO die)
But it failed hard which coincided with Intel releasing a major winner with their last planar architecture - sandy bridge.
As a result, AMD spent years circling the drain and their stock dipped below 2 dollars. Some people made good money buying around that time.
[0] https://en.wikipedia.org/wiki/Bulldozer_(microarchitecture)
Edit: oh, I missed the exclamation mark haha. Oh well. Too tired to even feel ashamed
This article reminds of this other one [1] posted about 1 month ago.
It’s an interview with some guys that just got done building an exascale supercomputer, in which it was originally estimated to need 1000 megawatts but ultimately only needs 60. The reporter asks about zettascale and the power requirements; they wave it off and say that the big question about whether it will even be possible in the next 10 years is getting the chip lithography small enough so that you can physically build a working zettascale supercomputer.
Half algorithm & half user manual?
Half algorithm & half class?
(solve (this :by strong-ai) (that :by weak-ai))- Half "hard maths" algorithms. i.e. cominbatorics, geometry, etc.
- Half "fuzzy maths" algorithms. i.e. heuristics, approximation, machine learning.
The idea being to solve the parts that can be easily solved by hard maths with those hard maths so that you can reduce the problem space for when you apply the fuzzy maths to solve the rest of the problem.
In other words, it's taking the problem, breaking out discrete pieces to solve with well established hard maths, using heuristics & numerical solutions to tackle the remaining known problems without "easy" analytical solutions, then using ML to fill in all the gaps and glue the whole thing together.
Interestingly the example is backwards (statistical reasoning first, hard reasoning second) compared to traditional usage of "hybrid" in AI and control contexts.
https://www.nature.com/articles/s41586-019-0912-1
and here
https://www.nature.com/articles/s42256-021-00374-3
> The next step will be a hybrid modelling approach, coupling physical process models with the versatility of data-driven machine learning.
The Frontier HPC system that AMD just delivered is aimed fully at that problem.
That or the algorithms use a combination battery and gasoline.
Conventional processing + AI together. Hybrid approaches.
I expect them to use Xilinix's AI engines primarily in their CPUs, APUs and GPUs - not so much FPGA.
----------
AMD is doing the right thing with Xilinx tech: they're integrating it into ROCm, so that Xilinx AI engines / FPGAs can interact with CPUs and GPUs. But there's no reason why these "internal core bits" should be shared between CPU, GPU, and FPGA.
Of course performance wise it doesn't touch the $1k+ graphics cards with crazy amounts of ram, but for students and if I need to do something quick on the go, its a really useful tool.
https://arxiv.org/abs/2202.11214
A neural network-based solution for weather forecasting: "FourCastNet is about 45,000 times faster than traditional NWP models on a node-hour basis."
It's supposed to. Investors want to hear this, not some crap about efficiency, when the entire world is talking about AI.
Edit: TBC, I care about efficiency and it's not crap, but that's likely the view of investors.
Oddly, to me this just sounds like efficiency gains by potentially introducing massive security holes i.e. the vector of the Meltdown/Spectres of the future. It also seems like they're trying to sell AI as some sort of secular qubit that they'll be error-correcting.
I love the battery life and performance of the hardware, not to mention the unrivaled build quality of the MacBook (screen, trackpad, keyboard).
In practice, however, MacOS limits the capabilities of the hardware such that I cannot daily drive my MacBook Pro as a work or personal computer (poor containerization support, an annoying development toolchain, and no _real_ support for video games).
When Asahi Linux is mainlined, stable and features full hardware acceleration - the MacBook running Linux will likely be the best laptop money could buy. Until then, please AMD, Intel, release some mobile hardware that's at least as good. It sucks so bad seeing what is possible with today's technology but that being exclusive to a company unsuccessfully determined to ring fence you into their API ecosystem.
docker on mac can be improved, but if i'm developing for other architectures it's much easier to just test natively. toolchain for all the languages i use is exactly the same as any other *nix.
gaming i'll give you a point, but that's why i have windows dual boot at home. /shrug
And it runs Office better than Linux or Mac.
WSL?
I also like the productivity customizations posseble in Linux while the same are difficult or impossible on Windows.
of course if your work is tied into the windows eco system having WSL is good to have.
WSL2's virtualized workflow just causes too many issues for me. WSL1 was better IMO but it wasn't significantly better than msys2 and also had issues (like you still need remote development tools to mount codebases inside editors) - unless you want to run/develop Linux binaries while on Windows.
For anything that isn't making basic non containerized applications (simple web applications, web servers), Windows is pretty good.
For anything more involved, requires multiple containers/compose/etc, I prefer Linux as it has the tools I need available natively and no gotyas or performance penalties.
That said, credit to Microsoft on WSL2. The auto-scaling hardware provisioning inside the VM has made containerized workflows on Windows much better. To me, it's just not better than running Linux inside VMWare/Hyper-V/VBox and "DIY"ing WSL2 yourself, something I had been doing for years before WSL2 anyway. WSL2 is more fool-proof then hand-rolling a Linux VM, so there is that.
Yup, been doing that for decades. It's called "vendor lock-in". And its unethicality has been discussed comprehensively for decades as well.
And when countries tried to move into an open document format - they'll bully the country, using the strong arm of the Uncle Sam, until they're back to MS Office again.
So when people are amazed by Bill Gates' charities - I don't. His money comes from the sufferings of countries.
I’ve met Bill. He wouldn’t be out of place among HN’s defective half: the dubious pro SaaS VC funded US university alumni…
Yup, it's NOT Linux.
Been using MacOSX since Powerbook up to first gen Macbook - and I finally gave up: I installed Linux on it instead.
Pretty much all servers are running Linux even back then. Using Unix on work computer/laptop causes way too many encounters with various quirks and glitches. It continuously drags down your productivity too many times everyday.
After much stress, I installed Linux on my Macbook instead - but then I encountered various hardware-related glitches instead.
So I finally gave up, and bought a Thinkpad.
Just run an always-on (headless!) Linux VM in the background, and don't use the host macOS for anything besides desktop apps (Slack, VSCode, Browser, mpv, terminal emulator but always ssh into the VM, etc). The same way you deal with a Windows machine.
This works good enough unless you work on hypervisors or other bare metal only tech. But hey, that's currently non-existing on M1 Macbooks (or undocumented and locked to Apple ecosystem) anyways.
> and no _real_ support for video games
That's the real deal breaker if you are into gaming.
Sometimes working local is just the easiest/fastest most convenient.
This is what I do. Specifically, I use Canonical's Multipass, and treat tmux as my "window manager," -- I even mapped my iTerm profile's Command+[key] to the hex code for tmux-prefix+[key], so that Command essentially feels like the Super key in, say, i3. For example, rather than having to type Control+p h, Command+h selects the pane to the left.
The Multipass VM is flawless. Closing the laptop doesn't shut it down, and with the tmux resurrect plugin, sessions persist between Mac restarts (which are rare). If I didn't know better, I'd think it was just a native terminal session.
If I need proper x86_64, I just ssh into my super beefy Linux NAS at home via Tailscale. Both Linux machines are identical in terms of dotfiles/etc, so it feels exactly the same.
I've truly never been happier with Linux. I no longer obsess over my window manager (which used to be a serious time sink for me), and I still get what is IMO the best desktop experience via my Mac. The only tinkering I do is checking out the occasional new neovim plugin, but I really enjoy doing that, as it has a tangible benefit to my dev workflow, and kind of feels like gardening, in a way -- I like the slow but persistent act of improving and culling my environment.
I still have a PC, but I actually recently uninstalled WSL2. It never felt truly finished or "right", and Windows Terminal can be incredibly sluggish -- keystrokes have far more latency than iTerm. I've actually started to embrace just letting "Windows be Windows," even learning Powershell (and enjoying it more than I'd expected).
I've also pretty much moved from gaming on a PC to PS5. So, for me personally, I don't really see a place for Windows anymore. Every single time I boot into my PC, something is wrong -- most recently, it literally won't shutdown unless I execute `shutdown /s`, and no amount of troubleshooting has been able to fix it. I know Windows like the back of my hand, and still it's a constant feeling of death by a thousand cuts.
On the other hand, for the missing apps and other stuff, I'm running a Linux VM via VMWare Fusion Pro. It works efficiently, and interoperability is good. Just add another internal network card to the VM, and keep an always on SSH/SFTP connection. Then everything works seamlessly.
Never had any problems for 7 or so years when developing that application and writing my Ph.D. in the process.
Because I'm using my laptop to, well, work on the go a lot; this is a showstopper.
I'm doing on that my 2014MBP, and it didn't cut the endurance to half, but need to re-test it for exact numbers. However, it doesn't appear in "power hungry applications" list unless you continuously compile something or run some service at 100% CPU load. Also, you can limit the resources it can use if you want to further limit it down.
The being said, the M1 Pro is a more than good enough piece of hardware that I'm willing to put up with this.
Would Wine (Proton) work out of the box on M1 Linux? My knowledge of Wine is limited. Based on the How Wine works 101 article [1]:
> The code inside the executables is “portable” between Windows and Linux (assuming the same CPU architecture).
because the CPU architecture is different, Wine would need changes to support it too - just like x32 and x64. Is that correct?
I use UTM and Docker on a daily basis and it is extremely smooth. What exactly is missing?
> an annoying development toolchain
What are you talking about exactly? For example most Python and Rust builds just work out of the box.
> no _real_ support for video games
https://docs.unity3d.com/Manual/Metal.html
These claims look like a bit of an exaggeration.
https://social.treehouse.systems/@marcan/109838053800961073
Asahi Linux is introducing support for some brand new Apple Silicon features faster than macOS.. M1 has a virtual GIC interrupt controller for enhanced virtualization performance. Linux supports it, macOS does not.. M2 introduced Nested Virtualization support. The patches for supporting that on Linux are in review; macOS still doesn't support it.Curious what will enable this in M2 vs M1. Looking at https://developer.arm.com/documentation/102142/0100/Nested-v... it appears to indicate nested virtualization is in Armv8.3-A and both M1 and M2 are ARMv8.5-A according to https://en.wikipedia.org/wiki/Apple_M1 and https://en.wikipedia.org/wiki/Apple_M2
Will be interesting to see what other arm cpus gain this with the Linux patches.
Have you used (non-nested) virtualization at all on Asahi? If so what is your experience with performance and overall thoughts so far?
Kernel integration and a virtualized filesystem that isn't bottlenecked by APFS. Docker is excruciating on Darwin systems.
> What are you talking about exactly?
Apple makes hundreds of weird concessions that are non-standard on UNIX-like machines. Booting up a machine with zsh and pico as your defaults is not a normal experience for most sysadmins, nevermind the laundry-list of MacOS quirks that make it a pain to maintain. For personal use, I don't think I'd ever go back to fixing Mac-exclusive issues in my free time.
> no _real_ support for video games
Besides Resident Evil and No Man's Sky (this generation's Tomb Raider and Monument Valley), nobody writes video games for Metal unless Apple pays them to.
For a while, MacOS had a working DirectX translation stack for Windows games, too. Not since Catalina though.
Docker is great as long as you don't use bind mounts. I use it daily for development in dev containers.
> Besides Resident Evil and No Man's Sky (this generation's Tomb Raider and Monument Valley), nobody writes video games for Metal unless Apple pays them to.
There are plenty of great games on macOS. Factorio, Civilization, League of Legends, Minecraft. But you're right that there aren't too many AAA games.
> Rust builds just work out of the box
Works great until you're the author of an application written in Rust and want to distribute MacOS binaries which are automatically generated in CI/CD.
The only (legal) way to compile a Rust binary that targets MacOS is on a Mac. So your CI needs a special case for MacOS running on a MacOS agent. Annoyingly, cross compiling CPUs architectures doesn't work so you need to an Intel and arm64 Mac CI agent - the latter being unavailable via Github actions.
To make things even more bizarre, Apple doesn't offer a server variant of MacOS or Mac hardware, which seems to indicate they expect you to manually compile binaries on your local machine for applications you intend to distribute.
https://machinelearning.apple.com/research/stable-diffusion-...
https://huggingface.co/blog/diffusers-coreml
Not surprisingly, they are quite a bit slower than those power-hungry nVidia GPUs.
Typing this on a ThinkPad X13, after years of MBPs, I beg to differ. The screen on the ThinkPad is better, even at a slightly lower resolution of 1920 x 1200. It's IPS like the Mac of course, but also Anti-Reflective (matte), which Apple hasn't offered since 2008?
The ThinkPad has Left, Right, and Scroll mouse buttons, as well as a TrackPoint stick. Not so on the Mac.
Finally, the keyboard. You're going to laud Apple for their quality keyboards, really? The ThinkPad has a nicer keyboard feel (subjective, I know), has actual Home, End, PageUp, PageDn, and Delete keys, two Ctrl keys, and praise Jesus, gaps between the function keys, so you can use them confidently w/o looking.
Intel is starting to make changes towards this structure but they haven't fully committed to it yet.
In particular, the RDNA3 graphics card line launch has been a dud.
The Nvidia 4090 turned out to be far ahead of the AMD 7900XTX.
Then there turned out to be an overheating issue on the AMD cards.
And now it just looks like both Nvidia and AMD are price gouging GPU buyers instead of competing with each other. They are deliberately keeping prices high and creating artificial shortages of GPUs, because that's what kept prices high during covid/crypto mining.
I used to be really cheering for AMD as the underdog, but I guess its true that none of these companies are your friend, they're just there to shake you down.
AMD had a chance to really pull ahead of Nvidia by being "the good guy" and actually offering end users great value for money, but instead they've chosen to emulate Nvidia.
Substantial customer good will has been lost by AMD.
you've seen evidence for this, or it's your opinion?
>AMD had a chance to really pull ahead of Nvidia
with a graphics card line that you just told us is far behind Nvidia's?
Probably alluding to the under-shipping of GPUs; stated by Lisa Su during the recent investor call[0]
[0]https://www.fool.com/earnings/call-transcripts/2023/02/01/ad...
Cheap and good enough can absolutely be a way to win.
If AMD actually competed on price then they'd be shifting the GPU market away from Nvidia.
They seem content however just be second best but making bank.
Nvidia relaunched the 4080 12gb as the 4070ti reducing the price by $100. There has to be one hell of a profit margin on the high end cards.
The take from multiple[2][3][4] journalists on that call is they're trying to avoid a "supply glut" and maintain high prices.
[1] https://seekingalpha.com/article/4574091-advanced-micro-devi...
[2] https://www.pcgamer.com/amd-undershipping-chips-to-help-prop...
[3] https://gamerant.com/amd-undershipping-graphics-cards/
[4] https://www.extremetech.com/computing/342781-amd-ceo-says-it...
they're talking to the public, i.e. investors, and they could easily be saying "our current sales figures are lower not because our product is not popular, but because there is currently a large inventory downstream to meet current demand. When that glut is cleared, expect our sales to resume."
if downstream sellers have sufficient inventory, the only way to induce them to buy more would be for AMD to drop prices. If AMD cards are in hot demand and selling out immediately, restricting supply would be artificially boosting prices. But if the downstream pipeline is full, it's not right to say that reduced demand from wholesalers while that glut clears is AMD artificially boosting prices.
> They are deliberately keeping prices high and creating artificial shortages of GPUs
That's simply not true. Check out TechTechPotato's youtube channel for the explanation on this, but under-shipping isn't price fixing. They're just shipping less to distributors because there's less demand, allowing the distributors and retailers to keep a stable amount on hand.
Seeing plenty of people choose 13th gen over Zen 4, the platform pricing for Zen 4 just wasn't very attractive. AMD [had to] significantly cut prices across the lineup by 20-30 %.
Nobody needs two-digit CPU core counts and 5~6GHz clock speeds to do their emails, communicate on Skype/Discord/Teams/Slack/Zoom/whatever, browse Facebook and Twitter, watch Youtube, and even play some vidja gaemz. An i3 or even a god damn Celeron is perfectly fine.
So at that point, Intel's superior stability (read: less jank) wins out by a hair and otherwise nobody really cares because there's no practical difference. The vast majority of people will just buy whatever's cheaper or just happens to be on the display table that day.
RDNA3's lower end chips, once they hit the market, are expected to further improve on this.
Most gamers won't upgrade to the current RDNA3 chips, because the current RDNA3 chips are top of the line, expensive, ~300w monsters.
I keep eyeing a new build, but realistically, I know it’s just a vanity project because so few games will take full advantage of the better hardware. My favorite games in the past years could have run on ten year old hardware.
Performance per watt has really come a long way since GCN.
I'm not sure if you're referring to AMD "undershipping", but if you read how they use that word it's pretty clearly a bad thing. AMD has been shipping less (to retailers) than what they could, or would like to.
7900XTX is on par or better than 4090 in alot of games, as well as some games favouring nvidia more than amd... (obviously not talking about Ray Tracing)
Coupled with a lower price...
I'm not quite sure what you're talking about since you're saying the oppisite is all reviews I've seen.
@ 4k gaming then 4090 on average is /much/ better than the 7900xtx, but looking at 1080/1440p that lead deminishes alot.
edit: At the end of the day tho its all too damn expensive now.
I just don't get why most people care? It's 60% more expensive, too. Even the 7900XTX is ludicrous at $1000.
Give me a $400 card from this generation that competes with a $500 card from the previous generation, and I'll call it a win.
A $1600 card winning anything seems like an irrelevant battle. Is the volume / sales for those cards high enough to be the real focus, when cards like the GTX 1060 were the volume leaders by a long shot?
My impression is that the simplest way to improve energy efficiency is to simplify hardware. Silicon is spent isolating software, etc. Time is spent copying data from kernel space to user space. Shift the burden of correctness to compilers, and use proof-carrying code to convince OSes a binary is safe. Let hardware continue managing what it's good at (e.g., out-of-order execution.) But I want a single address space with absolutely no virtualization.
Some may ask "isn't this dangerous? what if there are bugs in the verification process?" But isn't this the same as a bug in the hardware you're relying on for safety? Why is the hardware easier to get right? Isn't it cheaper to patch a software bug than a hardware bug?
If the compiler (like rust) can prove that OOB memory is never accessed, the hardware/kernel/etc don't need to check at all anymore.
And your proof technology isn't even that scary: just compile the code yourself. If you trust the compiler and the compiler doesn't complain, you can assume the resulting binary is correct. And if a bug/0day is found, just patch and recompile.
Removing these checks from the hardware is possible only if you can do without it 100% of the time; if you can trust that 99% of the binaries executed, that's not enough, you still need this 'enforced sandboxing' functionality.
A "software replacement for MMU" would thus need to solve fragmentation of the address space. This is something you would solve using a "heavier" runtime (e.g. every process/object needs to be able to relocate). But this may very well end up being slower than a normal MMU, just without the safety of the MMU.
Even in places where DMA is fully warranted, IOMMU gets shoe-horned in. I don't think there's any running away from costs to be paid for security (not the least for power-efficiency reasons).
Arbitrarily complex programs makes even defining what is and isnt a bug arbitrarily complex
Did you want the computer to switch off at random button press; did you want two processes to swap half their memory. Maybe, maybe not
A second problem to consider is that verification is arbitrarily harder than simply running a program -- often to the extent of being impossible, even for sensible and useful functionality. This is why programs that get verified either don't allocate or do bounded allocations. But unbounded allocation is useful
It is possible to push proven or sanboxed parts across the kernel boundary. Maybe we should increase those opportunities?
Also separate address spaces simplify separate threads -- since they do not need to keep updating a single shared address space. So L1 and L2 cache should definitely give address separation. Page tables is one way to maintain that illusion for the shared resource of main memory... Probably a good thing
That's not to say there isn't a lot of space to explore your idea. It is probably an idea worth following
One final thought: verification is complex because computers are complex. Simplifying how processes interact at the hardware level. Shifts the burden of verification from arbitrarily long running and arbitrarily complex and changing software; to verifying fixed and predefined limitations on functionality. That second one has got to be the easier to verify
I do not believe such OSes can ever be secure given how often vulnerabilities are found in web browsers's JS engines alone. Besides, AFAIK the only effective mitigation against all Spectre variants is using separate address spaces.
The amount of processing power available in a modern smartphone is truly mind-boggling. I'd love to see a chart showing the chip cost and energy cost of the power on an M1 chip in each previou syear. I would guess that 30+ years ago you'd be in the millions of dollars and watts of power but that's just a guess.
As we see from the modern M1/M2 Macbooks, these lower TDP SoCs are more than capable of running a computer for most people for most things. The need for an Intel or AMD CPU is shrinking. It's still there and very real but the waters are rising.
Some don't, but a phone does.
Google has purged 32-bit apps from the official Android app store, but as I understand it the Chinese OEMs that ship un-Googled AOSP ROMs with their own app stores haven't been as aggressive about moving to 64-bit.
M1 has 16 billion of them
AMD is below 10 billion
it's not that expensive
it's a tradeoff, you lose something there, you gain somewhere else.
x86 has a more complex decoder exactly because it was less powerful and had to save on computing power and energy consumption, not being a mainframe.
First of all, when you downclock anything, you're going to gain efficiency. If Apple downclocks M1, it can get even more efficient.
Second, most of these tests use Cinebench, which is highly optimized for x86, not ARM instructions. Geekbench should be used instead.
Third, the M1 is a SoC. Everything is on it. Everything is efficiently connected directly inside the chip.
The benefit of ARM is scale of multitasking due to not requiring the same kind of lock states that Intel's architecture requires, and can additionally scale much better than only one physical+virtual core pair.
I guess the only thing that's holding back ARM is Microsoft, as laptops are expected to run an desktop OS that people are comfortable with. Windows RT wasn't really a serious desktop OS and rather a joke made only for some IoT enterprises instead of end-users.
I wish there was more serious hardware than the standard broadcom or MediaTek chips, I'd definitely want some of that...be it as a mini ATX desktop/server format (e.g. as a competitor to Intel NUC or Mac Mini) or as a laptop.
With the ongoing energy crisis something like solar powered servers would be so much more feasible than with x86 hardware.
The high power ARM cores aren’t that small. If you took the M2 and scaled it up to 256 cores, it would be almost 7 square inches. You can’t just scale a chip like that, though, so the interconnects would consume a huge amount of space as well. It would also consume over 1000W.
The latest ARM chips are great, but some times I think the perception has shifted too far past the reality.
The actual cores are about .6/2.3 mm², and local interconnects and L2 roughly double that.
So with just those parts, 256 P-cores would be about 1.5 square inches, and 256 E-cores would be about half a square inch. And in practical terms you can fabricate a die that's a bit more than a square inch.
Of course it wouldn't use 1000 watts. When you light up that many cores at once you use them at lower power. And I doubt a 256 core design would have all that many P cores either.
As a rough estimate, you could take the 120mm² M1 chip, add 28 more P-cores with 110mm², 220 more E-cores with 300mm², 128 more MB of L3 cache with 60mm², 100mm² of miscellaneous interconnects, and still be on par with a high end GPU.
That sounds doable but is pushing it. A 128 core die, though, has nothing stopping it except market fit.
And the memory controllers aren't that big on the die. You could include a bunch more on a 128 core model.
Source: https://www.anandtech.com/show/16979/the-ampere-altra-max-re...
It's not Microsoft holding it back. It's Qualcomm.
Apart from their very latest SOC (designed by a bunch of ex-Apple employees, no less) their CPUs have are significantly worse than x86 in terms of general performance and have persistently lagged 4 years behind Apple in terms of performance (3 years behind x86). They sell for the same price per unit as x86 CPUs do, so there aren't very many OEMs that take them up on the offer given the added expense of having to design a completely different mainboard for a particular chassis.
As such, x86 is the only game in town if you're buying a non-Apple machine; Qualcomm's products aren't cheaper and perform much worse outside of having more batter life. Sure, Qualcomm owns Nuvia now, but that acquisition will still take some time to bear fruit.
It might be a very long time considering Arm is suing to get Qualcomm to destroy Nuvia's work.
I have no idea what you mean by this. The only x86 feature I can think of that might qualify as a 'lock state' is a bus lock that happens when an atomic read-modify-write operation is split over two cache lines. That has a very simple solution ('don't do that'--you have no reason to), and anyway, one can imagine more efficient implementation strategies
> can additionally scale much better than only one physical+virtual core pair
I have no idea what you mean by this either. Wider hyperthreading? It can be worthwhile for some workloads (and e.g. some ibm cpus have 4-way hyperthreading), but is not a panacea; there are tradeoffs involved.
Let's get real here, most things can't be parallelised at all. We must strive for better single core performance.
These days I only see a single core loaded up to 100% when I grep through a big directory or when I encounter a bug in some software.
Most of the time, it's either all cores are equally idling or equally doing something heavy (like building a big project).
That’s a bold statement as I type all day on an M1 Mac. My FT100 company just made the leap to them as dev machines.
Do you have any information that Microsoft is planning to support RISC-V at least as well as x86/x64? (That is to say, not with something like Windows RT, or Windows CE)
That would be tremendously good news, I shall add.
What you really mean is Arm isn’t the hot new thing anymore. Well it hasn’t been that for 20 years. Meanwhile billions of arm devices in leading edge nodes are being shipped. Oh well.
citation needed
https://www.theregister.com/2023/02/08/5_percent_cloud_arm/
"5% of the cloud now runs on Arm as chip designer plans 2023 IPO"
5% does not seems to me as a position that you want to change to something else.
ARM in every day computing outside of mobile phones and SBCs is just getting started as I see it.
I hope this never comes to pass, because a RISC-V future is a Chinese future.
A Chinese future will not be kind to western ideals that most of us hold dear.
Now, yes, China will just espionage and kangaroo court their way through and around such legalities anyway, but nonetheless RISC-V is less effort for more reward for China if it becomes at least on par with x86 and ARM.
Put more basically, it's a matter of national security. China can have an entire RISC-V ecosystem indigenously, unlike x86 and ARM.
RISC-V by contrast is much, much harder for any given country to regulate because of its free and open nature. At most the US and UK can embargo individual developments made within their jurisdictions, but they can't regulate the entire architecture. RISC-V doesn't have a kill switch named Intel/AMD or ARM.
https://www.reuters.com/technology/arm-china-says-its-ousted...
Though the rp2040 has largely ended my cheap risc-v addiction
Most, if not all SemiCoductor “folks” I know are very pragmatic. As in how a Real Engineer should be, unlike software engineers. And in my experience, only HN and the Internet are suggesting ARM is dead. Everything will be RISC-V.
It's getting traction coz chip companies don't want to pay ARM for licensing, not because it is particularly better at something.
ASUS UL30A-X5
Really an excellent computer, ran linux great (games didn't really exist yet though), and with tuning was coming in under 10W if the display brightness was turned down. First time I was able to get through flights without the system running dead.
I think in this case what's going on is that temperature rises increase resistance in a chip and therefore cause lower efficiency. If you can keep it cool, you can keep it more efficient. The move seems like a necessary one, a computer as powerful as that UL30A is probably inside the phone if you turn off the radio and display, that thing still had a giant battery and only lasted 10-12 hours.
I've seen AMD do some pretty impressive things, I wouldn't count them out. They're at least willing to attempt to compete on price.
AMD’s latest parts are actually quite close to M1/M2 in computing efficiency when clocked down to more conservative power targets.
They crank the power consumption of their desktop CPUs deep into the diminishing returns region because benchmarks sell desktop chips. You can go into the BIOS and set a considerably lower TDP limit and barely lose much performance.
Where they struggle is in idle power. The chiplet design has been great for yields but it consumes a lot of baseline power at idle. M1/M2 have extremely efficient integration and can idle at negligible power levels, which is great for laptop battery life.
At any rate, using single points to compare energy efficiency isn't a good comparison, unless either the performance or power consumption of the data points comparable. Like, the M1's little cores are 3-5x even more efficient when operating in an incomparable power class, and Apple's own marketing graphs show the M1's max efficiency is also well below its max performance [1]
Those perf/power curves are the basis of actually useful comparisons; has anyone plotted some outside of marketing materials? It might even be possible under Asahi.
[1] https://www.apple.com/newsroom/2021/10/introducing-m1-pro-an...
I can’t be bothered to chase down an actual comparison, but usually you’ll see something along those lines if you compare the benchmarks for the top tier chip with a slightly lower tier 65w equivalent.
Below 100w it’s more linear, but that might depend on undercoating and the like as well.
Every notebookcheck.net review. For example https://www.notebookcheck.net/AMD-Ryzen-7-6800U-Efficiency-R...
They also do the same to a lot more laptops they test.
Look at the multi-core results, Zen3+ comes pretty close.
Also the single thread result shows what GP said: AMD CPU drains too much power at idle.
Imagine if you're testing how energy efficient an EV and a gas car is. But you only run the test in the North pole, where the cold will make the EV at least 40% less efficient. And then you make a conclusion based solely on that data for all regions in the world. That's what using Cinebench to compare Apple Silicon and x86 chips is like.
[0] https://www.reddit.com/r/hardware/comments/pitid6/eli5_why_d...
Albeit for later releases this is less true since most customers have switched to GPUs...
It doesn't. As far as I know, everything is translated from x86 to ARM instructions - not direct ARM optimization.
Cinema4D is a niche software within a niche. Even Cinema4D users don't typically use CPU renderer. They use the GPU renderer.
The reason Cinebench became so popular is because AMD and Intel promote it heavily in their marketing to get nerds to buy high core count CPUs that they don't need.
What should I be buying to not have to ask myself that question?
idle means the computer is turned on, a Mac on idle consumes less power than an x86 on idle, they both consume ~zero in deep sleep.
We just stare at an article in a web browser. We look at a text document. We type a bit in the document. An app is doing an HTTP request. The CPU is doing nothing basically.
Once in a while it has to redraw something, do some intense processing of an image or text, but it takes seconds.
It's the 99% in idling that counts and there most laptop CPU's suck.
Even when watching a video the CPU is not (should not be) doing much as there are HW co-processors for MPEG-4 decoding built in.
It's quite embarrassing how AMD and Intel have screwed up honestly.
30 Years ago I don't think the compute power of a modern phone chip was available at any price, even in super computers.
On a tangential note, there are economists who think this increase in compute is somehow an increase in one of their measures - I don't recall which one. I disagree, because with that logic we all have trillion dollar tech in our pocket. Making a better product over time is expected, it's not some kind of increase in output.
It's an Apples to ThinkingMachine Oranges comparison but CPU-Benchmark[1] ranks the Apple A16 Bionic used in the latest iPhones, its GPU - in the "iGPU - FP32 Performance (Single-precision GFLOPS)" section - at 2000 GFlop/s.
GadgetVersus[3] reports a GeekBench score of the A16 Bionic at 279.8 GFlop/s. - SGEMM test of matrix multiplication, it seems.
AnandTech[4] was reporting the A15 architecture ARMv7 came in at 6.1 GFlops in the "GeekBench 3 - Floating Point Performance" table, SGEMM MT test result, in 2015.
[1] https://www.top500.org/lists/top500/1993/06/
[2] https://cpu-benchmark.org/cpu/apple-a16-bionic/
[3] https://gadgetversus.com/processor/apple-a15-bionic-gflops-p...
[4] https://www.anandtech.com/show/8718/the-samsung-galaxy-note-...
For many things, like neural networks, it's probably good enough.
http://compression.ru/download/articles/text/witten_1994cj_l...
> dig the life that both are scary and jamming lore whose biased top must be recoverable.
However it'll be a PR disaster if the parts have spotty availability in retailers, and will cause wild rumors to spread, from "AMD doesn't want you to have this CPU" to "AMD is going bankrupt".
Anyway, I hear some people saying you can just buy the 105W "X" variant and limit the power usage to 65W in the BIOS. Does it really give the same result, also in idle power consumption?
But the review-game (same reviewers who complain about power usage) drives AMD/Intel to get the last 5% of performance increase for 20% of power increase.
Meanwhile down at the mid-end you've got Apple chipping away at them with ultra efficient ARM chips.
Interesting to see the same fundamental issue at the nano scale...
I’m pretty sure people will suggest x86 options that fit into that description but again it looks like you have to scrape the internet for stuff like intel 1[2,3][1-9]00t or wathever the AMD alternative is.
Which I unlocked to ~300W and installed a water cooling system.
It can idle at around 5W, but a synthetic load (e. g. prime95) quickly makes it draw full 300W and then throttle a bit. Fun stuff, definitely wouldn't do it again due to enormous strain on the PSU and an unreasonably high power draw relative to performance.
If anyone knows a dataset suggesting what % of world energy usage is used by computing hardware (and what proportion of that hardware is idle vs fully utilised) I'd love to see it.
Also, less on topic, what's the environmental impact of writing your app in something programmer friendly but power inefficient? I strongly suspect some big tech companies will be suffering with this (namely anyone who is still on RoR at scale).
I know it's my PTSD, the words make sense but I really hope this is not code for "we can't get much more performance out of our architectures, so we'll focus in selling "efficiency".. "
I really hope I'm wrong, since the power consumed by this chips has gotten a bit out of hand ( not as much as the GPUs.. )
Having said that, Apple's M1 was the real one that changed things, also ARM on the server is starting to actually be a mainstream thing, so I get why AMD and Intel are sweating, I just hope they can pull it off because a World where you need to buy a +$1000 aluminum box attached to the CPU you want to get is not a better World.. or worse: Doing business with Qualcomm.
My Thinkpad X1 Nano Gen 2 with Alder Lake has only 50% battery runtime of what X1 Nano with 1160G7 accomplished. For pretty much the same performance. Power constrained to say 5 or 7W, the 1160G7 feels faster.
I hope that Lenovo can offer something so light as a X1 Nano with AMD inside. Technically it should be both possible and feasible, given that the AMD CPUs are much more efficient/performant at low power levels.
Also, IMO software needs to take a lot of responsibility for energy efficiency, we just pawn that off on the hardware vendors. I wonder what the carbon cost of javascript is, I don't think I'd want to see the results, or python for that matter.
Looking forward to even more efficient Zens, my next laptop will definitely be an AMD
Per my usual update schedule, I'm looking down the line to at least a Zen5(+?) upgrade in a few years time, so I hope they improve this kind of things in the future. However that's entirely up to OEMs deciding to make a good product, and a bit out of AMDs grasp.
“What is AMD’s next challenge?”
“Efficiency.”
“Sorry, the answer we were looking for was ‘less.’”
That’s a huge advantage
In mobile (anything with a battery really), it obviously doesn't work because you care about battery life.
In server, it doesn't work, since two of the main costs and limiting factors in data centers are power and cooling.
Having a clock that every component can agree upon means that components don't have to worry about each other anymore. Physics and information theory would suggest that removal of this centralized clock signal necessarily introduces additional latency in order to safely determine or modify system state.
What he should be doing is announce the next big thing. Which I imagine might include trendy things like tackling AI with some hardware/software stack that is energy and cost efficient and competitive with nvidia. Or a non intel architecture based chip intended for high end gaming/ar/vr type devices where energy efficiency and performance are going to matter more than compatibility with legacy PC hardware. AR is going to suck if you have to be tethered to a huge power supply or battery and carry a liquid cooling apparatus with you. This requires a different approach.
And even those things really should have been the focus for the last ten years. A slightly faster version of the thing they've been selling for the last ten years is not going to turn things around and there are only so many people still assembling PCs from parts that actually know and appreciate AMD as a brand.
2. This is a keynote at a supercomputing conference. The audience are people who do supercomputing. She's not going talk about consumer stuff.
3. She did announce new, integrated products that are in at least some niches better than what nVidia has.
4. AMD brand is flying very high on the server and supercomputing side right now. They have beat Intel for 3 generations straight now.
In general, you should read the article instead of commenting based on just headlines. Doing this just makes you look really stupid.