Intel Xeon Max 9480 Deep-Dive 64GB HBM2e Onboard Like a GPU or AI Accelerator
servethehome.com
servethehome.com
If my CI build time at least would go down.
Software is so slow compared to hardware. It's embarrassing that we haven't moved not even a hundredth of what hardware has the last 30 years.
Why get this?
CI is a fairly new thing. The idea of constantly doing all that compute work again and again was unfathomable not that long ago for most teams. We layer on and load in loads of ancillary processes, checks, lints, and so on, because we have the headroom. And then we reminisce about the days when we did a bi-monthly build on a "build box", forgetting how minimalist it actually was.
Not the same thing.
This never occurs voluntarily, and people will wail and thrash about even as you try to help them get out of their rut.
Builds being slow is one of my pet peeves also. Modern "best practices" are absurdly wasteful of the available computer power, but because everyone does it the same way, nobody seems to accept that it can be done differently.
A typical modern CI/CD pipeline is like a smorgasboard of worst-case scenarios for performance. Let me list just some of them:
- Everything is typically done from scratch, with minimal or no caching.
- Synchronous I/O from a single thread, often to a remote cloud-hosted replicated disk... for an ephemeral build job.
- Tens of thousands of tiny files, often smaller than the physical sector size.
- Layers upon layers of virtualisation.
- Many small HTTP downloads, from a single thread with no pipelining. Often un-cached despite being identified by stable identifiers -- and hence infinitely cacheable safely.
- Spinning up giant, complicated, multi-process workflows for trivial tasks such as file copies. (CD agent -> shell -> cp command) Bonus points for generating more kilobytes of logs than the kilobytes of files processed.
- Repeating the same work over and over (C++ header compilation).
- Generating reams of code only to shrink it again through expensive processes or just throw it away (Rust macros).
I could go on, but it's too painful...
A lot of what you describe originates from lessons learned in "classic" build environments:
- broken (or in some cases, regular) build attempts leaving files behind that confuse later build attempts (e.g. because someone forgot to do a git clean before checkout step)
- someone hot-fixing something on a build server which never got documented and/or impacted other builds, leading to long and weird debugging efforts when setting up more build servers
- its worse counterpart, someone setting up the environment on a build server and never documenting it, leading to serious issues when that person inevitably left and then something broke
- OS package upgrades breaking things (e.g. Chrome/FF upgrades and puppeteer using it), and a resulting reluctancy in upgrading build servers' software
- attackers hacking build systems because of vulnerabilities, in the worst case embedding malware into deliverables or stealing (powerful) credentials
- colliding versions of software stacks / libraries / other dependencies leading to issues when new projects are to be built on old build servers
In contrast to that, my current favourite system of running GitLab CI on AWS EKS with ephemeral runner pods is orders of magnitude better:
- every build gets its own fresh checkout of everything, so no chance of leftover build files or an attacker persisting malware without being noticed (remember, everything comes out of git)
- no SSH or other access to the k8s nodes possible
- every build gets a reproducible environment, so when something fails in a build, it's trivial to replicate locally, and all changes are documented
The only stack that routinely throws wrenches into pipeline optimization is Maven. I'd love to run, say, Sonarqube, OWASP Dependency Checker, regular unit tests and end-to-end tests in parallel in different containers, but Maven - even if you pass the entire target / */target folders through - insists on running all steps prior to the goal you attempt to run. It's not just dumb and slow, it makes the runner images large AF because they have to carry everything needed for all steps in one image, including resources like RAM and CPU.
Heh. docker.io package on Ubuntu did this recently, whereby it stopped honouring the "USER someuser" clause in Dockerfiles. Completely breaks docker builds.
No idea if it's fixed yet, we just updated our systems to not pull in docker.io 20.10.25-0ubuntu1~20.04. or newer.
One of the rules of this approach is to filter all extraneous variable input likely to disrupt the build results, especially and including artifacts from previous failed builds.
Most of the issues you’ve described are consequences of yet more issues such as not caching with the correct cache key.
I argue that all of the problems are eminently solvable. You’re arguing for leaving massive issues in the system because of… other massive issues. Only one of these two approaches to problem solving gets to a solution without massive problems.
> broken (or in some cases, regular) build attempts leaving files behind that confuse later build attempts (e.g. because someone forgot to do a git clean before checkout step)
But thanks to CI you will never fix such broken build system!.
Download gets corrupted (randomly) at some point. Corrupted file gets stuffed into the cache. Now, since everyone is using the cache, everything/everyone is broken because the cache is ‘always good’, and the key didn’t change!
So then, someone figured it out and turned off caching - and it fixed it.
So now caching is always off.
Or they just turn off caching and forget about it.
It turns out, it's all super crazy at every level. Things like Docker use incredibly slow algorithms like SHA256 and gzip by default. For example, it takes 6 seconds to gzip a 150MB binary, while zstd -fast=3 achieves the same ratio and does it in 100 milliseconds! The OCI image spec allows Zstandard compression, so this is something you can just do to save build time and container startup time. (gzip, unsurprisingly, is not a speed demon when decompressing either.) SHA256, used everywhere in the OCI ecosystem, is also glacial; a significant amount of CPU used by starting or building containers is just running this algorithm. Blake3 is 17 times faster! (Blake2b, a fast and more-trusted hash than blake3, is about 6x faster.) But unfortunately, Docker/OCI only support SHA256, so you are stuck waiting every time you build or pull a container. (On the building side, you actually have to compute layer SHA256s twice; once for the compressed data and once for the uncompressed data. I don't know what happens if you don't do this, I just filled in every field the way the standard mandated and things worked.)
This was on HN a couple years ago and was a real eye opener for me: https://jolynch.github.io/posts/use_fast_data_algorithms/
There are also things that Dockerfiles preclude, like building each layer in parallel. I don't use layers or shell commands for anything; I just put binaries into a layer. With a builder that doesn't use Dockerfiles, you can build all the layers in parallel and push some of the earlier layers while the later ones are building. (One of the reasons I wrote my own image assembler is because we produce builds for each architecture. The build machine has to run an arm64 qemu emulator so that a Dockerfile-based build can run `[` to select the right third-party binary to extract. This is crazy to me; the decision is static and unchanging, so no code needs to be run. But I know that it's designed for stuff like "FROM debian; RUN apt-get update; RUN apt-get upgrade" which is ... not needed for anything I do.)
The other thing that surprises me about pushing images is how slow a localhost->localhost container push is. I haven't looked into why because the standard registry code makes me cry, but I plan to just write a registry that stores blobs on disk, share the disk between the build environment and the k8s cluster (hostpath provisioner or whatever), and have the build system just write the artifacts it's building into that directory; thus there is no push step required. When the build is complete, the artifacts are available for k8s to "pull".
The whole thing is a work in progress, but with a week of hacking I got the cycle time down from 1 minute to about 5 seconds, and many more improvements are available. (Eventually I plan to build everything with Bazel and remote execution, but I needed the container builder piece for multi-architecture releases; Bazel will have to be invoked independently for each architecture because of its design, and then the various artifacts have to be assembled into the final image list.)
I'm a little surprised to see that 6x figure. Just going off the red bar chart at blake2.net, I wouldn't expect to see much more than a 2x difference, unless you're measuring a sub-optimal SHA256 implementation. And recent x86 CPUs have hardware acceleration for SHA256, which makes it faster than BLAKE2b. But those CPUs also have wide vector registers and lots of cores, so BLAKE3's relative advantage tends to grow even as BLAKE2b falls behind.
But in any case, yes, builds and containers tend to be great use cases for BLAKE3. You've got big files that are getting hashed over and over, and they're likely to be in cache. An expensive AWS machine can hit crazy numbers like 100 GB/s on that sort of workload, where the bottleneck ends up being memory bandwidth rather than CPU speed.
On a couple older Intel x86_64 systems without the extension, I get around 300 MiB/s hashing a ~15 MB file.
On a Ryzen 3700u laptop I get around 1250 MiB/s. And I also get around 1380 MiB/s on an an aarch64 ARM vCPU on AWS (t4g instance type).
I don’t understand this mentality. What, exactly, did you expect to get faster? If you run the same software on older hardware it’s going to be much slower. We’re just doing more because we can now.
From my perspective, things are pretty darn fast on modern hardware compared to what I was dealing with 5-10 years ago.
I had embedded systems builds that would take hours and hours on my local machine years ago. Now they’re done in tens of minutes on my machine that uses a relatively cheap consumer CPU. I can clean build a complete kernel in a couple minutes.
In my text editor I can do a RegEx search across large projects and get results nearly instantly! Having NVMe SSDs and high core count consumer CPUs makes amazing things possible.
Software is improving, too. Have you seen how fast the new Bun package manager is? I can pick from dozens of open source database options for different jobs that easily enable high performance, large scale operations that would have been unthinkable or required expensive enterprise software a decade ago (even with today’s hardware).
> Why get this?
If you really think nothing has improved, you might be experiencing a sort of hedonic adaptation: Every advancement gets internalized as the new baseline and you quickly forget how slow things were previously.
I remember the same thing happened when SSDs came out: They made an amazing improvement in desktop responsiveness over mechanical HDDs, but many people almost immediately forgot how slow HDDs were. It’s only when you go back and use a slow HDD-based desktop that you realize just how slow things were in the past.
(Make it work, make it right, make it fast)
Those things aren't free. But, I don't know the relative performance hit there compared to good old software bloat.
OS/Application have been slower way before any of those were a thing.
Are we though?
Our computers are orders of magnitude faster than they were. What new features justify consuming 100x or more CPU, RAM, network and disk space?
Is my email software doing orders of magnitude more work to render an email compared to the 90s? Does discord have that many more features compared to mIRC that it makes sense for it to take several seconds to open on my 8 core M1 laptop? For reference, mIRC was a 2mb binary and I swear it opened faster on my pentium 2 than discord takes to open on my 2023 laptop. By the standards of 1995 we all walk around with supercomputers in our pockets. But you wouldn't know it, because the best hardware in the world still can't keep pace with badly written software. As the old line from the 1990s goes, "what Andy giveth, Bill taketh away."[1] (Andy Grove was CEO at the time of Intel.)
My instinct is that as more and more engineers work "up the stack" we're collectively forgetting how to write efficient code. Or just not bothering. Why optimize your react web app when everyone will have new phones in a few years with more RAM? If the users complain, blame them for having old hardware.
I find this process deeply disrespectful to our users. Our users pay thousands of dollars for good computer hardware because they want their computer to run fast and well. But all of that capacity is instead chewed through by developers the world over trying to save a buck during development time. Every hardware upgrade our users make just becomes the new baseline for how lazy we can be.
Slow CI/CD pipelines are a completely artificial problem. There's absolutely no technical reason that they need to run so slowly.
But really, it's easy to forget the massive leaps we make every year.
That said, I'm really looking forward to the day we can embed LLMs and generative AI into video games for like world generation & dialog. I can't wait to play games where on startup, the game generates truly unique cities for me to explore filled with fictional characters. I want to wander around unique worlds and talk - using natural language - with the people who populate them.
I'm giddy with excitement over the amazing things video games can bring in the next few years. I feel like a kid again looking forward to christmas.
Sure, there's a balance to be made between cutting wood and sharpening the saw. Who do we blame when the boss-man won't allow anyone to sharpen the tools even though we're obviously wasting outrageous amounts of time? You blame the people that won't allow those investments to be made.
When you multiply that across an entire industry, add some trendy fashionable tech (that's also just fast-enough to be tolerable), and this is how we end up in the shitty circumstance you describe.
And yet I still wouldn't trade my fancy IDE and slow CI pipelines for a copy of Turbo Pascal 7, as fast as it would be!
Maybe because the stuff you do isn't bottlenecked by compute? In my case every hardware upgrade resulted in a big productivity improvement.
Better CPU: Cut my C++ build times in half (from 10 to 5 minutes if I change an important .h)
Better GPUs: Cut my AI training time by a few X, massively improving iteration times. Also allow me to run bigger models that more easily reach my target accuracy.
I have made builds in the past where you had to judge, tight enough? I thought it was unsettling. I believe my last couple builds had a more cam-lock feel to the heatsink though where it was tightened to a point where it had an obvious force-threshold stop to it.
Impressed with OpenFOAM results, as that's a typical workload for our users. However, the AMD system is basically equal.
[0] https://www.ixpug.org/images/docs/ISC23/McCalpin_SPR_BW_limi...
[1] https://macperformanceguide.com/blog/2022/20220618_1934-Appl...
is a good punchline from the report. The Apple cores have that problem but a lot worse. They are slow. It goes back to this flawed idea in the community that you can "just" "add" "more memory," when parts like the H100 have their memory size matched to the architecture (physical and software) of the dozens of CPUs on them.
I'm not sure why the conception persists that a 60W laptop part would be comparable to a 300W server part of the same process generations, let alone this particular part.
The only thing you can say about the larger chips is they have much wider channels, but they are still just regular DDRs.
if this pricing system is making an inefficient market for cpus then maybe it could be disrupted somehow
Your working set of data won't spill out from the L3 to "the L4" when it grows too large.
64GB of HBM2e as ram is more performant than 128GB of DDR5 with HBM2e cache, and often the cached variant has no speedup compared to a standard Intel configuration.
Also OpenFOAM loves cache: https://www.phoronix.com/benchmark/result/amd_ryzen_7_5800x3...
IMHO those results don't contradict the what I said. Of course if the workload entirely fits in the 64GB HBM, there's no point in using it as a cache, just use it directly. But if you need to address more RAM(any big DB, fs, etc.), and you don't want to manage the tier manually, then the caching mode could shine.
It's a benefit in 2 socket configurations when the memory the CPU needs is connected to another socket. The data will be cached on the other socket's HBM for the entire system without filling up that other socket's L1-L3 cache for data it hasn't requested yet.
Another way to see it is the last time Intel messed around with L4 cache under the EDRAM section. https://www.anandtech.com/show/9582/intel-skylake-mobile-des...
The entire GPU and EDRAM complex get moved out of the CPU caching scheme.
For those confused by the title, Intel released a Xeon that includes 64GB of high speed RAM on the chip itself, configurable as either primary or pooled memory, or a memory subsystem caching layer.