HNHacker News
TopNewBestAskShowJobs

afr0ck

438 karma · joined August 9, 2019

Karim Manaouil

k.manaouil@gmail.com

submissionscomments
afr0ck··on Building a Linux GPU Driver for the M4 Mac Mini in One Month
And this kiddo is probably barely 22yo
afr0ck··on Salesforce Global Outage
Classic
afr0ck··on The Qualcomm Oryon CPU is the first mobile CPU to reach 5GHz
The video editing stuff was about the FlexCache. Cores can pull each others cache. This is not novel btw, it's used by recent IBM CPUs.
afr0ck··on AI;DR (AI; Didn't Read)
I think the main reason many people (including me), very often, lack the motivation to read content that is likely generated by AI is the suspicion that it comes from a place of intellectual laziness. Another reason, based on personal experience, is that AI content may suffer from too much verbosity, too much jargon and over-confidence, which makes the reading experience feel fake and border-line irritating. In many cases the content may have very little to no nuance, which is ultimately a waste of time. As an anecdote, someone posted a blogpost on Linkedin on using agents to implement a driver to access PCIe devices over TCP/IP. I was intrigued because that's not an easy task for several reasons, like handling PCIe interrupts and DMA. For exmaple, how does the remote machine map the device's PCIe BARs? And when it issues I/O to the devices registers, how are these reads and writes transferred to the remote device. In the end, this is just some virtual memory. In a local machine, this is either directly mapped to the PCIe physical addresses or some IOMMU virtual address space which is then translated by the hardware upon CPU/device/VM access.

After reading the long verbose promising article, in the end, the guy (with the help of the agent) only managed to implement access to the PCIe config space so that lspci on the remote machine works and shows the remote PCIe device, but that's all. It never addressed the issues above nor even mentioned them. The code was AI generated. The article was AI-written. The article never made a reference to DMA, interrupts, MSIX-X, IOMMU, IOTLB, virtual memory, etc, but it made big claims on next-gen datacenter disaggregated architecture, boosting GPU utilization, reducing large scale inference costs, etc.

Anyway, you get my point: big long beautiful words, but zero nuance.

afr0ck··on AI;DR (AI; Didn't Read)
I think the main reason many people (including me), very often, lack the motivation to read content that is likely generated by AI is the suspicion that it comes from a place of intellectual laziness. Another reason, based on personal experience, is that AI content may suffer from too much verbosity, too much jargon and over-confidence, which makes the reading experience feel fake and border-line irritating. In many cases the content may have very little to no nuance, which is ultimately a waste of time.

As an anecdote, someone posted a blogpost on Linkedin on using agents to implement a driver to access PCIe devices over TCP/IP. I was intrigued because that's not an easy task for several reasons, like handling PCIe interrupts and DMA. For exmaple, how does the remote machine map the device's PCIe BARs? And when it issues I/O to the devices registers, how are these reads and writes transferred to the remote device. In the end, this is just some virtual memory. In a local machine, this is either directly mapped to the PCIe physical addresses or some IOMMU virtual address space which is then translated by the hardware upon CPU/device/VM access.

After reading the long verbose promising article, in the end, the guy (with the help of the agent) only managed to implement access to the PCIe config space so that lspci on the remote machine works and shows the remote PCIe device, but that's all. It never addressed the issues above nor even mentioned them. The code was AI generated. The article was AI-written. The article never made a reference to DMA, interrupts, MSIX-X, IOMMU, IOTLB, virtual memory, etc, but it made big claims on next-gen datacenter disaggregated architecture, boosting GPU utilization, reducing large scale inference costs, etc.

Anyway, you get my point: big long beautiful words, but zero nuance.

afr0ck··on Qualcomm to Acquire Modular
Why you say that? Nuvia made a massively great success with Oryon CPUs which are now all over the place.
afr0ck··on Oracle shed about 20k roles globally in the last year
It's also an opportunity, from a different perspective
afr0ck··on PgDog is funded and coming to a database near you
Is this vibe-coded?
afr0ck··on Did the Linux memory management maintainer "just quit"?
David Hildenbrand, another very involved memory management legend is picking up the role. It will be fine.
afr0ck··on The next two years of software engineering
I created my first Linux from scratch when I was a freshman in college in a third world country (not India). Fast forward few years later, I now write Linux kernel code for a living. Not sure what you did wrong, bud, to end up miserable like this.
afr0ck··on Same-day upstream Linux support for Snapdragon 8 Elite Gen 5
That's not how operating systems work. KVM is both an interface and a hypervisor. Just as we have different hypervisor implementations for amd, intel, arm and others all abstracted behind the same KVM interface, there is no reason the same can't be done for Gunyah. Userspace does not have to know anything about that. KVM already supports svm and vmx for amd and intel on x86. Why is something similar can't be done for Arm? Plus now there is pKVM.

I just don't understand this argument of a separate interface. The only reason you want to do that is to decouple from the KVM community, but that introduces a shit tone of duplicated effort and needless fragmentation to the virtualisation software ecosystem hindering your users from enjoying the existing upstream tools they already know about. In other terms, vendor locking and shitty downstream experience.

afr0ck··on Same-day upstream Linux support for Snapdragon 8 Elite Gen 5
I worked at Linaro, who was contracting for Qualcomm. Qualcomm were pushing for some protected hypervisor called Gunyah (which had its own Linux interface and needed a new qemu port) that apparently no one liked. I tried to port it to KVM [1], but upstream folks (mostly Google) outright rejected the port. Otherwise KVM would have been available on QCOM boards. You can still try it. I have a Linux kernel and a Qemu port on my github [2,3]

[1] https://lore.kernel.org/kvm/20250424141341.841734-1-karim.ma...

[2] https://github.com/karim-manaouil/linux-next/tree/gunyah-kvm

[3] https://github.com/karim-manaouil/qemu-for-gunyah

afr0ck··on Linux Career Opportunities in 2025: Skills in High Demand
Linux kernel + bootloaders + firmware

The Linux kernel side is mostly device trees, device drivers and the like.

u-boot is very famous as a bootloader in the embedded space

Firmware for board bring up and devices

afr0ck··on Living my best Sun Microsystems ecosystem life in 2025
There are Qualcomm laptops now I believe (at least that's what I heard when I was last working for them). NXP also made some boxes (I own a bunch of them). The server market is also growing with Ampere and Cavium (now Novell) which I have both.
afr0ck··on AMD's EPYC 9355P: Inside a 32 Core Zen 5 Server Chip
NUMA is only useful if you have multiple sockets, because then you have several I/O dies and you want your workload 1) to be closer to the I/O device and 2) avoid crossing the socket interconnect. Within the same socket, all CPUs shared the same I/O die, thus uniform latency.
afr0ck··on GigaByte CXL memory expansion card with up to 512GB DRAM
I think Meta has already rolled out some CXL hardware for memory tiering. Marvell, Samsung, Xconn and many others have built various memory chips and switching hardware up to CXL 3.0. All recent Intel and AMD CPUs support CXL.
afr0ck··on GigaByte CXL memory expansion card with up to 512GB DRAM
CXL uses the PCIe physical layer, so you just need to buy hardware that understands the protocol, namely the CPU and the expansion boards. AMD Genoa (e.g. EPYC 9004) supports CXL 1.1 as well as Intel Saphire Rapids and all subsequent models do. For CXL memory expansion boards, you can get from Samsung or Marvell. I got a 128 GB model from Samsung with 25 GB/s read throughput.
afr0ck··on Without the futex, it's futile
It's not that deep. The futex was developed just to save you from issuing a special system call to ask the OS to put you on a wait queue.

The whole point is that implementing a mutex requires doing things that only the privileged OS kernel can do (e.g. efficiently blocking/unblocking processes). Therefore, for systems like Linux, it made sense to combine the features for a fast implementation.

afr0ck··on What does Palantir actually do?
>right-wing religious revival rooted in Christianity, combined with technological acceleration and a reimagined political order that prioritizes heroic individuals and hierarchical impulses

Sounds like a fast path to totalitarianism a la 1930.

afr0ck··on Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?
Inference runs like a stateless web server. If you have 50K or 100K machines, each with a tons of GPUs (usually 8 GPUs per node), then you end up with a massive GPU infrastructure that can run hundreds of thousands, if not millions, of inference instances. They use something like Kubernetes on top for scheduling, scaling and spinning up instances as needed.

For storage, they also have massive amount of hard disks and SSD behind planet scale object file systems (like AWS's S3 or Tectonic at Meta or MinIO in prem) all connected by massive amount of switches and routers of varying capacity.

So in the end, it's just the good old Cloud, but also with GPUs.

Btw, OpenAI's infrastructure is provided and managed by Microsoft Azure.

And, yes, all of this requires billions of dollars to build and operate.

afr0ck··on UK to buy F-35As that can't be refueled from RAF tankers
France has its own independent military production including jet fighters (Rafale), tanks, ballistic missile, nuclear submarines and nuclear heads.
afr0ck··on Making TRAMP go Brrrr
I use vim with mutagen for syncing files. It's simple and works fine, but you have to duplicate storage.
afr0ck··on Karol Herbst Steps Down as Nouveau Maintainer Due to Linux's Toxic Environment
All of this drama is coming from the Rust email thread. Please, explain to me how this is irrelevant? I have nothing against Rust in the kernel, but Rust people are definitely dramatic.
afr0ck··on Resigning as Asahi Linux project lead
This is not how the kernel works. You cannot rely on someone's "commitment" or "promise". Kernel maintainers was to have very good control over the kernel and they want strong separation of concern. As long as this is not delivered, it will be very hard to accept the Rust changes.
afr0ck··on Several Russian developers lose kernel maintainership status
> "some Russian-sounding names are banned, but we still have to demonstrate there is a due process".

That's not true! There are still many Russian maintainers in the kernel, but they are not based in Russia. They only banned individuals, based in Russia, who are employed by sanctioned companies.

afr0ck··on Intel and AMD form advisory group to reshape x86 ISA
Yes, when support is not upstream.
afr0ck··on Asterinas: OS kernel written in Rust and providing Linux-compatible ABI
Linux kernel is not complex. Most of the code runs lock-free. For example, the slab allocator in the kernel uses only a single double_cmpxhg instruction to allocate an object via kmalloc(). The algorithm scales to any number of CPUs and has NUMA awareness. Basically, the most concurrent, lowest allocation latency allocator you can get in the market, which also returns the best objects for the requesting process on big memory systems.

The complexity on the other hand is architectural and logical to achieve scale to hundreds of CPUs, maximise bandwidth and reduce latency as much as possible.

Any normal Rust kernel will either have issues scaling on multi-cores or use tax-heavy synchronisation primitives. The kernel RCU and lock-free algorithm took a long time to be discovered and become mature and optimised aggressively to cater for the complex modern computer architectures of out-of-order execution, pipelining, complex memory hierarchies (especially when it comes to caching) and NUMA.

afr0ck··on Intel and AMD form advisory group to reshape x86 ISA
What you wrote doesn't make any sense. Arm has DTB [1]. Most SoCs re-use a lot of hardware IP blocks and they require very little modifications to DTB files in the kernel and device drivers to get them working. PCIe and USB support discoverability so no issue from that side.

Arm ecosystem is cleaner in my experience and learned from the mistakes of the past. Arm CPUs are still not as fast as high-end x86 chips, but it's just a matter of time before that market is also eaten by Arm.

[1] https://community.arm.com/oss-platforms/w/docs/525/device-tr...

afr0ck··on AMD's Turin: 5th Gen EPYC Launched
NUMA gives you more bandwidth at the expense of higher latency (if not managed properly).
afr0ck··on AMD's Turin: 5th Gen EPYC Launched
So, essentially, you're just doing cache eviction in software. That's obviously a lot of overhead, but at least it gives you eviction control. However, there is very little to do when it comes to cache eviction. The algorithms are all well known and there is little innovation in that space. So baking that into the hardware is always better, for now.
Page 1 of 6Next →