AMD Is Currently Hiring More Linux Engineers
phoronix.com
phoronix.com
I've experienced the following issues on my laptop (Dell G5 SE):
* Crashes on suspend/resume. [1]
* Kernel warning on boot [2]
* It crashes for me without amdgpu.runpm=0 kernel config [3] (fixed?)
* Crashes when inserting and removing USB-C monitor (fixed?)
* The onboard GPU works, but the dedicated GPU crashes with vsync. (didn't test recently)
1. https://gitlab.freedesktop.org/drm/amd/-/issues/1222
On my work computer with a Ryzen 7 "Renoir" GPU the problem was not the drivers but the fact that the latest Ubuntu LTS ships with an older kernel, so having to run latest mainline.
I have a feeling this sort of situation will become more and more problematic, especially when new architectures are released. Somewhere the Linux community may have to step up for more easily having recent kernels.
The experience of running old hardware on Linux is much better than that of running new hardware.
Except for certain products at least (I run latest nvidia compute cards and the drivers just work).
- Upgrade to _the very latest_ kernels. I'm not kidding, you need to start following mainline stable builds. In Ubuntu or ubuntu-based derivatives check out the ubuntu-mainline-kernel script: https://wiki.ubuntu.com/Kernel/MainlineBuilds I'm on 5.10.4 right now but probably should uprade. You need a _minimum_ 5.4 kernel to fix some egregious instability issues, but in general keep your kernel bleeding edge updated.
- Disable MWAIT instructions. This was the final missing piece of making my machine stable. The TL;DR is that AMD put out errata recently that MWAIT is basically 100% broken on these early Ryzen processors: https://www.amd.com/system/files/TechDocs/55449_Fam_17h_M_00... You can read much more in this enormous kernel bug report: https://bugzilla.kernel.org/show_bug.cgi?id=196683 The fix is to add a "idle=nomwait" flag to your kernel startup.
It's a night and day difference with these fixes in place. I still sometimes get a touchpad lockup on wake from sleep, but it's maybe 1 out of every 50 times and kind of expected for linux. No longer do I have random kernel panic-level lockups every hour. The little laptop just flies and flies with these Ryzen processors.
Took 2 years for a BIOS update that solved the final problems allowing it to work without custom kernel perimeters.
Ubuntu 20.4 was the first version of Linux that was totally stable.
2 weeks after everything was perfect - the fan died :(
I got a RX 6900 XT in december and have not had any significant issues with the drivers. YMMV on less bleeding edge distros.
One one had, the performance really sucks.
On the other, there's no way we would've had Spotify, Slack, Teams, and all these big apps supported on Linux.
But seriously -- wine should have equal performance to running a program "natively" on windows because it essentially is running the program natively, just with different system DLLs that call back to the linux kernel instead of the NT executive.
Their implementation may be slightly slower in places, but it's not a problem inherent to Wine itself, and could always be fixed.
I know there's a stigma against running win32 apps on linux, and possibly rightly so, but there really isn't a reason why Wine couldn't be a legitimate runtime environment for linux. You can make 100% open source software that targets Wine, and never needs to use or link to proprietary software.
Wine has even been ported to architectures where there have never been native windows ports, like the ppc64le port of Wine.
(Wine/win32's binary interface is also easier to intercept and automatically translate/emulate calls for non-native architectures, which is the basis of WoW64 and x86 on ARM emulation, and in Wine land, projects like Hangover https://github.com/AndreRH/hangover )
That is not guaranteed. Windows programs and Win32 APIs are writting for and optimized to run on the the NT kernel which has different perfomance characteristics from the Linux kernel. Some example:
- Wine needs to emulate a case-insensitive filesystem on top of a case sensitive filesystem, which is less efficient thatn using a filesystem / kernel FS layer that is designed for this.
- Some locking primitives are different enough that Linux will need additional syscalls to let Wine reach the same perfomance when emulating the Windows ones [0]
[0] https://www.phoronix.com/scan.php?page=news_item&px=FUTEX2-2...
I just meant that, in general, Wine should be very close to as fast on Windows, since it can usually implement win32 DLLs without much emulation/translation needed. Excepting all the bits that do :-)
[0] https://www.collabora.com/news-and-blog/blog/2020/08/27/usin...
Electron app memory usage: 150 MB
Native app memory usage: 0 MB (because you never ship it)
1: https://twitter.com/jamonholmgren/status/1105876480930734086
If the app had only essential features, it could have functioned well on web.
Back when Russian Facebook rival vk had an unlimited offering of pirated music, their music player on the web was the best. You have a list of songs, a search bar, and play/pause.
Nowadays Spotify web client takes a second or two to register mouse click. This is just sad.
Whenever that happens, I just start doing "4 billion... 8 billion... 12 billion"
We waste soooo much horsepower.
https://news.ycombinator.com/newsguidelines.html
> Comments should get more thoughtful and substantive, not less, as a topic gets more divisive.
And yes, Gnome is still sometimes completely unresponsive...
Meh... Where are the remote jobs? I bet if you open an office in Oregon (where there's tons of Intel Open Source people) you'll find a lot of candidates. Or, you know, act like the old times of 2020 and hire remote employees?
Or just how often development hardware self-bricks and needs to be recovered.
Doing hardware work remotely is a lot slower. Possible, but slower.
Edit: Or how there is that one person who is really good at assembly and if you can walk down and ask them a question it'll save you a couple hours (days) of debugging.
And then there is that one person who is really good with an oscilloscope and while in theory you shouldn't need to have decode messages by sight from waveforms, well, this one person can and it sure is useful time to time....
Sure, but minor fixes soldered in place aren't uncommon.
My experience is with firmware for embedded and consumer electronics, maybe video cards are sufficiently complex that none of the things I saw are even possible!
Lots of issues resolving power states, clock trees, monitoring buses, docs from suppliers were always subtly wrong, power management chips never quite worked how they were supposed to, things like that.
would guess since this is AMD and low level Linux, there may be a good deal of interfacing with physical hardware, in addition to any other preferences they might have
[0] https://www.phoronix.com/scan.php?page=news_item&px=Mesa-21....
If I were AMD or Intel, I'd keep investing in HPC, but be terrified that my lead there was just one good ARM BLAS library away from evaporating.
And those are big cores. For a ton of little cores, looks less clear-cut for me because of the communication overhead. (could work out though, but won't be exactly easy).
(mandated by Arm ServerReady SR/SBSA, you are not allowed to ship Device Tree there)
(even for 32-bit Windows on Arm, back in the Windows RT times, this was the case. It's a shame that Microsoft decided to lock those down back then)
On more affordable Arm machines, Honeycomb LX2K has an official UEFI firmware (with ACPI) and NVIDIA's Jetson AGX Xavier has an official (but experimental) UEFI+ACPI option.
Those aren't complete yet in mainline when using ACPI. For the LX2K, onboard networking currently doesn't work when using mainline kernels. And for the AGX Xavier, you lose PCIe (until the quirk to support it there is merged, or the ECAM through PSCI patchset) and integrated GPU support.
When using a Windows on Arm laptop, most Qualcomm drivers only have device tree bindings, not ACPI at least at this point. As such, you might need to load a device tree at boot for a more complete experience. (you can look at repos on https://github.com/aarch64-laptops if you have one)
For SBCs, forget about almost anything not named "Raspberry Pi", for which... a third-party UEFI + ACPI firmware is available at https://rpi4-uefi.dev, which allowed the original RPi4 to be certified as SystemReady ES.
A lot of the things that actually drive making people implement UEFI + ACPI is that Windows, RHEL and VMWare ESXi require UEFI + ACPI to boot on 64-bit Arm.
E: A64FX seems like a really cool chip, but it sure is much easier to get time on AWS than Fugaku. :)
It can see a future where x86 is the modern equivalent of a “mainframe”
I want to be wrong though...
The new M1 Mac is fairly open too in that you can boot non-MacOS OSes on it. It differs from the iPhone and iPad in this way as the Mac is made for a different market niche.
The best way to fight closed devices is to not buy them regardless of their CPU ISA.
ARM is just a different ISA with somewhat superior efficiency characteristics due to fixed length easier to decode instruction encoding. That's about it.
But overall people assign too much weight to ISA, especially in case of M1 and forget other aspects of the CPU design.
X64 decoders are really complex. I've heard 5% of the die. Unlike other parts like the ALU, SSE/AVX, AES engines, etc., the decoder is almost 100% "hot" all the time. This adds significant energy overhead vs. a simpler decoder.
The other problem is that it's hard to decode lots of X64 instructions at once without this energy cost ballooning. Apparently 4X decode is kind of a magic number. Intel cranked it to 6X in Ice Lake and newer with some magic, but beyond 6X is gonna be hard.
The Apple M1 has 8X decode, meaning that it can fully saturate 8 execution units for a lot of instruction level parallelism. Nothing in ARM makes it that hard to go higher. If there's a performance advantage to it we will see 12X, 16X, or even higher decode pipelines.
Another big win for ARM is looser memory semantics, allowing for more memory write reordering and more efficient cache designs.
Overall ARM is just easier to optimize than X64. X86/X64 is only fast because vast amounts of money has been spent optimizing these designs. If the same amount were spent optimizing ARM, we'd be way ahead of where we are now.
P.S. The only area where X64 still really shines is its big complex vector units, but these massive vector instruction sets are to some extent a hack to get around the difficulty of decoding X64 instructions. Think of huge SSE/AVX instructions as big macros. I'm sure ARM will get 256-bit and maybe even 512-bit vector extensions at some point, at which point the gap will shrink a lot in this area too.
P.P.S. SMT / hyperthreading is a hack to get around the difficulty decoding lots of instruction in parallel. It's a way to keep a big parallel core busier. It comes with its own overhead though, which increases power consumption, and is a security minefield. Are there even any SMT ARM cores? It's just not as big of a win when you can add more decode width easily.
EDIT: To answer my first question, it was 32-bit fixed-length. The ISA for POWER10 additionally supports aligned prefixed 64-bit instructions. See section 1.6 in https://wiki.raptorcs.com/w/images/f/f5/PowerISA_public.v3.1...
It is kind of important for me I do not have to use closed source blobs or operating systems. And I think I'm not the only one.
Btw, I've yet to see a Graviton instance that compiles our Rust project faster than my 3090x workstation. Maybe there is, but it costs a fortune. And what Amazon does here is it just kicks out all competition who cannot build their own CPUs. Then they can ramp up the prices.
In my experience and testing Rust lang is lagging a bit with ARM support. It's not that the language doesn't have good support (it does, it's great and very easy to cross compile), it's that the ecosystem around rust is _very_ x86 focused. Like I saw a simple JSON parsing library decided it would be cool to add SSE and SIMD support and now it doesn't work well on ARM (which has a totally different set of SIMD extensions). I have a strong suspicion your rust projects depend on libs with similar issues--but be ready because once those libs support similar optimizations for ARM, it's going to be super fast.
If you're working on online services and NOT looking at or using ARM servers right now--you are in big trouble and already too late to the game.
I only know about unaligned reads and the weaker memory model, but both of them are already undefined behaviour in C/C++/Rust.
If it is for their Ryzen processors... please hire more. I am thinking about going back to Intel laptops after the pitiful support in the current Linux kernel.
Random freezes when I am actually using the laptop are the most common. Infuriating.
Scrolling a webpage in Firefox using touchpad gestures is a risky business, more than half of the freezes happen when scrolling.
The fact that hibernation and sleep also don't work is just the icing on the cake.