Frame – Linux X server in Assembly
isene.org
isene.org
If someone had written this program manually, the strategy would have been very different. With a good macro-assembler (and nasm is good enough) one should define a great number of macros, to encapsulate all the tedious boilerplate, especially for things like function prologues, epilogues and invocations.
With a well written macro library, an assembly program can be almost as compact as a C program, instead of containing many text lines for each equivalent high-level language statement.
Such an assembly source with good macros can be read and understood much more easily than raw assembly language, like in this "frame.asm".
Otherwise, this is interesting work.
The other day, colleague showed me a (pretty basic) terminal emulator written in one-shot by Opus. Kicker is - that was compiled to a 30 KB static binary. That's right. No libX11, no libXfont, not even libc.
static inline long read(int fd, void *buf, long count) {
register long rax asm("rax") = __NR_read;
register long rdi asm("rdi") = (long)fd;
register long rsi asm("rsi") = (long)buf;
register long rdx asm("rdx") = count;
asm volatile(
"syscall"
: "+r"(rax)
: "r"(rdi), "r"(rsi), "r"(rdx)
: "rcx", "r11", "memory"
);
return rax;
}In my memory the syscall ABI has changed a few times (i386 had int $0x80, then sysenter, then abstracting it in the vdso, then amd64 has 'syscall'), so it may be easier to let the kernel header provide the mechanism.
My terminal emulator using the same binding also started out hand-written but Claude overhauled that too recently and it knows vtxx escape codes far better than me too.
On a long enough time frame with enough tokens invested there’s probably not a difference, but being written in assembly by an LLM doesn’t imply optimal to me. I’d almost prefer having an LLM rely on higher level abstractions offered by a programming language rather than rolling everything itself. After reviewing a lot of LLM code, even at Fable and Sol levels, I just don’t trust that LLMs are writing optimal code. Assembly makes it harder to even review.
I do it find it very fun and entertaining. This is a component and I’m grateful that it was shared.
Writing maintainable assembly is at odds with writing fast assembly in most circumstances.
A key optimization that's hard to pull off is inlining.
An optimizing compiler can see that a method is small enough that it can be pulled into the caller, it can then further eliminate from that smaller method branches that can't be executed due to the nature of the caller (Imagine calling a function with a `bool` parameter and you send in `true` at the call site).
To make the code faster in hand written assembly, you have to do the inlining, but that makes writing more structured code a lot harder. You are duplicating logic paths in the name of performance.
Not to mention the fact that the compiler gets updated and knows about more instructions and architectures then you do or then you could have. Hard to write the FMA instruction if it didn't exist when you were writing the assembly in the first place.
https://news.ycombinator.com/item?id=8508923 https://news.ycombinator.com/item?id=36618344 https://news.ycombinator.com/item?id=41922295 https://news.ycombinator.com/item?id=44186843 https://news.ycombinator.com/item?id=44177446 https://news.ycombinator.com/item?id=36949314
It's true that in specific functions a human can do better than a compiler at optimizing for a target platform (sometimes, not always). That's a place where someone could reasonably drop down to assembly and get better performance.
But the case I'm specifically pointing out, one of inlining, is something that compilers do better than humans. This isn't "propaganda from compiler companies" (btw, not a thing. The most common and popular compilers are opensource and not owned by any single company). There are other cases like this where compilers are just more likely to get things right than humans are. They have a lot of heuristics about common assembly patterns that few humans can be expected to have memorized.
Each of the cases you pointed out are cases where the compiler does a bad job at optimizing a single function for whatever reason. They are not examples of a whole application written in assembly outperforming compiled high level languages. And each of the cases almost certainly took the human a considerable amount of time to figure out and prove their solution was better than the compilers.
What you've done is cherry pick when compilers fail and you are using that as evidence that they always fail.
Compilers sometimes fail, something I'm happy to admit. But on the whole for a whole application the compilers will get a lot more right than a human possibly could because they can output unmaintainable assembly.
That entire point is moot here because the abstraction in this case is a massive ball of crap and it's used everywhere. So you never get the benefit you think you would for using a lower level language. That's why generally LLMs are best with python and even better with a "harness" (domain specific framework and language)
Modern macro assemblera are fun
I see the growing trend of words losing all meaning is still going strong in 2026. I wonder what human communication will look like in the near future?
Note that I do think reading is superior. This is not to diminish anyone who chooses to listen to books. Some people do it because of accessibility, some because of time (can listen while commuting, or in the gym), or they just choose to listen to some books they don't care as much whilst still read the ones they do.
This is all great and people should be encouraged to do what works for them - but please don't pretend it's the same. Sometimes I even think we need a different word for reading e-books.
Also, has anyone run it successfully? I got as far as building and running with --display and then running `DISPLAY=:7 dwm` and `DISPLAY=:7 alacritty`, but I can't seem to focus the window to actually type. Given that the author posted a picture of the thing actually running a live environment and claims to actually be using it, I'm pretty sure this is a me problem but I haven't been able to figure out where it is. Mouse works, too.
by claude code. So this was only possible since no human had to bear looking at X original source code.
For some of the trickier ones it might be worth diving into an X11 implementation, but you can also defer most of the trickier ones other than getting the event handling right (most of the old school drawing API mostly matters if you want to run 30+ year old software that hasn't been updated much since; that said it might be nice to be able to run twm and xeyes - I have a whole separate file in my own X11 server for legacy stuff required to run twm and xeyes and similar ancient software, but not useful for anything else I actually care to run...)
Working: tile, dwm, pcmanfm-qt, feh, dmenu, glass
Not: alacritty, st
So... Thus far, anything but a terminal other than glass. I'd think the problem was alacritty doing fancy things with APIs that frame doesn't have, but st??
EDIT: Testing with xtruss shows st uses RenderCompositeGlyphs8(), part of the RENDER extension. I don't have Alacritty installed, but I think Alacritty uses the GPU, at least by default? Looking at the source for Glass it seems to use PolyText16 - the "old school" old server-side font rendering API. A lot of older X11 apps would work fine with PolyText16, and a lot of newer X11 apps does all the text rendering client side into a shared memory buffer without requiring GPU support, so it's not hard to end up with a set of applications where none of them would run into either gap.
EDIT2: I've looked at the Frame source, and it does seem to have support for the RenderCompositeGlyph calls and other supporting request, so not sure what the issue with st is. It's not doing anything unconventional.
I'd try rxvt (PolyText8/16) or xterm (ImageText8/16). If either/both works it's likely an issue with the RenderCompositeGlyphs support.
Isene and I have relatively similar philosophies on this, except I have Claude burning tokens on optimizing and fixing my Ruby compiler now because I still want things in a high-level language, and my entire stack is Ruby instead of asm. But I love what he's doing - I just don't love x86 asm...
Turns out a functioning X server is a relatively simple piece of software. It's mostly just tedious. And most of the bulk is protocol handling that Claude can handle really trivially.
There are other tedious bits too, like all of the details around exactly how to propagate which events, but the protocol encoding is certainly one of the most tedious bits.
I've had all my side projects being written in x64 for the last 6 months and it is shockingly effective.
The question is, if one takes their time and has enough competence to do that.
Have you ever compiled something by hand? You should try sometime, it's an illuminating experience. Humans find it hard because you have to remember a lot of details while simultaneously paying attention to a different large set of information while also generating instructions. It's tough, but not impossible, it takes humans a lot of time and effort. How might a computer fare if it could remember everything and pay attention to multiple inputs and outputs at once? That's what an LLM does.
How can a generic LLM generate better assembly than a dedicated compiler, whose sole purpose is to generate assembly code. With people pedantically adding every optimization imaginable and unimaginable to produce the most efficient code possible. And you have the audacity to say LLMs, which write garbage non-trivial amount of time, are capable of producing better assembly.
This has got to be either a masterful ragebait, or a person with very low knowledge of modern compilers, because even an LLM would not write something so stupid as this.
LLMs generating "assembly that runs a great deal more efficiently" is a ludicrous claim that cannot be substantiated outside PEBCAK situations.
(Oh, and the "compiler" will also refuse to generate certain types of programs.)
How often there are stories "Google/Apple/Microsoft [but mostly Google] deactivate my account for some stupid reason and I lost everything: my business, my tax data, my client data, my developer account and access to my apps, everything".
With LLM/Coding agent providers it will be worse!
I suspect (with zero proof or understanding) that this has something to do with how well C maps to assembly. It's not a stretch to say the model's vector space maps this chunk of assembly with that line of C. And we all know how much C code exists online.
It's also far better than me (as someone who has done assembler since the Commodore 64) at using gdb to debug it, despite being effectively stuck using it in batch mode (which I didn't even knew existed). Watching it write elaborate scripts to dig into a code generation bug in my compiler is something.
I feel like the problem used to be that it'd struggle with the ambiguity of flow that is much more apparent in a high level language. But clearly that's not a problem any more.
I couldn't once get any of the SoTA models from a month or two ago to correctly execute more than the first 5% of the instructions for a fizzbuzz (compiled from C with GCC). As I recall, one of the Qwens did the best and would only mess up "a little bit", but that's of course enough to derail everything (can someone remind me again why we think natural language is good for interacting with precise machines?). I didn't think it'd go very well, but failing at decoding something as well-documented as RISCV is not very impressive!
Most models would also start gaslighting me when I pointed out their mistakes. To their credit, they'd very often cheat by deducing that the code was for fizzbuzz, and try to fake the execution. Always badly though. (This despite explicit instructions to execute the code faithfully instruction by instruction and not be informed by their overview of the code).
I honestly don't understand how people can work like that. I had fun because the whole thing was a joke, an art project. Doing serious work in that way must be so ridiculous.
But then again, I don't use LLMs very much and might be holding them wrong.
Sortof like suckless.org, but vibecoded on the fly.
I've found his Rust android development framework quite interesting, and will probably try using that to vibecode some software I've always wanted to have for my own personal use, but I've never had the time to code properly, especially considering it won't be useful for anyone else. Of course, I'll probably release it under a FOSS license, in case somebody finds it useful anyways.
Or:
$ objdump --disassemble /usr/lib/xorg/Xorg
Maybe I'm getting too old for this, but I really don't see the point in having AI generate assembly code for this.Everything that can be done with LLMs is better than without. Why use tooling at all when you just can use a LLM? Maybe just make the LLM generate machine code for your tager directly.
Just like making digital systems that are "better" than analog ones from the past because of digital.
The fact that you have chosen to say this about generated amd64 assembly is telling. This is a terribly pointless exercise. Further, the llms are bad at more niche languages especially. Even experimenting with having them write C, which is quite a bit more common than writing straight assembly, I have seen them fall over quite severely.
No dependencies and better performance? Fantastic.
A Dockbar I am working on https://github.com/edumucelli/docking/ is a pain to build on Wayland. It already supports a lot of composers and even mutter/gnome with that Gnome shell extension, but at what cost ...
You can always use XWayland if you feel that way.
So you can't really "fix it", short of patching e. g. Gnome's compositor, with patches that will never be accepted upstream.
And so much of it really was needless drama. Go read through many of the wayland protocol extension issue tracker threads. The amount of intentional heel dragging, willful ignorance, and purity spiraling is off the charts (IMO obviously).
To be clear I like and use wayland in its current state. There are plenty of valid criticisms of X11. IMO a redo was not a bad idea per se and it eventually turned out well enough (provided your app doesn't depend on the 5% that suffers an arbitrary rejection).
I've never quite found that Linux is more optimized on battery-powered machines for energy savings, even though supposedly there is a lot of room to tweak and optimize settings -- from selecting a low resource window manager/DE to turning off various services to switching up power management utilities. But this does seem like an approach that might produce that kind of fruit?
Many distros already try to push good defaults, but you can do a whole lot when optimizing for a mobile experience. You can also do some fun stuff with it, like running a script[1] when going from ac->bat power to, e.g., turn of a service, lower refresh rate or reduce brightness.
[0]: https://linrunner.de/tlp/index.html [1]: https://linrunner.de/tlp/usage/run-on.html#run-on-ac-run-on-...
For example, I recently got another 1 hour out of my old laptop's battery because I didn't realize for the intel video card driver I needed to add some modprobe flags to get it to load up a firmware binary blob. Doing that enabled hardware video decoding, faster performance, and lower power usage.
There's a bunch of setting like this that you need to make sure are turned on to get the best battery performance. Some OSes are better about toggling them than others and mine (gentoo) let's you discover later that you forgot to turn them on :).
The machine I'm typing on is the 2nd newest in the fleet -- it's a work box -- and it's an i7-8550U, an 8th gen "Kaby Lake" chip.
Also disable hardware SMT as a kernel option in grub config. Then the cores can clock down way more often and L1 data cache size doubles.
XFCE and X11 tripled my laptop battery life vs. whatever Wayland+GNOME Ubuntu (2024?) brought by itself.
Powertop and tlp also help.
Happy camping.
EDIT: the lower heat dissipation also halved boot time. That one surpised me the most.
EDIT2: Disable "atime" with ext4 option "noatime". Saves a lot of power, heat, trimming, and re-writes on your SSD/NVMe.
For "faster shutdowns" manually run systemctl start fstrim.service. Not exactly sure why fstrim.timer seems unreliable.
COBOL and 4 GL dreams coming into reality.
AI wrote*
recently i also rewrite most of the app's underlying core function to rust, just like the guy do for the phone
perhaps i should also do more stuffs given codex reset too quickly
Very inspiring and I applaud your efforts.
I haven't written assembler in years but this inspires me to do a small project for the fun of it.
Working on it:
https://github.com/X11Libre/xserver
And they also deleted old code too. A lot of the old code could probably be removed, but is it really that relevant whether you have 4 millions line of code or 2 million lines of code? C is in my opinion too overbose. Rust is even worse. Which language would yield fewer lines of code without speed penalty? C is king largely because of the speed gains. We don't see people use python for an xserver.
For example I had Codex port Kiwi Cassowary C++ to Nim: https://github.com/elcritch/kiwiberry
I wish mine had no fan too except me.