Anyone can validate the performance on a real PC of the era? If it's confirmed, LLMs may be one of the ways of creating highly optimizing compilers in the future.
Anyone can validate the performance on a real PC of the era? If it's confirmed, LLMs may be one of the ways of creating highly optimizing compilers in the future.
Take the irow loop in vga12.inc[1] (of course I'm going to look at the graphics code). Each iteration does push di; rep stosb; pop di; add di, ROW_BYTES (and some other stuff). Why not save the push/pop and add (ROW_BYTES - count), which could be stored in a register (dx is free here)? Just the push+pop is 15+12 cycles.
[1]: https://github.com/jggonz/os8088/blob/1f2fae44180fadaf85368c...
It's a weird world we live in now.
Many still haven't understood that dynamic compilers, and machine learning based optimisations are equally not deterministic, which is why benchmarks are hard to implement properly.
It's been a long time since I've worked on a machine I could describe as "understandable."
Aside: One of the niceties of using agentic LLMs to do optimization work is letting them do the drudgery of creating a pile of microbenchmarks for different scenarios (scalar popcount on a this-shaped vector, or SIMD on another? which wins? sometimes the answer is surprising!). I've built all sorts of bespoke tools / harnesses for forcing them to evaluate their findings empirically, and it's amazing what can be done that would take me weeks by hand.