KolibriOS – operating system written entirely in assembly language
kolibrios.org
kolibrios.org
Generally, higher level languages are usually built upon non-zero-cost abstraction overheads, library bindings, memory management and runtimes.
Just like in a HLL, if you really need it you can still write reusable data structure libraries, but Asm's really tiny constant factor and effort involved in adding complexity forces you to think about whether you really need it first.
In other words: in the amount of cycles spent initialising a hashmap and inserting a few dozen or hundred items (perhaps involving memory allocations, etc.) just so you can get (once again, a relatively large factor each time) asymptotically constant-time lookup in an HLL, you could've gone through the whole set many times already in a tight loop of less than a dozen instructions, that entirely fits in the L1 cache along with the data too.
...except the additional complexity of doing so? If you have to write every single instruction, you start thinking more about whether you have to write each one.
'Constant factor' is entirely irrelevant to this discussion.
It's entirely relevant to the real world.
C compilers will convert constant divides into reciprocal multiplication. All the handwritten assembly I have seen uses the divide instruction.
You could try to "decompile"(!) this project into a higher level language and compile the result if you were really curious and wanted to try this exercise in the opposite direction. I suspect a lot of the things done in this code aren't even representable in a HLL or something a C compiler could be coerced to generate (without cheating and using inline Asm.)
Your claim about compilers using "SIMD tricks" to "get close in speed" makes no sense either. If the HLL is using SIMD instructions and this magic assembler that you're talking about isn't, how would their performance ever be equivalent?
Would someone writing assembly tend to avoid using external libraries unlike with HLL?
That decision is entirely up to the developer and the project requirements. If you're not writing a bare-metal program there's no reason why not to use external libraries. It's still perfectly possible.
You can absolutely not write a program much larger than 256 bytes in the style you use to write those. It would be utterly incomprehensible and unmaintainable. The small size actually works in your favour here, and allows you to use incredibly questionable tricks.
I remember back in the day where I was following a small amateur indie game scene (1998-2002), a guy released a game he'd written entirely in assembler, just for fun. He could have easily have written it in C instead, but he didn't.
The binary was significantly smaller and the game really smooth.
You don't write assembler the same way you write C or C++.
A compiler is working under various constraints, for instance regarding interoperability. When you craft everything by hand, you really have no constraints other than your imagination. You can benchmark stuff and learn and adapt. Compared to that a compiler is a sophisticated idiot.
That doesn't mean all of us should start working in assembler. But don't look down at people that do - instead find inspiration, and perhaps try to make the compilers less dumb.
See figure 7.
how would their performance ever be equivalent?
The majority of general-purpose code can't use SIMD because it's really branchy, mundane "business logic" type of code, and that's where handwritten Asm's code density really shows an advantage.
I see that you claim to have "written lots of bare-metal assembly" and have posted a link to your OS in Asm, so I looked at the code...
You are not using Asm the way Asm is supposed to be written. You are writing code like a compiler, which totally misses the point of using Asm. Now it is obvious why you don't understand --- because you've never seen what "real Asm" looks like.
To elaborate, one thing that stands out is lots of stack manipulation, barely using the registers at all. Putting everything on the stack is what stupid compilers do. This is no good for speed nor size.
I have been reading and writing Asm for a few decades. See some of my other comments if you'd like to learn more... here's a quick sampling:
https://news.ycombinator.com/item?id=8248172
https://news.ycombinator.com/item?id=12332672
https://news.ycombinator.com/item?id=12356450
Case in point: I saw a web forum once, written entirely in assembly language. The author claimed high performance, but looking at the sources, I saw that it uses no buffering, and even a simple webpage results in thousands of write syscalls. Whatever they saved in more efficient function prologues, they lost back in wasted context switches.
Edit: I have a little bit of experience in the area of operating-system development in assembly. This is purely a hobby project that I work on half-heartedly. Nevertheless, it does demonstrate what I know about bare-metal x86 assembler.
Some of the other low level systems (like the Mac) had a trap system that wasn’t so far away in cycle count from user code. But in these days of needing 10,000 cycles to bridge a system call it’s best to do whatever you can to avoid calling the OS.
Abstractions like functions are still present, just not in the way that you might think of them in a higher-level language. You use 'branching' instructions to jump from one part of the code to another. Either by referencing labeled sections of the code to perform 'absolute' jumps, or by jumping 'relative' to the current instruction. This forms the building blocks that can be used to implement more abstract constructs such as loops, functions and conditionals. This is a gross simplification, I hope it helps answer your question though.
You can write tests like in any language, they just happen to be more tedious.
If you make use of powerful macro assemblers, like MASM and TASM, you can write what looks like high level functions.
You can have a look here for a couple of MASM examples,
Optimised ASM can still be tested but one problem is that when the code doesn't work, it can simply crash. This makes it a pain to debug.
You might as well ask why dogs bark or puzzle with furrowed brow over the croaking of frogs.
Then I sort of turned the problem around in my mind.
When is a compiler prevented from optimizing? Maybe pointer aliasing?
Another thing I wondered. I've written hello.c and it's about 84 bytes of C and 8.3k as an executable. Would hand-coded assembler be that large?
maybe it's that compilers CAN do well, but because of requirements, they can't do some things well and necessarily create a lot of "boilerplate infrastructure".
The answer is generally no. A lot of that overhead is due to the C std library. If memory serves from a blog post I read long ago, with very very aggressive tuning, you can get a hello world binary down under 50 bytes (which is smaller than the ELF header). A “normal” ASM coded “hello world” binary could easily be in the range of a few hundred bytes without anything special.
hello.c: 4 loc, 84 bytes
hello.s: 30 loc, 517 bytes
hello: 8.3k
so it must be data structures for linking.
oh, and if I use -static, hello becomes 844704 bytes. :)
I just compiled my own 'hello.c' on Linux, with the optimisation flag set for 'code size', and the binary weighed in around 8.3k. I then removed the `printf` call and the inclusion of `stdio.h`. The resulting binary was 8.1k. This is exactly as expected. Why would the inclusion of a call to a dynamic library bloat the final binary size?
Here’s two good links to read on it:
https://www.muppetlabs.com/~breadbox/software/tiny/teensy.ht...
Nice job anyhow...
BSF (Bit Scan Forward) is the most famous which returns an undefined result if your operand is 0.
Many instructions leave the state of bits in the FLAGs register undefined. E.g. AAA will leave SF,ZF, and PF undefined.
Then of course there is the “reserved bits” but those are easier to avoid.
http://visual6502.org/wiki/index.php?title=6502_Opcode_8B_%2...
Some is inherent to the ISA itself.
And some undefined behavior is errata that differs from CPU to CPU. Check out the errata documentation for any processor released in the last few decades. For example, x86. You can trigger things that happen on one x86 CPU that won't happen on another.
Some things, one should just accept.
IoT devices with memory measured in single digits KB.
Plus, someone needs to write the Assembly that those compilers generate.
So I never had to deep dive into stuff like writing AVX by hand, just basing my remark on some comments that occasionally read about.
dav1d, the open source AV1 decoder, has now more asm than C code. It's one of the most recent open source projects with significant asm work ongoing.
The asm version outperforms the C version (full optimizations enabled) by 4,5x on AVX2.
We have similar results in SSSE3 (3,5x) and ARM64 (4x).
We're not talking about a few percents, we're talking about multiple times faster.
And AV1 is a standard, so there are no algo shortcuts that can be made: it is either compliant or it is not.
Sure, in most cases, it is not needed to write asm; but there are cases, notably for multimedia, where writing asm by hand is a lot faster, and that includes codecs and game engines.
I'm not sure that operating systems are so suited to ASM optimization (in the sense that you may not reap so many benefits). Maybe one could optimize for size (so you can make super tiny OS) ?
Maybe one could optimize for size (so you can make super tiny OS) ?
Yes, and remember... optimizing for code size is optimizing for speed. =)Smaller code and data = more CPU cache hits, which are orders of magnitude faster than fetching from RAM. So even if "all" you do is make code smaller, you can get more speed...
Not always. I think loop-unrolling is a common perf-optimisation technique. Even if the perf-gain is positive in this case, I highly suspect that it doesn't outweigh the cost of maintaining asm code(vs C/other higher level code).
edit: formatting
That's what created multi-kilobyte memcpy() implementations, which barely beat REP MOVSB but cause huge icache bloat.
Now we have CPUs with 256KB instruction cache and up. You can compile your whole OS at insane optimizations.
In Java, that might involve adding more classes and abstractions to make things conceptually simpler. In C, that might involve using structs to keep relevant data together. In C++, it might involve using a std::set to keep a list of 3 constants, Etc.
It turns out that when you have to write the asm by hand and you see that double pointer dereference is a pain, you avoid data structures with double dereferences (as C data structures often end up with). You avoid classes and abstractions (as are common in Java), because they involve a lot of boilerplate. You use compile time constants for that set of three things rather than any complex hashing that C++ would do, etc.
There are lots of small things like this, and it turns out the effect adds up significantly.
I strongly disagree :-)
My experience is that optimizing assembler leads to reorganize your code or even your algorithm into a form that you'll CPU will be most efficient at executing which almost always translate to unbelievably intricate, super hard to modify code. I've done stuff on 6502, 80x86, MMX and the assembly parts, optimized for speed and ended up with impossibly tricky code. Worst, sometimes I have to adapt my data structures to allow optimization, which becomes even tougher. So, I prefer to leave assembly code at the "onyl if necessary" level
But now, to be honest, optimizing assembly code is incredibly satisfying to me :-) I've got the feeling to use a CPU to its maximum capacity. Also, using assembly comes after thorough algorithmic study. So once I'm at the ASM level, I've maxed out my own capabilities ! How happy me !
So if you have the chance to do that, just give it a try !
I'd assume that when you maintain larger codebases (i.e. full OS and application suite), that you start writing practical and maintainable code instead.
But even so, I'd say that ASM doesn't help. Clumsy code in ASM is worse than clumsy code in high-level language :-)
It's not subliminal. It's packing a backpack that someone else has to carry, vs packing a backpack that _you're_ going to carry.
Until the day the compiler can smartly trim down all the unity framework to just 512 bytes because it notices I'm not using most of it, hand coding in asm will always work out smaller and faster, if one puts in enough effort.
If you compile hello world in C, it's not going to be very big either.
> if one puts in enough effort
"enough effort" is a cheat. I can replace almost any program with a smaller javascript version, if I put in "enough effort". It's not a reasonable way to compare anything.
You should actually try that, and compare sizes with the asm hello world.
Hello world in C becomes anywhere from 120KB to 1.2MB even if just using puts. Even setting the entry point to a function with no arguments doesn't help.
A tiny hello world ends up needing to ignore the standard library and use a console output function from the OS.
If you go to an embedded point of view where there is nothing already supplied, you can get a hello world, even with a basic printf, down to that size.
It is an order of magnitude bigger than the assembly version.
It prints (via write syscall) "hello world.\n" and then calls exit(0).
Assembled with fasm, linked with ld -n, then stripped.
The difference is between just using asm, or going out of the way to try to coerce a C toolchain to minimize the bloat.
When I tell it to emit the assembly with "-S" I get 28 lines of code when I have a feeling assembly is much shorter.
I do not know, to be honest.
I was just pointing that multimedia and game engines can get a lot of helps on ASM by hand.
There are some SIMD constructs that you can efficiently express in C by writing intrinsics, but that isn't substantially different from writing ASM.
Then there are some trivial cases where a loop can be unrolled and packed into SIMD instructions automagically, which retains readability, but greatly limits what you can write. You'll need to read the generated machine code to make sure you didn't mess up.
The benefit of just writing ASM is to not have such a translation layer between you and the processor. That same code that you carefully wrote to make it through the optimization passes for one compiler will likely get messed up in another compiler.
Frankly, no.
Improving the compiler could get a few dozen of %, which is huge already. But 300%+, no.
In fact, any scatter-gather operations were non-existent in Halide.
From what I remember these operations were introduced at some point, but we moved away from Halide.
Also, it was not quite simple to transform a loop that draws sprites over the entire image into a loop that draw sprites over part of image and draws many parts of image in parallel (change nesting). Hand-written CUDA version of the algorithm ended up with exactly that.
Thus, if you need some partial-derivatives-numerical-kernel, Halide is good for you. If you are working on the video decoding, Halide is not that good for you. If you are working on video encoding, Halide will be more of a nuisance than a helping hand (early exits from loops, computable access ranges, etc).
Do you think it would have worked to organize tiles and threads outside of halide and use halide for isolated parts that are already organized into arrays?
> transform a series of angle-amplitude pairs (radioastronomy) to the image
On the suggestion in your second part: why use Halide then? Should it be responsibility of Halide to work out the best loop nesting and best use of threads?
Again, Halide was put aside and we used CUDA for final version, exactly because of inability of Halide to do good work in our case.
I don't know about "should", but it seems to me that it would still be valuable, even if working out the threading and organization.of the data into an array.
> Again, Halide was put aside and we used CUDA for final version, exactly because of inability of Halide to do good work in our case
I didn't say anything about that. I'm not sure why you are restating it.
I would expect a much smaller gap as AVX2 is also available through intrinsics on C ?!
That's trivial in ASM I believe, but don't know how well it's possible in C/C++, if at all.
https://software.intel.com/sites/landingpage/IntrinsicsGuide...
You can also try to make a union whose one field is a __m64, but beware not to trigger UB conditions.
One can, for example, write:
typedef int v4si __attribute__ ((vector_size (16)));
v4si a, b, c;
[…]
b += 3; // add 3 to each of b’s elements
c = a + b; // pairwise addition
Using this gives up some control over assembly; it doesn’t guarantee vector instructions get used, but it makes it easier for the compiler to generate them.Also, I’m fairly sure there are vector instructions not covered by this extension.
The answer and parent were in no way only talking about OS.
Be curious to know what you're thinking of here.
I wonder how long it will take for compilers close that gap. I assume compilers will eventually produce assembly code that outperforms anything written by human.
[1] https://www.muppetlabs.com/~breadbox/software/tiny/teensy.ht...
Whatever, the Kolibri ISO is 64 MB. Nuff said.
Likely due to the inclusion of source code and some non-core stuff actually written in c.
https://old.reddit.com/r/programming/comments/dygsvm/heavyth...
I was skeptical, so I looked into his zlib implementation. It was 18% faster than gzip v1.6. The author said, "all I did was hand-compile the reference implementation, an 18% reduction in user-space time is a big advantage IMO."
I am not skeptical any more! More details of the benchmarking are in the linked thread.
// uppercase without if
*c -= (*c-'a'<26U)<<5;
Adding loop unrolling to that C-optimized version via asm provided another 20% performance boost. And I didn't even add vectorization using AVX.While compilers provide a lot of optimization directly, there are huge performance gains in simple functionality by helping the compiler with easier to optimize code or by using asm directly.
Also notably the highest GCC optimization level I tested (-O3) reverted the optimization of the optimized C code, while -O2 kept them, resulting in -O2 being much faster. So only using asm guaranteed the performance gain.
Oh yeah and I'd just like to mention as an aside that assembler statements make more sense to me than bootstrap css codes - a sad statement to the art of reasonable brevity.
For really tight algorithms I've still seen humans beat compilers by a lot, though it takes a lot of skill and some domain knowledge of the CPU.
For code that benefits a lot from vectorization, human coders still beat the crap out of compilers. I am not aware of any vectorizing compiler or JIT that can even approach what a human coder can do with SSE, AVX, or NEON (ARM's equivalent). I think these CPU extensions are just too complex for current generation compilers to effectively deal with. They require too much abstract understanding of what's actually happening in the code and the CPU to use really effectively.
I'm a bit surprised that there's been so little attention paid to the opportunities for applying deep learning and other advanced techniques to compiler optimization. It seems like this area is ripe for a new wave of innovation, but it's not happening. I do believe that many companies would pay for an advanced compiler capable of generating code that was significantly faster than stock compilers, but the speedup would have to be more than a few percent to justify spending money on it.
"If you really want guest addons for KolibriOS to be available, you may offer your help in porting them" http://wiki.kolibrios.org/wiki/Setting_up_VirtualBox
Stated differently: is this a useful thing, or is it an exercise in "how far can we push a pure-assembly project"? The latter would of course be fine, but I'm quite curious to know which it is.
;)
The website also claims that the operating system has a "word processor." A more appropriate term for it would be "text editor," it appears to be less functional than Notepad on Windows.
Everything's fast in the early stages, when there's not so much of it yet.
At the time the whole system booted off a floppy disk in a split second (despite the slow read rates and seek time of floppy disks!)
Menuetos abandoned its 32bit version (the one that was open) to focus on 64bit, which is not, and even has a clause in the license that prohibits disassembly.
I'd say RPI it's the best in that regard. Also fully fledged Linux there, not some pet project OS.
I'm not sure if it's actually more expensive to find VIA et al based systems for cheaper than NUCs?
But there's Vortex86.
Not much educational use to get from a system that is closed source and has a clause forbidding disassembly in the license.
Regardless of being closed source, I'm going to guess that it would still be highly educational to learn Assembly programming on a system that makes it a first class citizen where you just boot into the OS that already is entirely built around ASM. For really simple stuff, a simple microcontroller might be better, but I'm sure some folks are more interested in building desktop apps.
Educational use most often means you learn how to use a software, not how it's made.
Autodesk gives free educational licenses for its software and I saw no one saying it's of no use.
Seriously, if there's a platform to which a so tightly packed OS could be immensely beneficial, are all those small Linux-capable ARM boards that cost like two beers. Even the most hardware limited ones (256 MB RAM, 32bit single core, etc.) that still can run a complete Linux environment in some contexts would literally scream when served something that fast and tight.
From https://www.raspberrypi.org/forums/viewtopic.php?f=55&t=2209... :
What is RISC OS?
RISC OS is a new and different OS for the Pi. It isn't Linux, it isn't Unix, it isn't based on any other OS. It's the first ARM OS, begun in 1987 by the team who designed the original ARM processor. It's also a descendant of the OS used on the early 1980s BBC Micro... those who remember the BBC Micro might find some of its commands familiar. BBC BASIC is only a few keypresses away.
What's interesting about RISC OS?
It's small. It's fast. RISC OS is a full desktop OS, where the core system including windowing system and a few apps fits inside 6MB. It was developed at a time when the fastest desktop computer was an 8MHz ARM2 with 512KB of RAM. That means it's fast and responsive on modern hardware. The memory taken by apps is usually counted in the kilobytes. A 700MHz 256MB Raspberry Pi is luxury - what to do with all that memory?
RISC OS is also a lot simpler than modern OSes such as Linux. The pace of development has been a little slower than other OSes, which means there are fewer layers getting between you and the system. It's much easier to get stuck in and change things. It's also easier to understand. As a formerly closed-source OS, most of the interfaces are documented in a series of books called the Programmers' Reference Manuals (PRMs) which are included on the RISC OS Pi distro. That means you can change a lot of things without having to work on the OS code itself (which is available if you want it). It's very modular, so you can mix and match components, and the communications between modules are carefully documented.
RISC OS gets out of the way. It's a 'co-operatively multi-tasked' OS. While that means one misbehaving application can stall the system until you kill it, it also means you easily write apps that take over the whole machine - for example controlling hardware where you need predictable timing. RISC OS is a single-user OS, which means there's very little security - not great for internet banking, but very handy when you want to dig around and program the internals of the OS.
As a full desktop OS, there's also plenty of traditional desktop software available like drawing programs and desktop publishers. Features you've come to expect on a desktop like scalable fonts and printing are supported. RISC OS was big in UK education in the 1990s, and there's a large back catalogue of educational software.
:)
But there is an OS being written in Rust. https://www.redox-os.org/
If this could run Bash, rclone, and whatever Ethernet hardware I have in my old laptops I think it would be really useful.
I wonder how much effort would be needed to port the GNU Core Utilities? Probably well outside my comfort zone.
Still it should be worth a look!
I haven't followed Menuet for a few years. Not sure if now they have GCC or not. If not, probably you won't see your favourite apps run on it.
Hmm...
“Have you ever dreamed of a system that boots in less than 10 seconds from power-on to working GUI, on $100 PC?”
Too show it's possible, to have some fun. I think they're perfectly aware that this won't replace their current browser-host.