Writing a “bare metal” operating system for Raspberry Pi 4
github.com
github.com
Usually hobbyist OS projects don't start directly on hardware. Many of them never reach the point of running on hardware.
But, you can certainly wait to start on hardware until you've got it started on qemu. Debuggability is a lot better unless you've got specialized equipment for your hardware setup
One of the biggest issues people bring up when someone asks "Why can't the OS handle garbage?" is that the OS doesn't know what is garbage and what isn't. But if you only ran languages that ran on the OSes virtual machine, then it would know the memory model. Java and Racket both have histories of languages being built on top of them, so if you don't like the base language you could create your own.
I'm not saying this would be practical or useful (other than maybe as a proof of concept) but it could be fun!
But then I don't think garbage collection can be done purely in hardware. You still have to have something resembling a runtime, and it has to be written in a lower level language.
But from what I heard, Java tried SecurityManager at the language level and it failed. So I'm not sure. Also I haven't researched capability-based security deeply, so I'm not an expert in this field.
Worth noting that a language-based OS is not necessary for capabilities -- they were invented in the context of the usual hardware memory-protection in the 1960s.
The Nerves project is perhaps a more pragmatic approach that is in production use that I really like which uses a fairly light Linux to get the VM up and running: https://www.nerves-project.org/
With Nerves you mostly write Elixir or Erlang to build your device functionality.
at this point, I think I'd prefer boolean queries like altavista so I can search the word vectors myself. maybe some meta info so I can include/exclude based on various tags and links.
I so miss those days! I don’t remember exactly when it fully became useless for exact specific queries like that, but I think it was around 2007-2009.
Actual GitHub repo for anyone looking for the files: https://github.com/rust-embedded/rust-raspberrypi-OS-tutoria...
I’ve worked with some ARM SoCs from NXP and I could have sworn that one core comes up first and the others get released from reset later, with a “bringing up secondary CPUs” message printed.
* read the CPU main ID register
* if core 0, branch to primary-core bootup code
* otherwise, go into a loop (eg "read x from known location for this core, if x is non zero branch to x, else keep looping")
* core 0 releases each secondary from the loop when it is ready -- this is when core 0 prints that "bringing up secondary CPUs" message
(There are a bunch of minor variants on this, eg waking secondaries by sending them an interrupt so they can sleep via wfi insn instead of busy looping, but the basic approach is always the same.)It is also possible to do this in hardware -- you can have an SoC with a power controller so secondaries start powered off or held in reset, and the primary core prods the power controller to start each secondary.
On 64-bit Arm the common standard is that this is all handled by the firmware (which implements a standard ABI called PSCI), and the OS code just makes SMC calls into the firmware for "power on the secondary". (The firmware does something like the above under the hood.)
Enrollment is over for it as of now, so better luck next semester!
Also, I cannot believe how straight-forward bootstrapping is on a modern ARM chip compared to x64, where things become painful quickly thanks to decades worth of backwards compatibility. Enabling the A20 line[0] is my favourite example of the nonsense you have to deal with.
Where does the rpi4 store the firmware necessary to read from the sd card where this software is (presumably) stored?
[1] https://github.com/isometimes/rpi4-osdev/blob/master/part1-b...
The hard part is the linker-script to get this working right :-) https://github.com/isometimes/rpi4-osdev/blob/master/part1-b...
Doesn't necessarily have to involve assembly!
Getting into C (or Rust) assuming the presence of some kind of system firmware (BIOS, UEFI, u-boot, coreboot, etc.) isn't too difficult in the grand scheme of things.
Not to toot my own horn much, but here's an example I did of getting into Rust on a 386EX SBC i had hanging around. I actually yanked out the BIOS chip and this replaces it. Please forgive any poor Rust practices, this was written in a hurry.
https://github.com/teknoman117/ts-3100-images/tree/master/ru...
I discovered a mind-melting bug where replacing the RTC clock chip / battery-backed RAM can erase the BIOS. This SBC uses the same flash chip for both user storage and the BIOS. The partition between the user area and the bios is stored in the CMOS ram, so if there is any junk in it, the BIOS might misidentify the flash boundaries and erase itself...
So, I wrote this to recover the boards. Bonus points were that I only had an 8 KiB EEPROM hanging around so it had to fit in 8K initially.
https://github.com/isometimes/rpi4-osdev/tree/master/part2-b...
https://raspberrypi.stackexchange.com/questions/10489/how-do...
You don't handle the CPU from the reset vector like you would in a microcontroller or system firmware environment. There is an entire loader stack that finds a boot device to read a kernel image from that exists under you.
That's not to take away from this series at all, it's just the parent comment was asking about how it was so easy to get into a kernel image written in C on an SD card without any apparent SD card or FS logic.
There are a bunch of other headers in that file, but the "start_of_setup:" label is what's invoked by the bootloader, and "calll main" transitions to C. So 32 lines of code, by my count.
The main chip of the RPi4 has a small amount of code in a built-in ROM which runs on boot. In the normal boot flow, that code loads the bootloader from the EEPROM chip, but it can also read a recovery image from the SD card. See https://www.raspberrypi.com/documentation/computers/raspberr... for details (or https://www.raspberrypi.com/documentation/computers/raspberr... for how it was on the older RPi devices).
Are there any efforts out there to create open-source firmware for this chip?
Some have attempted to reverse-engineer... This was a good read: https://blog.quarkslab.com/reverse-engineering-broadcom-wire...
If I am following along this material, then I don't need all the digressions with close enough descriptions of the tools. Like, if I am reading a home building tutorial, don't explain what a hammer is.
No one can ever learn anything if no one writes this type of stuff down. sometimes the network of knowledge and laying out the precise line of questions to answer is the best way of charting out and communicating info for learners.
So thank you for writing this!
But for real - I’ve had way more luck grokking embedded rust than all of the bare metal C examples i’ve looked at. C breeds dense bittwiddling and code that relies on inscrutable compiler behavior. There are easier ways to learn how these systems work at a bare-metal level.
Also hint: C doesn't breed bit-twiddling, writing software that actually interacts directly with hardware does.
One operation per line. A comment for every operation. Shifts that explicitly say if they are wrapping or overflowing. Rust uses ! instead of ~ but if I had my way it’d be a named function like bitwise_invert().
// curval &= ~(field_mask << shift); // original line
// pseudo-rust version
let shifted_mask = FIELD_MASK.wrapping_shl(shift); // be clear about what kind of shift we’re doing
let invered_mask = shifted_mask.bitwise_invert(); // use a fictional invert fn to avoid single-char operators.
let shifted_val = curval & inverted_mask; // new variable instead of mutating the existing one
Ideally those comments would say WHY we’re doing those ops rather than what’s notable about them - but i didn’t dig into the code enough to write explanations.And then we let the compiler crush that into an efficient lil one liner like the author of the original code did manually.
My goal was to demonstrate some basic principles to get code running on bare metal, encourage curiosity, further my own knowledge and document my findings.
I appreciate that more self-documenting code might be desirable, but to some people (me included) a large number of lines can be as off-putting as more esoteric syntax. I acknowledge, however, that it is very hard to please everyone!
I must say that I’m very comfortable in C, only because it’s where I landed up as a kid. I actually find it way less confusing than more “modern” languages (I kinda skipped OO etc.!), and I enjoy the “control” it gives. Maybe you can’t teach an old dog new tricks after all! ;-)
...which means that it's serving its purpose well. If you're a beginner and it's hard to understand, that's absolutely normal. If it's easy, you probably already knew. That's how learning is supposed to work.
If you would choose such a deliberately verbose style, especially the splitting in multiple lines is the worst, the written code would become really unreadable, as too much space would be filled with text that does not provide any information, obscuring the important parts.
Normally the name of the register, the mask constant and the shift constant have informative names that should indicate all that needs to be known about the operation done and any other symbols should occupy as less space as possible on the line of code.
It does not matter if the compiler inlines them, encapsulating the bit field operations obfuscates the code instead of making it more easily understandable.
It is not possible to make the name of the function to provide more information than the triplet register name + bit field name (the name of the shift constant) + the name of the configuration option (the name of the mask constant).
Encapsulating the bit operations into a function just makes you write exactly the same thing twice and when you are reading the code you must waste extra time to check each function definition to see whether it does the right thing.
The C code would look just like a table with the names, where the operators just provide some delimiters in the table that occupy little space.
Replacing the operators with words makes such code less readable and concatenating the named constants into function names or using them as function arguments brings no improvement.
The only possible improvement over explicit bit operations is to define the registers as structures with bit-field members and use member assignment instead of bit string operations.
Unfortunately the number of register definitions for any CPU is huge, so most programmers use headers provided by the hardware vendor, as it would be too much work to rewrite them.
For almost all processors with which I have worked, the hardware vendor has preferred to provide names for mask constants and shift constants, instead of defining the registers as structures, even if the latter would have allowed more easy to read code.
I think I see what you're arguing. That this:
reg1 &= ~(width_mask_1 << shift);
reg2 &= ~(width_mask_2 << shift);
reg3 &= ~(width_mask_3 << shift);
// etc...
is clearer than something like this: reg1 = ClearBitField(reg1, 1, shift);
reg2 = ClearBitField(reg2, 2, shift);
reg3 = ClearBitField(reg3, 3, shift);
// etc...
If that's what you're arguing, I simply don't agree. `ClearBitField` is descriptive and readable. It avoids creating all those width_mask_n constants, since you specify the width as input to the fn. You don't have to go digging into `ClearBitField` because you wrote a unit test to confirm that it does what it says on the label and handles the edge cases.On top of that, the code inside `ClearBitField` can be as verbose or as compact as you desire, because it's contained and separated from the rest of the code.
Real register names are usually very long, to indicate their purpose, so you would not want to repeat them on each line.
This can be avoided by redefining ClearBitField.
Even so, writing an extra "ClearBitField" on each line does not provide any information. It just clutters the space.
Anyone working with such code is very aware that &=~ means clear bits and |= means set bits.
When reading the table of names, the repeated function name is just a distraction that is harder to overlook than the operators.
The way to improve over that is not adding anything on the lines, but using simpler symbols by defining the registers as structures, i.e.:
register_1 . bit_field_1 = constant_name_1;
register_2 . bit_field_2 = constant_name_2;
register_3 . bit_field_3 = constant_name_3;
Unfortunately, like I have said, the hardware vendors seldom provide header files with structure definitions for the registers and rewriting the headers is a huge work.
However, if you are able to rewrite just the register definitions that you use, that would be better spent time than attempting to write functions or macros for these tasks.
However, that's because I've written a fair bit of C code, and so when my brain goes into "C mode", the symbols &, =, ~, <<, etc. all have clear and unambiguous meanings - whereas ClearBitField does not. Additionally, the pattern ~(foo << bar) is a common C idiom, so beyond the individual symbols, my brain recognizes the whole pattern so it's "semantically compressed" (easier to think about) for me. This would not be the case for a beginner.
Which style is better depends on an individual's preferences and experiences - there's no "right" answer.
This is a stellar example of one of the many reasons why code-as-text is a huge mistake - because structure and representation are conflated and coupled together. A sanely written programming language represents code as code objects, and you can configure those code objects to be displayed however you like, whether that's baz &= ~(foo << bar) or ClearBitField(baz, 1, bar).
All the extra verbosity simply obfuscates the actual intent of the code: clear all bits in field_mask (shifted to the left by some offset). That's pretty easy to see at-a-glance from the C code (some comments could make that clearer, but this is simple enough that any experienced systems programmer will know what this does without comments).
I agree that Rust embedded code is often more readable than C, but that's done by creating abstractions to manage complexity rather than just by adding more words. For instance, one could write a wrapper struct that provides a less-tedious interface than a bitfield (like `curval.set_field(false)`).
I have a preference for verbosity in code, and I know that many people don't share my preference. That's alright - there's no exact right way to write that code. But my point was C encourages you to write code that relies on knowing secrets about specific hidden behavior in your compiler. `shl` isn't more clear than `<<`, but `wrapping_shl` and `overflowing_shl` ARE more clear, because it makes us explicitly aware of behavior that `<<` doesn't surface.
As for clarity, I agree an abstraction would be best. And Rust encourages those abstractions where C discourages them. I'd still argue that the inside of that abstraction should be the verbose version, but other than the wrapping_shl that's mostly just a style/preference thing.
let multiplied = a.wrapping_mul(x);
let y = multiplied.wrapping_add(b);
The "terse" equation I can instantly recognize as a linear function, while I'd have to stare at the more verbose version it for a while to figure out what it does. In my opinion, bitwise operators work the same way: if you're working in a domain where you have to write thousands of simple bitwise operations, a bit of shorthand can make the code much more expressive. fn calc_linear() -> int {
math(x) -> y {
let * = saturating_mul
let + = wrapping_add
y = a*x+b
}
return y
}
There’s clearly a war between explicit and verbose going on here. Not sure the solution, but I firmly believe we can do better than C on this one.The three lines of rust, by comparison, are ugly, long-winded and far-and-away harder to grok.
That’s probably because I’ve been an embedded systems engineer, and if you think C is terse, you ought to try verilog… but then I wouldn’t have the temerity to suggest that something that works well for me is how everyone else should do it.
field_mask was probably constructed as ((1 << width) - 1) instead of as a manifest constant. So you can just do:
ClearBitField(input, width, shift) { return input & ~(((1 << width) - 1) << shift) }
Now you just use that everywhere you would clear a contiguous bitfield which is a pretty common operation when operating on hardware. Now all your bit-twiddling is isolated to a single well-defined generically useful function instead of repeating it a billion times.
We know this is a generically valuable operation since this is basically a C implementation of the ARMv8 bfi (b)it(f)ield (i)nsert instruction with a fixed 0 argument or in assembly:
BFI X{n}, XZR, #shift, #width