My Fear of Commitment to the First CPU Core
jabperf.com
jabperf.com
Ever since, when I do any sort of CPU or core pinning or grouping, I leave the first core or two alone if possible. It's convenient to say the first two cores are for all the ancillary stuff, but the last 30 are for our workload.
Now I realize why I never witnessed something like this in all benchmarks I performed... I always ignore the first CPU core.
Love your user name too! Plover! Zorton!
Very Linux. It's complex, very complex, and full of footguns like this.
The real answer is to use seL4[0]. Among its extensive formal proofs, it has proof of WCET (Worst Case Execution Time), therefore unlike Linux and other hopelessly non-deterministic unorthogonal toy OSs, seL4 can actually guarantee latency bounds.
There's... A LOT... of work to do before seL4 is going to be anywhere near usability parity with something like Linux, unfortunately.
Rather than make a general purpose OS, we decided to use it more like a unikernel or "library OS" where you're trying to make a well defined kind of "appliance" image to deploy to specific hardware rather than try to fake being a POSIX-y shaped OS. At some point we plan to leverage the WASM ABI to help with stability across versions & toolchains and stuff a WASM runtime in there too, so that support for "applications" written in something other than Rust is easier. Though, it's a pretty low priority at the moment since it's not at all core to our business presently.
the GP was talking about guaranteed low latency execution. You're orbiting a different star here.
As is proposing seL4 as replacement for Linux in general compute...
"Throw away 95% of your workflow and rebuild everything from scratch" is not good advice
I was meaning to circle back and see how difficult it is to throw together uni-kernel type work with it, but didn't see much in the way of documentation for how to ... use it.
Is the general idea here to target use of sel4+Rust for embedded work? Or is it feasible to think about this stuff in the context of server workloads on e.g. x86_64 machines?
I have a fantasy of writing a database engine that sits close to baremetal, completely managing its own memory for page buffers and so not engaging in negotiations with the Linux virtual memory system at all.
But such an application would need a robust networking and block device I/O stack. Something I'm not sure I can expect in sel4 land.
Our current plan, whenever we can get to it, is to implement a SQLite VFS layer (they did a really nice job abstracting away the database from the storage) to be able to host SQLite in a FerrOS task so that your local filesystem is a SQLite database.
Just email me (nathan@auxon.io) or post some issues on GH if you want to connect with folks who work on FerrOS. Heck, depending on what happens with some of our SBIRs we'll probably end up hiring a couple people to work on it full-time to do the stuff I mentioned.
I love the idea of seL4, and was sorely disappointed to see Australia cut their funding. Decades of exploits have me convinced that microkernels are the only real path towards security.
Not necessarily. Check out the CHERI project, which has a reasonably good handle on exploits https://www.cl.cam.ac.uk/research/security/ctsrd/cheri/
Well, seL4 is now managed under a foundation as part of the Linux Foundation umbrella. The seL4 Foundation. Ultimately, I think it might be for the best. Now there are more seL4 researchers and practitioners out in industry working to integrate and build upon seL4 as a building block that never would have existed because they'd previously all been living under the same roof.
An horrible story, covered in detail in Gernot's blog[0].
Fortunately, seL4 foundation has picked up steam, and is no longer dependent in any way on the leeches that almost killed it.
Some HFT systems are written in Java - they just give it loads of RAM, disable the garbage collector and restart it once a day.
They profiled the memory leak and gave the chip enough RAM that it wouldn't fill up before the missile exploded. Problem solved!
The thought was that you would rather overheat and perform the calculation than protect the circuit and lose the ship. It's kind of a hail mary strategy. It definitely gave me a visceral feeling for how one should be ready to re-frame their systems analyses, even if apocryphal.
If the mock data is good, you can do the whole thing in testing and fail tests if new code introduces GC events. However the Java code you have to write end up being very C like.
Though perhaps one could write a program in a way that gets reliably jitted?
In this context, Linux is fairly sensible as it has decent knobs for getting out of the way and is well supported for all of your non-fast-path-code.
You should team up with the Hurd folks. They have, what, 30 years of the same impotently shaking their fists at the rest of the world touting how their perfect little academic exercise and its tens of users is the future of computing if everyone else on the planet wasn't so st00pid. Really. Any day now.
They really have the angry old CS professor schtick down. You could learn a lot.
I use every chance to drive people away from the Hurd.
That project is a waste of time. It has failed to move away from Mach or address the Hurd critique document. It has no direction and thus no future.
Instead, I try to promote efforts that merit the attention, such as seL4 or Genode.
The talk is ostensibly about generating graphs, but it also describes a patch to make Linux prefer the previously-used core for any task, to make massively parallel tasks not jump around and end up with two threads on the same core.
I doubt OS's ever purposely utilize the second core lol.