The server chose violence
cliffle.com
cliffle.com
This sounds good when the system is small and tight and applications are written mostly by people who designed the whole system.
But as an application developer, I’d be somewhat scared to interface with third-party code over an IPC model where the other service can at any time send back an instant death pill to my process.
I guess I just don’t trust other app developers that much. The world is full of terrible drivers and background processes written by stressed-out developers harassed by management. They’ll drop in a bunch of potentially unsuitable default REPLY_FAULTs if it means they get to go home before 8pm.
I think that's intentional because that's what Hubris is aimed at.
Swift death to deviance is a way to keep the system tight. The designed scope probably keeps it small anyway. Scopes have a way of creeping, but I don't think people will want to force tasks into Hubris that would be better on the host rather than in its embedded controllers.
The server says "that client is bad!" so the kernel kills it. The problem is really that the two didn't understand each other.
For service, think "OS interface". If you make a bogus kernel call on a monolithic kernel, it would be reasonable for the OS to kill you. Also note that when you say "process" it might be different than you think because threads all share the same address space on hubris.
The Dennis Nedry approach to counting dinosaurs in Jurassic Park.
The reason why servers can shoot back at their clients is reliability, not security. Errors are thought to originate from bugs, not from deliberate attacks. The extreme reaction of the kernel ensures that developers find them as soon as possible.
Of course, there is an overlap with security, and this can be a useful fallback measure in the event that a process tries to do something that it isn't supposed to do.
These are both correct.
Well, I mean, Hubris is general in the sense that, if you're doing an embedded system and you can deal with the constraints it has, like the latter, it can work for your projects. But it's not trying to be anything other than a good embedded OS, or to handle any project.
It’s funny to wake up and read this after falling asleep reading about algebraic effects.
If you squint the right way, this is a kernel that lets a server perform an effect that the client cannot handle.
I feel like this would make code reuse and composition much harder, but provides a much simpler execution model. Definitely the right trade off in a static embedded system. You can always just vendor and modify a task if you need to reuse it.
As an example: trying to close() invalid FD is a a non-fatal error which is very often ignored. But it is actually super dangerous, especially in multi-threaded apps: closing wrong fd will harmlessly fail most of the time, but 1% of time you'll close a logging socket or a database lock file or some unrelated IPC connection.. That's how you get unreliable software everyone hates.
However, in your example it’s the kernel that is deciding the request (message) is bad. In Hubris it is the message receiver.
This is a bit contrived, but imagine you’re receiving some stringly typed data from an external source and sending a message to a parsing task that either throws or messages you back with a list of some type t. Maybe it is returning ints and you as the client know that if something isn’t parsable as an int you want it to treat it as a ‘0’ because you’re summing the list. Somewhere else you want to call the same task, but you want strings that can’t be parsed to be treated as ‘1’ unless they can’t be parsed due to overflow (in which case you rethrow) because you’re taking the product.
In some situations it’s natural for the client to know more than the server about how to handle errors. With this nuke from orbit model, there’s some forced coupling between the client and server (mutual agreement over what causes a REPLY_FAULT).
I propose HTTP 499 “Shame on you.” A client receiving 499 (perhaps on a request that it must have originated with a specific header like “Strict: true”) must terminate, in a language-dependent manner, the task which issued the request.
It perfectly balances the “WTF... But actually, hey” that one sees in those contexts.
On Linux, sure it’s not possible to directly crash another program you’re talking to via a socket alone (ignoring bad data on the socket).
But you can absolutely kill them. Anything running as root can kill anything else. Can even reboot and bring down the whole system.
Maybe a bit harder and a bit more unusual, but at least for containers, root privileges are common. And yeah, sure, there’s a cgroup there are you’re more limited. But you get the idea.
It’s also a bit different from the (conventional?) wisdom about being “liberal in what you accept, conservative in what you emit” though that’s a bit more tied to networked systems.
Though, maybe it’s inevitable that a system has to be liberal in what they accept.
How else can you change the api slightly without breaking existing programs?
Correct. From [0]:
"Hubris is an aggressively static system. The configuration file defines the full set of tasks that may ever be running in the application. These tasks are assigned to sections of address space by the build system, and they will forever occupy those sections.
Hubris has no operations for creating or destroying tasks at runtime. Task resource requirements are determined during the build and are fixed once deployed. This takes the kernel out of the resource allocation business. Memories are the most visible resources we handle this way, but it applies to all allocatable or routable resources, including hardware interrupts and memory-mapped registers – all are explicitly wired up at compile time and cannot be changed at runtime."
This is correct, yes.
One of Einstein's famous quotes is, "...as simple as possible, but no simpler." I'm pretty sure this design violates the latter portion. I'm not interested in operating environments that can tolerate no real-world chaos, and I'm not aware of any commercially viable realms which would either. What -- push it back to the init system to keep trying again? But by what mechanism would that strategy be able to understand the fault that occurred, in order to try again better?
Anyway, kudos for purity of conviction (I guess).
For more details on the thinking here and what it looks like in practice, see (e.g.) [0] and [1].
Watchdog timers will happily kill/restart your processes that don't poke them often enough. Even in my hobby exercises I've seen I2C busses hang up often enough(and bring the whole system down!) when some protocol bit goes wrong that I think the design is actually quite inspired. As I understand it this isn't talking about known error cases(that are handled) but protocol mismatches and other things that shouldn't ever happen.
Many other comments touched on it but it's a purpose built OS, much in the same way I'm not going to build a UI in Erlang, Hubris seems well positioned for the space that it occupies.
I think the general idea is to apply this to problems which are clearly the result of an invalid program state, and therefore not reasonably recoverable. They are either caused by bugs, an attack, or corrupted hardware. In all cases you shouldn't continue, because there's something seriously wrong with the caller. If the caller continues, it could only cause more damage.
It sounds a bit like Erlang/OTP's "let it crash" philosophy. Erlang is used in quite a bunch of mission-critical hardware and is famous for its reliability, so it might not be such a huge dealbreaker in practice.
Which was based partly on ideas from Tandem Computers' NonStop / Guardian. Hardware and software were fail-fast i.e. they would work correctly or stop, so they couldn't corrupt data. If there was a problem, the whole processor / process would be stopped, and a backup took over, which seems somewhat similar to the "supervisor" tasks in hubris.
Quite a bit of a different use cases - an embedded os for microcontrollers vs large OLTP applications. They both could be considered "mission critical", at least for the people who own/make money with them.
In this case your application is one that is a little more rigorous in checking what it accepts. So it has a security benefit, but not the kind you think it does: an attacker is not set back because you destroyed their progress, it’s that you made certain invalid states that were previously possible to chain into more desirable invalid states no longer work. So an attacker will look elsewhere instead of trying to do that.
EDIT: yet another factor is sometimes you may not even have access to the system you need to troubleshoot. Being able to reason about code execution without observing it is a useful skill (and still a debugger is a useful tool).
Processes keep state to analyze abuse of various kinds, and killing a process presumably wipes its memory. Unless there’s some way to retain state across restarts?
> On Hubris, if you break a system call’s preconditions, your task is immediately destroyed with no opportunity to do anything else.
Oh, yeah. I've long thought EBADF and EINVALs (and EFAULT, I guess) should basically always be fatal.
That's a bona fide remote procedure call, isn't it?
I'm not saying that's because they're fundamentally impossible, but because they have a track record of tripping up language designers and it's good to cross check against the experiences.
Recommended languages are Java (ultimately a failure despite vast effort), and Haskell and Erlang where they work, but a lot of work of very different kinds was put in to make it work. I definitely get Erlang vibes from this piece so it's possible the preconditions for correct asynchronous exceptions are met or can be met here. But they are very subtle and have a tempestuous history of working 99.9% but it being literally impossible to get to 100%. This could be a big, big, big trap.
Your Erlang vibes are there for good reason, it's certainly an influence.
Erlang solves it by locking what things it has that can have that problem behind other execution contexts that don't get killed when the main one dies, so they can still clean up. ("Ports", in their terminology.) Haskell solves it by being a functional language and beating the collective community's head in it for several years. (Immutability helped a lot, laziness took out back.)
If that sounds impossible... hey, great! Then I just pattern matched on something that wasn't a match. If that doesn't sound impossible, then it may be worth a look around.
Synchronousness may not really matter, I've kind of thought that "asynchronous exception" is not a good name for the issue for a while, but it's what it gets called. It's really about one execution context lobbing errors/exceptions into others. Although being synchronous would avoid the worst timing issues.
Tasks in Hubris are independently compiled programs, not threads in a shared context. So I don't believe that it's an issue. You don't share locks between tasks, you create a task that holds the shared resource, and have the two tasks that want to share it talk to that task, patterns like that.
Why is Java a failure? Recent JVMs have come a long way, and GraalVM makes it somewhat comparable to Go-like languages.
I understand the historical hate and how Oracle bought it, but it really isn’t that bad of a language if you’re using modern Java.