HNHacker News
TopNewBestAskShowJobs

mycoliza

269 karma · joined December 20, 2017

submissionscomments
mycoliza··on Oxide Computer raises $445M (SEC Form D)
there are even images of it on our website!
mycoliza··on Futurelock: A subtle risk in async Rust
In reply to your edit, that section in the RFD includes a link to the full example in the Rust playground. You’ll note that it does not make any use of ‘select!`: https://play.rust-lang.org/?version=stable&mode=debug&editio...

Perhaps the full example should have been reproduced in the RFD for clarity…

mycoliza··on Futurelock: A subtle risk in async Rust
An analogous problem is equally possible with streams: https://rfd.shared.oxide.computer/rfd/0609#_how_you_can_hit_...
mycoliza··on Futurelock: A subtle risk in async Rust
> No, just have select!() on a bunch of owned Futures return the futures that weren't selected instead of dropping them. Then you don't lose state.

How does that prevent this kind of deadlock? If the owned future has acquired a mutex, and you return that future from the select so that it might be polled again, and the user assigns it to a variable, then the future that has acquired the mutex but has not completed is still not dropped. This is basically the same as polling an `&mut future`, but with more steps.

mycoliza··on Futurelock: A subtle risk in async Rust
Indeed, you are correct (and hi Matthias!). After we got to the bottom of this deadlock, my coworkers and I had one of our characteristic "how could we have prevented this?" conversations, and reached the somewhat sad conclusion that actually, there was basically nothing we could easily blame for this. All the Tokio primitives involved were working precisely as they were supposed to. The only thing that would have prevented this without completely re-designing Rust's async from the ground up would be to ban the use of `&mut future`s in `select!`...but that eliminates a lot of correct code, too. Not being able to do that would make it pretty hard to express a lot of things that many applications might reasonably want to express, as you described. I discussed this a bit in this comment[1] as well.

On the other hand, it also wasn't our coworker who had written the code where we found the bug who was to blame, either. It wasn't a case of sloppy programming; he had done everything correctly and put the pieces together the way you were supposed to. All the pieces worked as they were supposed to, and his code seemed to be using them correctly, but the interaction of these pieces resulted in a deadlock that it would have been very difficult for him to anticipate.

So, our conclusion was, wow, this just kind of sucks. Not an indictment of async Rust as a whole, but an unfortunate emergent behavior arising from an interaction of individually well-designed pieces. Just something you gotta watch out for, I guess. And that's pretty sad to have to admit.

[1] https://news.ycombinator.com/item?id=45776868

mycoliza··on Futurelock: A subtle risk in async Rust
Yeah, a coworker coming from Go asked a similar question about why Rust doesn't have something like the Go runtime's deadlock detector. Your comment is quite similar to the explanation I gave him.

Go, unlike Rust, does not really have a notion of intra-task concurrency; goroutines are the fundamental unit of concurrency and parallelism. So, the Go runtime can reason about dependencies between goroutines quite easily, since goroutines are the things which it is responsible for scheduling. The fact that channels are a language construct, rather than a library construct implemented in the language, is necessary for this too. In (async) Rust, on the other hand, tasks are the fundamental unit of parallelism, but not of concurrency; concurrency emerges from the composition of `Future`s, and a single task is a state machine which may execute any number of futures concurrently (but not in parallel), by polling them until they cannot proceed without waiting and then moving on to poll another future until it cannot proceed without waiting. But critically, this is not what the task scheduler sees; it interacts with these tasks as a single top-level `Future`, and is not able to look inside at the nested futures they are composed of.

This specific failure mode can actually only happen when multiple futures are polled concurrently but not in parallel within a single Tokio task. So, there is actually no way for the Tokio scheduler to have insight into this problem. You could imagine a deadlock detector in the Tokio runtime that operates on the task level, but it actually could never detect this problem, because when these operations execute in parallel, it actually cannot occur. In fact, one of the suggestions for how to avoid this issue is to select over spawned tasks rather than futures within the same task.

mycoliza··on Futurelock: A subtle risk in async Rust
As a member of (Eliza, Sean, John, and Dave), I can second that debugging this was certainly an adventure. I'm not going to go as far as to say that we had fun, since...you can't have a heroic narrative without real struggle. But it was certainly rewarding to be in the room for that "a-ha!" moment, in which all the pieces really did begin to fit together very quickly. It was like the climax of a detective story --- and it was particularly well-scripted the way each of us contributed a little piece of the puzzle.
mycoliza··on Futurelock: A subtle risk in async Rust
We actually don't use Rust async in the embedded parts of our system. This is largely because our firmware is based on a multi-tasking microkernel operating system, Hubris[1], and we can express concurrency at the level of the OS scheduler. Although our service processors are single-core systems, we can still rely on the OS to schedule multiple threads of execution.

Rust async is, however, very useful in single-core embedded systems that don't have an operating system with preemptive multitasking, where one thread of execution is all you ever get. It's nice to have a way to express that you might be doing multiple things concurrently in an event-driven way without having to have an OS to manage preemptive multitasking.

[1] https://hubris.oxide.computer/

mycoliza··on Futurelock: A subtle risk in async Rust
What could `tokio::select!` do differently here to prevent bugs like this?

In the case of `select!`, it is a direct consequence of the ability to poll a `&mut` reference to a future in a `select!` arm, where the future is not dropped should another future win the "race" of the select. This is not really a choice Tokio made when designing `select!`, but is instead due to the existence of implementations of `Future` for `&mut T: Future + Unpin`[1] and `Pin<T: Future>`[2] in the standard library.

Tokio's `select!` macro cannot easily stop the user from doing this, and, furthermore, the fact that you can do this is useful --- there are many legitimate reasons you might want to continue polling a future if another branch of the select completes first. It's desirable to be able to express the idea that we want to continually poll drive one asynchronous operation to completion while periodically checking if some other thing has happened and taking action based on that, and then continue driving forward the ongoing operation. That was precisely what the code in which we found the bug was doing, and it is a pretty reasonable thing to want to do; a version of the `select!` macro which disallows that would limit its usefulness. The issue arises specifically from the fact that the `&mut future` has been polled to a state in which it has acquired, but not released, a shared lock or lock-like resource, and then another arm of the `select!` completes first and the body of that branch runs async code that also awaits that shared resource.

If you can think of an API change which Tokio could make that would solve this problem, I'd love to hear it. But, having spent some time trying to think of one myself, I'm not sure how it would be done without limiting the ability to express code that one might reasonably want to be able to write, and without making fundamental changes to the design of Rust async as a whole.

[1] https://doc.rust-lang.org/stable/std/future/trait.Future.htm... [2]: https://doc.rust-lang.org/stable/std/future/trait.Future.htm...

mycoliza··on Oxide’s compensation model: how is it going?
As an Oxide employee, let's just say...it does, and it is :)
mycoliza··on How oxide cuts data center power consumption in half
We also sell computers... :)
mycoliza··on Oxide Cuts Data Center Power Consumption in Half
The big piece of copper is fed by redundant rectifiers. Each power shelf has six independent rectifiers which are 5+1 redundant if the rack is fully loaded with compute sleds, or 3+3 redundant if the rack is half-populated. Customers who want more redundancy can also have a second power shelf with six more rectifiers.
mycoliza··on How oxide cuts data center power consumption in half
We write the security mitigations. We patch the CVEs. Oxide employs many, perhaps most, of the currently active illumos maintainers --- although I don't work on the illumos kernel personally, I talk to those folks every day.

A big part of what we're offering our customers is the promise that there's one vendor who's responsible for everything in the rack. We want to be the responsible party for all the software we ship, whether it's firmware, the host operating system, the hypervisor, and everything else. Arguably, the promise that there's one vendor you can yell at for everything is a more important differentiator for us than any particular technical aspect of our hardware or software.

mycoliza··on Who killed the network switch? A Hubris Bug Story
We did it as a result of my suggestion, and I'll freely admit that it does seem pretty goofy!

The motivation, though, was that we wanted to write tests for the function that would run on a developer's machine using `cargo test`, and the Hubris kernel currently only compiles for the targets that we actually run Hubris on (various Cortex-M targets). So, if we want to write tests for that function, we would either need to move it to a crate that doesn't contain Cortex-M-only code, or litter the whole kernel with `#[cfg(not(test))]` attributes so that most of it doesn't compile when building tests. This felt much less unpleasant than conditional compilation. We're hoping that, eventually, other complex-but-not-architecture-specific kernel code will end up in the `kerncore` crate as well so that we can write tests for it, too, so eventually it won't be a crate with only one function...

I do think that there's room for tooling improvement to make writing host-platform tests for `#![no_std]` Rust that gets cross-compiled to another platform, but I don't have any particularly concrete ideas for what that would look like. For now, at least, putting it in a separate crate lets us have our tests --- and those tests let us ensure that some of the function's edge cases are handled correctly, so I do think it was worth it.

mycoliza··on Minimal, allocation-free OpenMetrics implementation for no-std/embedded Rust
in my use case (a hobby project; https://github.com/hawkw/eclss), the device and Prometheus are on the same LAN and i have Prometheus set up to discover scrape targets using multicast DNS. this is…probably not what you’d do if you were shipping a consumer project, i think, but it’s a nice setup for a home network.
mycoliza··on Minimal, allocation-free OpenMetrics implementation for no-std/embedded Rust
hi, i wrote this thing!

if you’re wondering why it’s weird, half-finished, or under-documented, it’s because i wrote it in a couple hours to scratch my own itch, and really didn’t expect to be at the top of hackernews today! if this is something that other people are actually interested in using, i’d be happy to clean it up a bit and add some of the missing stuff…

mycoliza··on Tokio Console
A lot of the Tokio console UI was inspired by `htop`, which provides a pretty similar overview of processes and threads. It doesn't really have the same ability to inspect things like `pthread_mutex` and timerfds etc in the same way that the Tokio console can inspect the state of `tokio::sync::Mutex` and `tokio::time::Sleep`, though; although I wonder if something like that could be possible with eBPF...
mycoliza··on Tokio Console
It's https://typeof.net/Iosevka/ (I took the screenshots).
mycoliza··on Tokio Console
Yes, we've designed the overall architecture of the system to be modular so that the telemetry can be consumed by a number of different UIs --- we'd love to see someone write web interfaces and/or native GUIs for the console data. I have basically no web development experience whatsoever, though, so I went with the terminal app, because not having to learn JavaScript first made it a lot easier to get started :)

We're also thinking about factoring out the Tokio Console command-line application's internal data model and client code into its own library (https://github.com/tokio-rs/console/issues/227) to make it easier to build other UIs on top of that.

mycoliza··on Tokio Console
At my day job (https://buoyant.io/), we're using it to write a reverse proxy/load balancer --- one of the use-cases where you really, absolutely do need asynchronous concurrency. :)
mycoliza··on Tokio Console
Go has pprof (https://github.com/google/pprof), which I've heard good things about --- and, the pprof data model was one of the influences I looked at when designing the Tokio console's wire format. But, I'm not sure if pprof has any similar UIs to the one we've implemented for the Tokio console; and I haven't actually used it all that much.
mycoliza··on Tokio Console
Yeah, currently, the console knows how to detect TrueColor, but in this case, I just used the ANSI 256 palette rather than picking a better one when TrueColor is available...we should probably fix that!

Side note, it turns out that detecting what color palettes a terminal supports 24-bit colors is surprisingly fraught. There are a couple env variables that may be set...but not every terminal emulator will set them. And then you can use `tput`...but the terminal may not have correct data in the tput database. So that was fun to learn about!

mycoliza··on Tokio Console
Thank you, that's really nice to hear!
mycoliza··on Tokio Console
Something like that would definitely be useful! It's not really in scope for this project, which is intended as a telemetry and diagnostics tool, but I can imagine a Rust REPL being useful. Of course, in order to do that, you'd need to implement a general-purpose Rust interpreter, which seems like a fairly large amount of work.

In practice, I personally just use REPLs mostly for quick testing out of a small expression or something...and honestly, I usually just use the Rust playground (https://play.rust-lang.org/) for this. Small examples are compiled fast enough in the playground that it's kind of a REPL-like experience for testing stuff out semi-interactively...but it's not the same as connecting to a running application and running new code inside of that application. That's something that seems very difficult to add to Rust, a compiled, statically-linked language with limited support for hot reloading...

mycoliza··on Tokio Console
Tokio is a non-profit, community-supported project, although many of us work on it as part of our day jobs. If you want to support Tokio development, you can contribute to it on GitHub Sponsors (https://github.com/sponsors/tokio-rs) and on OpenCollective (https://opencollective.com/tokio).

I've also recently started accepting donations on my personal GitHub Sponsors page (https://github.com/sponsors/hawkw) if you're interested in supporting my open-source work in particular.

mycoliza··on Tokio Console
Yeah, the goal is for `valuable` to replace `tracing`'s (currently much more limited) `Value` trait entirely, when we release `tracing` 0.2. Before making a breaking change, though, we want to release opt-in support for `valuable` in the current v0.1.x `tracing` ecosystem, so people can start trying it out and we can figure out if there's anything missing.

You can follow the progress of that here: https://github.com/tokio-rs/tracing/pull/1608

I believe it's currently just waiting for a crates.io release of `valuable`!

mycoliza··on Tokio Console
(primary author of the console here) you're right that there are too many digits of precision right now...the reason for that is that it's actually _not_ a thoughtful design at all, though I appreciate you saying that it is; I just picked an arbitrary number when I was writing the format string and didn't really think about it. We probably don't want to display that much precision --- for smaller units, we probably don't want any fractional digits, for larger units like seconds, we probably want two digits of precision maximum.

Regarding the color scheme, glad you like the idea. Because it's a terminal application, the choice of the colors was constrained a bit by the ANSI 256 color palette (https://www.ditig.com/256-colors-cheat-sheet); I wanted it to be obviously a gradient, so I just picked colors that were immediately adjacent to each other in the ANSI palette. It might be better to pick colors that are one step apart from each other, instead, so they're more distinguishable visually...but there's kind of a balancing act between distinguishability and having a clear gradient. We'll keep playing with it!

mycoliza··on Tokio Console
Hi, I'm one of the main authors of `tokio-console`, so if folks have any questions, I'm happy to answer them!
mycoliza··on Tokio 1.0 – async runtime for Rust
Rust is intended as a systems programming language, it's for people who are writing "the next nginx". It turns out that there are also a bunch of people who want to write webapp servers in Rust, too, but that's never really been the goal.
mycoliza··on Diagnostics with Tracing, a Unified Instrumentation System for Rust
This is a great question, and the answer is "it's complicated". The core `tracing` libraries don't use either; instead, they provide an interface for `Subscriber`s (the pluggable component that collects & records trace data, kind of like a logger but fancier) to implement a way of tracking whatever contexts they care about. The typical approach is for the subscriber to track a current span per thread, but they could implement something else.

`tracing` instruments futures by wrapping them with a future combinator that enters a span each time the future is polled; the `#[tracing::instrument]` attribute will do the same thing under the hood when used on an `async fn`. This is kind of analogous to the Go-style context parameter, in that the contexts are stored in structures or on the stack, except that users don't have to manually pass the context around.

The core library provides an option to set the `Subscriber` that collects trace data in a scope; this does use thread-local storage. However, the default dispatcher can also be set globally (like the `log` crate), and the use of thread-locals is feature-flagged so it can be turned off by `no-std` users.

Finally, I have some thoughts on an abstraction for "context-local" storage that allows the user to customize the context that's used to shard the data. This could be used like a user-space version of OS thread-locals when threads are present, but it could also be used by bare metal code for (say) having a context for each CPU core. This would allow subscriber implementations to track a span per thread by default, but let embedded or kernel-mode users override this without having to reimplement the rest of the subscriber logic. This is still in the early stages though.

Hope that all makes sense; I'm happy to answer any further questions!

Page 1 of 2Next →