The Safety Boat: Kubernetes and Rust
msrc-blog.microsoft.com
msrc-blog.microsoft.com
It's often about how fast one can deliver maintainable software that works well enough.
Go was so "easy" that I was immediately productive, but this resulted in often suboptimal and messy code that I had to refactor over and over again due to concurrency, abstraction, or performance issues.
True productivity is hard to measure.
Pretty sure that Go isn't equally restrictive (while still being garbage collected).
It is a better option than keeping using C, and it would have been great in 1996, but that is about it.
To learn Rust you simply need to understand how values are kept track of by the compiler. Once you develop an intuition for this it's the same as any other modern imperative programming language.
Going from something like Java or Python to Rust, one would have a lot to learn.
It is a data race, not a race condition.
> and which passed the race checker for Go
No, it is not. https://github.com/helm/helm/pull/7820#issuecomment-60436062...
There is a comment by issue author which is literally a go data race detector warning. Like "WARNING: DATA RACE".
But you're right, this is just choice of terminology. :)
var i int
doneCh := make(chan struct{})
go func() { i = 1; doneCh <- struct{}{} }() // a
go func() { i = 1; doneCh <- struct{}{} }() // b
<-doneCh
<-doneCh
At the end of the program, i is always equal to 1 no matter which order a or b wrote to i. But it's a race because you are assigning to a shared variable without synchronization. A small modification to the program creates a race condition: var i int
doneCh := make(chan struct{})
go func() { i = 1; doneCh <- struct{}{} }() // a
go func() { i = 2; doneCh <- struct{}{} }() // b
<-doneCh
<-doneCh
Is i 1 or 2? It depends.It is correct for the race checker to complain about the first program, because after a bit of hacking the first program can very easily change into the second program.
(And I tried it, and it does complain.)
But still, it detects this error.
Here's what happens. Delete takes a ResourceList. It delegates to "perform" and then "batchPerform". perform calls batchPerform in a separate goroutine, which calls a helper function in another goroutine for every resource in the ResourceList. The helper function is defined in Delete and updates a data structure defined in Delete. This is a classic case where some synchronization is necessary. The function runs multiple times in multiple goroutines, and updates a single shared structure. (Perhaps not obvious because it delegates to two helper functions, and the list that the function is executed on is a "ResourceList" not a []Resource, so it isn't clear that there is a "for { go func() }" loop anywhere; the programmers did their best to make it non-obvious that a loop is occurring.)
The confounding factor here is that batchPerform tries to synchronize with a WaitGroup, but it's faulty and not enough to protect the data integrity. batchPerform creates a WaitGroup, but only calls Wait() on the WaitGroup when the "kind" of an individual resource is not equal to the "kind" passed to batchPerform. I am guessing that it's very natural to craft some test data where this condition is met, and the for loop in batchPerform only runs the function once at a time (perhaps a ResourceList of length 1). In that case, there is no race condition for the race detector to detect.
All in all, if I were reviewing this code, it would not be checked in its current form. Splitting perform and batchPerform doesn't make sense to me, and they both implement faulty synchronization logic in a slightly different way. (batchPerform uses "for { wg.Add(); go f() }; wg.Wait", perform does "for range x { go func() { ch <- f() }() }; for range x { <- ch }". I consider these pretty much exactly equivalent, but neither prevents f() from running concurrently with itself. The only reason this passed the race checker is because batchPerform doesn't actually use the WaitGroup in the normal way, instead degrading to "for range x { wg.Add(); go f(); wg.Wait() }", which DOES prevent f() from running concurrently with itself, with certain inputs.
The root cause is that the caller of Delete isn't really sure about the semantics of "perform". Does it protect the body of the callback function? There is no documentation, and the author thought "yes". But the answer was "no". In general, the convention in go is to consider something thread-unsafe unless it's marked as thread safe. When you see something like "var foo Foo; f(list, func(bar){ foo = bar })" your spidey sense should be concerned about synchronization. But in this case, the code went out of its way to hide the existence of a loop and the existence of parallel processing, and so the programmer made a mistake. A bug or at least VERY confusing use of WaitGroup in batchPerform allowed the tests to pass. Should the compiler detect this? It would be nice. But a code reviewer should have been super concerned about this implementation.
Fwiw, it's too bad the commit message didn't say something like "Since we're doing delete on many resources in parallel, we need to hold a lock while updating errs/res.Deleted". The reviewer was also obviously confused at first.
[1] https://github.com/helm/helm/pull/7820/commits/edb2b7511bcb9...
I’ve heard people brag that Haskell is a great language because it’s supposedly easier to write correct code.
Rust has this same reputation?
Rust has many advocates now at places like Mozilla, Amazon and Microsoft that have delivered critical software in Rust that they believe has made it safer.
The fine print is that nobody claimed it's easy to write Haskell code that compiles.
My understanding of how async/await works in Rust is that you can have multiple async runtimes in one Rust program. Is that not the case?
It's definitely easier to only deal with one runtime. Ideally, we should have some kind of abstraction to allow crates to support both runtimes (e.g. a trait that'd allow creating an async TcpSocket of the right "kind" for your runtime), but AFAIK this is not currently done.
That's correct; we're still working on these abstractions. It's the end goal that most folks have in mind, though.
Other major abstractions that are missing so far include async versions of the Read and Write traits, a Stream trait for the async equivalent of the Iterator trait, and perhaps a way to spawn new tasks.
This series of interviews covers these in more depth: http://smallcultfollowing.com/babysteps/blog/2020/04/30/asyn...
That's really the case for any language where an eventloop is not part of a builtin runtime (like it e.g. is with Javascript or Dart). E.g. in C++ we also have boost asio, libuv, libevent, wagle, seastar,GUI framework eventloops in GTK, QT, etc.
The thing is once you are in async land, nothing is interoperable anymore in most environments. Whether that's ideal or not is a separate discussion.
What I experience however somehow is that Rust users raise a lot more concerns about interoperability than I've seen so far in other ecosystems. I might stem from the fact that those users often never used another native async environment.
In Rust the amount of people that actually work on the low level details and try to make things better is likely < 5. But there are a lot of expectations from everyone else about having perfect interoperability.
If your application has two parts or binaries that are completely separate you could potentially use two different runtimes, but otherwise I don't think it would make sense. And even then, it would just be a mess.
Right now, your runtime is essentially picked for you by your dependencies.
There is ongoing work to standardize more runtime interfaces so that more libraries can be runtime-agnostic.
t. C++ developer with a mixed std::string/QString/BSTR codebase.
This has absolutely nothing to do with generics.
None of std::string or QString are generics. They are just an example of historical alternative implementations for 'reasons' ( portability/speed) that create a lesson the long term
“average” and “proficient” are both very variable in that statement, imho.
It is true that we have a chat room with a bunch of folks, of which I’m part.
Even just by being availabletto answer questions or help with code review. Doing some pair programming sessions would probably be useful too.
Edit: one other big factor - presumably in their environment you have coworkers to get advice from. That’s huge when you’re first starting.
The Rust onboarding experience is incredibly explicit and once things start to click and code compiles, you're on the train.
I know that wasmtime can execute a WASM module and give it access to a file system. Can that filesystem contain a socket that the WASM module can interact with?
Conceivably you could compile all of the CPython runtime into WASM, just that you’d be left with a big binary that gets passed around all the time over the wire.
Imagine having JetPack Composer, SwiftUI, Qt designer, or WPF/UWP Blend in Rust.
Rust could totally do the same thing, and you could probably make it way easier to use than the mess that is Qt.
Remember that not only is GUI development with proper tooling very interacting, instead of the FOSS alternatives of code-compile-check visually, there is also the whole eco-system of third parties selling component libraries, with no control how they get integrated into the component toolbox.
So whatever solution one comes up with,it needs to be more productive than forcing users to scatter Rc<RefCells<>>, or fix their code that broke compilation, just because moving a widget on the GUI tree invalidated the borrow checker assumptions.
Don’t get me wrong, all technical arguments are correct and rust does have advantages for cloud software. But this also comes quite handy for MS. :)
While the possible security benefits of Rust is interesting in software like Kubernetes, it seems like this whole blog-post is an implicit RIIR proposal for the Kubernetes ecosystem from a Microsoft software engineer which isn’t going to happen anytime soon.
> Rust has made great progress in the past year with its async story, but there are still some issues that are being worked out.
On top of that, there are still many crates that aren’t using async-await yet and most are not even 1.0, thus are not stable. I would not touch such crates if they are still immature or even unsafe.
Realistically, a Rust Kubernetes is possible but practically the effort of a production ready version is measured in years.