How Turborepo is porting from Go to Rust
vercel.com
vercel.com
Maybe I've just had a bad sample size, but I just haven't experienced a big enough win by using alpine to justify the weird scenarios that come up on occasion.
The idea is nice. A small, stripped down container that will load quickly and have very few maintenance issues (due to basically zero dependencies). But my debian-slim images work well enough and when I do hit a problem there's more community around it and it's more straightforward to fix.
Personally, I would love more of them to consider FreeBSD. It's got all of those features, a linux compatibility layer if necessary, a more permissive license, etc. I'd just love to see a lot more developers helping over there.
I see this as a net disadvantage. The great insight of the GPL licences is forcing changes to be contributed back to the project, companies can't easily privatize a public effort.
Whatever breaks on Linux/musl, is also likely to break on other operating systems.
Whatever breaks loudly elsewhere, is also quite likely silently broken on Linux/glibc.
If our default response is to double down on barely patching it enough to limp along, no wonder it keeps breaking.
You could probably do something similar in a traditional distro by creating an alias package that prints a warning on install, but then you actually have to watch the install logs to see that.
DNS TCP resolution doesn't work (https://christoph.luppri.ch/fixing-dns-resolution-for-ruby-o...). It's a feature that it doesn't retry truncated lookups or obey the relevant RFCs (https://twitter.com/RichFelker/status/994629795551031296).
Alpine has a 128k stack size (https://ariadne.space/2021/06/25/understanding-thread-stack-...), while most other OSes have at least a 512k stack.
Quite often fixing a build or a package on another system or architecture is a matter of a 3-5 line diff, and I've contributed quite a few over the years. It usually boils down to an incorrect assumption by the author/maintainer, stuff like "#ifdef __linux__" (when what you actually mean is: "any UNIX-like system with X11"), or hardcoding CC=gcc (where CC=cc just works).
There’s been a lot of discussion of the marshaling and unmarshaling costs of Json (and really any other text format). Should we be looking more strongly at binary formats, and if those have better outcomes than Json?
I’m guessing they didn’t want to introduce more change to the Go code than necessary, but I’m wondering if people have any horror stories to share about any of the binary formats available?
But that’s my question, if you choose a binary format, what are the issues to look out for?
Example, I work with Java, Rust, Go, and Python at work. I have specifically had issues with Avro (not my choice) which is really well supported in Java, but not much else.
Binary data support is pretty nice too for avoiding multipart request bodies.
That said, there was a comment in their post about not wanting to tightly couple the Go and Rust since they weren't sure they could accurately represent their structures in both langs.
I don't know anything about this turbo tool, but it seems relatively low-throughput, and they intend to migrate it all to Rust anyway, so it's probably not a big deal.
That said; I'd have done this with binary comms and not JSON, ProtoBuf or not.
For example, it prioritizes backwards and forwards compatibility -- which is not a concern for IPC where you control both ends. So no real optional or required fields. Comparing structures (Go) is awkward and uses reflection. Structs embed a mutex and preserve bytes for unknown fields etc...
MessagePack seems to make a better set of tradeoffs by comparison.
In my case, I was thinking about storage costs for JSON data; but thinking about this use case: Isn’t it true that the CPU would have to spend basically three times as much time making FFI calls? Assuming that practically speaking, field names make up 2/3rds of the bytes in their real-world JSON data.
Though, I know that another team at the same company loved it, but they were all in on java (and groovy and kotlin). As I understand it protobuf meshes with Java's type system much better (and probably Go's too?).
---
I did some benchmarking on various formats/rust libraries when we were deciding on what the protocol should be used for ^ and, IIRC, JSON fared better than you'd have expected. My assumption is that JSON just has had way more eyes/hands on the implementation than anything else. The things that I can remember beating it were (1) bincode, (2) cbor, and (3) protobuf, then JSON was #4.
In retrospect I wish we'd have gone with bincode, but we weren't sure if we were going to need some java code to interact with this service at the time and, at least at the time, bincode was rust-only (and possibly not even stable between compiler releases?). It would have been much faster to develop with and there was a whole bunch of overhead from protobuf that didn't really do anything for us since we were talking over a unix domain socket within a single system.
Whether the codegen/libraries for a particular language provides a more idiomatic binding for these well-known wrappers is up to the implementation - for example, golang libraries have conveniences added for well known libraries: https://pkg.go.dev/google.golang.org/protobuf/types/known. Rust libraries may have the same; I'm not as familiar with the ecosystem there.
JSON is awful in every way.
An idiomatic Go structure does not map well to an idiomatic rust structure, and similarly other languages.
But it's good enough and has very wide and mature support for many languages and a good story for backward compatibility which is important when you have files you want to read 2 years from now or when you have different version of your clients and servers co-existing.
There _could_ be a better interchange format in theory (and many have been proposed), but a combination of maturity and network effect (and the fact that many of the alternatives to protobuf do improve some things but do other things worse than protobuf) make it very hard to chose an alternative.
EDIT: ah, about discriminating zero value from unset value (aka options), they're back in proto3 (experimental) https://github.com/protocolbuffers/protobuf/blob/main/docs/i...
https://github.com/golang/go/issues/13492
golang cannot be compiled as a C static library with musl.
Yes, there's overhead in starting a new process to "just call a function", but I think this approach is still underutilized.
[0]: https://github.com/vercel/turbo/blob/c0ee0dea7388d1081512c93...
In one of our implementations, we use the `amazon_kcl_helper.py` script (python) to download and configure the KCL Multilang Daemon (a java application) which calls out to our application (a golang binary) via STDIN/STDOUT.
And honestly, it's fine. The Multilang Daemon manages shard ownership, spins up one binary per shard, fetches records, passes them to the binary, updates checkpoints, etc. Our binary just focuses on our actual application logic. And the python script does a good job of setting everything up.
It just feels wild to involve all three.
Inputs are arbitrary string STDIN and tokenized command line args (argv/argv). Outputs are arbitrary string STDOUT/STDERR and well-formed simple uint8_t return codes.
For input, most tools use the command line args for specifying tunable parameters. These are easy to adjust on the shell or in shell scripts, as opposed to deserializing & modifying & reserializing input files or STDIN. Most use the getopt convention (`myprog --input myfile -b 123`), which makes them well-formed enough to be a stable API and describable by simple grammars. As they are (usually) position independent key-value pairs (with some value arrays for multiple files etc), it makes it easy to add options without breaking existing callers. Look at the mountain of shell scripts out there that continue to run even as tools are updated.
For output, nothing similar exists. It would be interesting if UNIX tools could also output argc/argv and shells had builtin functions to easily parse (getopt style) and index into those. Or even just have a flat key-value envvar style return list. (I guess you could have a convention of the last line of STDOUT being dedicated to that and piping it into some tool that sets a bunch of prefixed env vars). Would have made everyday output parsing from random tools a whole lot easier instead of grep/sed/awk and the dealing with changing output as people updated the tools ("ifconfig" versus "ip" well-formed grammar).
Arrays and nested data structures are where everything gets complicated, and you need something like Powershell (powerful but clunky IMHO) or JSON and full programming languages. jshn is cool but verbose as its necessarily just a POSIX shell "extension".
I'm currently learning Go but have had my eye on Rust as well and have been wondering about the strengths and weaknesses of each language and how they compare and where each language is best applied.
The linked blog confirms my first impression that Go is well suited for networking applications and simplicity whereas Rust is well suited for OS/low-level applications.
Does anyone here have any experience using Go and Rust? What are your thoughts on the two languages?
IMO the most compelling use case to use Go over Rust is if you're writing a networked service or networking code.
I've written a couple protocol handlers in rust (that is parse the packet and handle the logic before giving the data to some other bit of code) and I found Rust really nice for that - real enums, pattern matching, and the use of traits can make interacting with binary protocols (particularly of the state machine variety) really nice with ergonomic interfaces. Meanwhile go's net/IP type (etc) often leave me frustrated.
On the other hand, goroutines and channels are such a nice way to handle a lot of message passing around a server application that I find myself reaching for mpsc channels and using tokio tasks like clunkier goroutines in rust often.
I think I'd draw the line more along "how often I need to work with byte arrays that I want to turn into semantic types" - if I'm going down to the protocol level and not doing a ton of high level logic I'll reach for rust, if it's more high level handling of requests/responses (http for example) I'll often reach for go.
(like the parent, I'm more inclined to reach for Rust over go if all else is equal, but I enjoy go too).
My experience so far with mpsc/tokio tasks is that Fearless Concurrency TM is pretty solid. There is some clunkiness (mostly inconsistency around async styles), but I've been bit by bad channel usage in golang enough times that rust felt refreshing.
When I analyzed Rust around the same time, I noted the Rust standard library [1] did not have HTTP support, and it wasn't a first class consideration. I think Hyper [2] was around, but I've never analyzed it deeply (though based on GitHub stars it seems to be popular). Protobuf [3] is also extremely easy to work with in Go.
Given the differences in standard library and applications I see created in both ecosystems, your analysis seems right, though you can probably develop network(ed) applications in both fairly well at this point.
[1]: https://doc.rust-lang.org/std/
Most recently, I'm trying to build a non-allocating, minimally copying `Lines` iterator which reads from an internal `Reader` into an internal buffer and then yields slices of the internal buffer on each call to `next` (each slice represents a single line, and it's an error if a line exceeds the capacity of the internal buffer); however, as far as I can tell, this is unworkable without some unsafety because the signature is `fn next(self: &mut Lines<'a>) -> Result<&'a [u8]>` which doesn't work because Rust thinks the mutable reference to self must outlive 'a, and explicitly setting the mutable self reference to 'a violates the trait. If I forego the trait and just make a thing with a `next()` method, then I can't use it in a loop without triggering some multiple mutable borrows error (each loop iteration constitutes a mutable borrow and for some reason these borrows are considered to be concurrent). The only thing I can think to do is have a `scan()` method that finds the next newline and notes its location inside the `Lines` struct and a separate `line()` method that actually fetches the resulting slice from the buffer.
I don't run into this in C or Go, and for all of the difficulty of battling the borrow checker, I'm not getting any extra safety (in this case).
There are downsides to this approach because it uses internal iteration while most things in Rust use external iteration. But shit happens. Go doesn't even have a first class concept of iterators as an abstraction (yet), and its standard library contains patterns for both internal (sync.Map) and external (bufio.Scanner) iteration.
> The only thing I can think to do is have a `scan()` method that finds the next newline and notes its location inside the `Lines` struct and a separate `line()` method that actually fetches the resulting slice from the buffer.
Yup that works too. That's basically the design used by the `streaming-iterator` crate: https://docs.rs/streaming-iterator/latest/streaming_iterator...
> I don't run into this in C or Go
Of course you don't. Neither C or Go even have abstractions called "iteration" at all. They have patterns for them. And Go lets you iterate over a fixed set of built-in types. (Currently. Maybe it's changing: https://research.swtch.com/coro)
Besides, Go has a garbage collector. You should expect all sorts of patterns involving memory/copying to change when you move from a language with a GC to one without.
However assuming trying to keep to 0-alloc odds are good you'd probably just be telling the user to git gud as the slices would become nonsensical on every refill of the buffer.
Also because of the missing middle: in Go (or Java, or C#, or even python) you could hand out slices to a large buffer, but discard the buffer rather than refill it in place, relying on the GC to clean things up (and possibly give you the same buffer on the next allocation). So you get an efficiency middle ground where you allocate more than strictly necessary but not for every line, and maintain correctness.
In Rust that’d require some sort of Rc projection which I’m not sure even exists?
My goal isn’t “implement an Iterator trait”, I just want to iterate over lines in a file, so I don’t especially care that Go and C don’t have an iterator type.
> Besides, Go has a garbage collector. You should expect all sorts of patterns involving memory/copying to change when you move from a language with a GC to one without.
I’m not allocating, so the GC doesn’t matter. The Go and C versions look essentially the same—it’s only the Rust version that I had a hard time with because of the borrow checker.
> I just want to iterate over lines in a file
No... you don't. You specifically said you wanted to do this without allocating and minimal copying. That's a different problem than "just iterate over lines in a file."
> I’m not allocating, so the GC doesn’t matter.
Of course it does. The GC manages the lifetimes for you. In your C code, you manage the lifetimes yourself and likely rely on the caller to not fuck things up.
Yes, of course. The point is iterating over the lines of a file given the aforementioned constraints and not “implementing Rust’s Iterator abstraction”. Whether or not Go or C have iterator abstractions is immaterial.
> Of course it does. The GC manages the lifetimes for you. In your C code, you manage the lifetimes yourself and likely rely on the caller to not fuck things up.
GC only manages allocations. There are no allocations here to manage.
You're missing the forest for the trees. The GC is integrally tied to lifetimes. If you have a `*Foo` in Go, that may or may not be on the stack. It might be on the stack, thus no allocation, if the Go compiler can prove something about its lifetime and usage. Same deal with `[]byte`. Maybe that's tied to an array on a stack somewhere. Maybe not. AFAIK, in order to get guarantees about this in Go you need to drop down into `unsafe`.
> Yes, of course. The point is iterating over the lines of a file given the aforementioned constraints and not “implementing Rust’s Iterator abstraction”. Whether or not Go or C have iterator abstractions is immaterial.
That's why the very first thing I said to you was to point out the closure approach. I even linked you to real code (that I've written) that does it. Yet 'round and 'round we go.
Rust has a much better type system, Go has a much better standard library. The problems with Rust are that the async solutions are shifting sands and that most people "cheat" by just shoving everything into a Box on the heap instead of properly using lifetimes on the stack. Both of those lead to a lot of inconsistency and sort of defeat the purpose of using Rust in the first place.
The main problem with Go is verbose error handling, it may also be too high level for some systems level tasks Rust may be better suited to.
Some people want the performance.
Some people want the safety guarantees and the checking the compiler does for correctness.
Some people want the expressive type system (and combined with the above, the ability to encode constraints into data types).
I don't think that people who choose based on the third but don't need super tight performance are using it wrong, they just have a different set of priorities than those that are using it for very high performance.
I was actually a bit astonished by it. The most trenchant concrete example they could give is... Go's standard library abstraction around file permissions is a bit inconvenient for their use case? This is easily fixed in Go in about a day without a complete rewrite; you just write some new stuff directly against the relevant syscall libraries and use build tags to control the different OS builds.
Just because an abstraction exists doesn't mean you have to use it! It is not a requirement of Go that you must use that particular abstraction, it's just something provided by the standard library. We once wrote a library for Perl using the openat and other *at functions because we needed the security guarantees, and it worked fine, despite the "standard library" not working that way.
My read is that they're basically doing this because they want to, not because Go forced them to. That's fine. There's nothing wrong with that, if they want to pay the price for migration. (Were someone proposing it to me in real life I'd expect a better reason than "I don't like how it handles file permissions in the standard library", though.) However, as an external observer using it as a grounds to decide it's much weaker than may meet the eye; I wouldn't overprivilege it.
The real reason to prefer Rust here would be something on the lines of "We're doing so much crazy concurrent stuff that we need the guarantees that Rust provides via compiler but Go only provides via common practices." The latter can carry you a long ways but it does eventually give out and become insufficient for a codebase, and when that's the case Rust becomes one of the short list of options. "Go common practices" is actually pretty high up on the set of "ways to do reasonable concurrency", reaching above that requires a pretty significant shift to Erlang/Elixir, Haskell, or Rust, and given the goals of the project that would basically leave Rust as the only acceptable performance choice. It's possible that this would qualify for them, though I'm not sure... from what it sounds like they're doing, a task management system that ensures that only one worker is doing a given task and everyone else waits on the relevant output without redoing it would be the core of their architecture, and even if you use that thousands of times per execution that doesn't necessarily mean the code is complex internally.
Not when pitching the project to hard-nosed non-technical managers burdened with personal responsibility for P&L results.
I have a mental category for these posts:
We saved some of cloud / server side cost with Rust, so you user can spend more on local desktop /cloud cost with Javascript
Sometimes cost savings can be imaginary but since Rust is cool it is all fine.
Rust takes a bit more effort to initially get results, but once you get there, you work in a language that gives you expressive ways to create abstraction (this is a matter of taste maybe), generates generally faster code, and does not incur overhead by garbage collection. Also, there are ways to write asynchronous code in an elegant way. The price is that you spend more time learning about structuring programs in a way that the compiler accepts. Once you are past that point, you'll be at least as productive as in Go.
I used to dabble with Go and like it, but once I had to write more code, I found it more tedious compared to Rust. But as I said, more readable to the uninitiated, that was one of the Go design goals.
I’m currently back to Python and losing my mind about doing error handling, as well as enforcing correct usage of my library API on the call site. Doing these things well (best?) in Rust is baked into that language’s DNA (Result, Option, newtype pattern, type state pattern, …). It’s painful to go without once you’ve seen the light.
Pydantic enforces types when you're working with a Pydantic objects. However, it doesn't help with functions that accept vanilla types as parameters. validate arguments checks that the callers parameters matches the function's type hints.
This is helpful because a static analyzer like mypy won't catch the wrong type being passed in all situations.
I've worked on an extremely large golang monorepo, but I don't have production Rust experience. They're vastly different languages that they shouldn't really be compared. The only commonality they have is that they both compile to native code, that's about it. Other than that, golang is basically a python/ruby/perl/php competitor, whereas Rust is a C++ competitor.
First, the two languages were initially revealed around the same time, and they were both new AOT-compiled languages targeting a bit downstack from your Java and C#, which was in opposition of the trend since the 90s.
Second, Rust was specifically targeted at "system programming" from the start, as from the start Mozilla desired using it to displace C++ in the browser; Go was also initially billed as a "system programming" language, though for a completely different interpretation of the term.
Finally, not only did the initial rust announcements very much emphasise concurrency (something Go also does), the initial rust was a much higher-level language, with a mandatory runtime, green threads, and (plans for) a GC.
Even though Rust changed drastically between the initial announcement and the 1.0 release, and went significantly downstack (aiming much more squarely at C and C++), those initial impressions have left lasting traces in the global consciousness.
And while the two have generally different niches, they very much do compete when it comes to... CLI utilities.
Our biggest struggle with golang was that our Ruby/FP backgrounds just clashed with golang's style. It wasn't terrible, but we always felt it was more verbose and clunkier to compose and reuse things than we expected. I'm sure generics are making this better, but even with generics the general feel of the type system just felt off to us.
Our monolith is http/graphql + postgres in Rust and we like it. Some learning curve for the team, but most of our "application code" is pretty easy Rust and people can jump in and work on it quickly enough. Most work is defining in/out types and implementing a few key traits for them. Our client engineers are comfortable doing it. Our "app infra code" is a little more technical, but we've appreciated the strong type system as we've built out pubsub systems and other important bits of infra.
Ultimately though, it's just a culture thing. I'd recommend one of jvm/rust/golang to anyone trying to build apps that handle significant traffic. Which of those 3 is just whichever matches up to your team and their personalities.
I've been programming for 12+ years and it gets real boring when you've to type/copy same thing again and again, and Go doesn't want to help you here.
Using alpine for achieving slim containers is a waste of time and frankly stupid.
2. ultimately getting rid of cgo would probably help a lot anyway, cgo calls are pretty slow
Also JSON is used when delegating to Go because the Rust binary does not handle the command anyway.
It’s slow compared to a function call in C, but it’s still measured in nanoseconds. If you’re not calling C functions from a tight loop in Go it’s unlikely to be noticeable.
Sure I can. Maybe YOU can't, due to the license you have chosen, but just say so.
"You" here is a generic pronoun, not "you, reader"
It's an issue from a licensing point of view, but it's also a rather poor idea from a technical point of view.
on glibc's part.
These blog posts make me laugh. Pretty sure they could achieve whatever they are doing with any language, heck even PHP would probably work. Just not cool though is it.