Rust in production at Figma
blog.figma.com
blog.figma.com
EDIT: Quote from the OP:
> One of them is called `error-chain` and another one is called `failure`. We didn’t realize these existed and we aren’t sure if there’s a standard approach.
`error-chain` is in maintenance mode these days; `failure` is its spiritual successor and seems that it's on the path to becoming a community standard eventually.
1. A case study of Rust used in something performance-critical. People pushing or just assessing it like to see those.
2. They say they're all in on modern C++ but still trigger undefined behavior and crashes a lot in practice. Member's of Rust team regularly claim that happens despite C++'s safety improvements over time. In this case, it was happening in C++ but not in Rust with C++ developers using both. Far as learning complexity, they're both complex enough that C++ coders should be able to learn Rust. So, the main drawback is in both languages. Looking at complexity vs reliable iterations, Rust provided an advantage in knocking out problems that were slowing the C++ coders down sometimes by "hours" of debugging.
3. On parallelism and concurrency, their prior method was getting a single-core implementation working first that they transformed into something for multi-core. This was giving them a lot of trouble. That there's lots of ways to implement concurrency in C++ means I can't be sure if it was due to the language, what library/framework they were using, their own techniques, or some combo. Also, I'm not going to guess since I don't use C++. :) Regardless of what they were doing, the alternative approach in Rust let them get a lot of stuff right on the first try. So, it was easier for those C++ coders to do multicore in Rust than in C++. It's evidence in favor of Rust teams claim that Rust makes concurrency easier with less debugging given they were newcomers immediately getting good results. However, I don't think it's evidence of anything comparative between Rust and C++ concurrency without knowing what they were doing in C++. As in, some C++ coders might be getting better results with different methods where gap between them and Rust might be anywhere from zero to smaller than this case.
4. Finally, handling platform differences was so much easier as newcomers using Rust's tooling than it was for them as C++ veterans that they still saved time overall in Rust despite having to implement Rust support for game platforms themselves. That's strong evidence that Rust's package manager or platform tooling is excellent with some anecdotal evidence it's better than C++ at this cross-platform, use case.
So, there's a summary of what I got out of the case study for anyone that might find it useful.
EDIT: That Reddit thread has some great comments about floats, too. Problems, solutions, Rust vs C++ handling, and so on.
Wikipedia says
> In business, a white paper is closer [than the government meaning of a policy document] to a form of marketing presentation, a tool meant to persuade customers and partners and promote a product or viewpoint.
I do think that some data would be good, but at the same time, the kind of data that's relevant here is really hard to get in any concrete terms, so might actually be harmful.
> a "paper" of any sort
Remember that this isn't a paper of just any sort: it's a whitepaper. So it has a specific audience: CTOs, VPs, and other technical management types. Audience is important for what you have to include in any work.
That said, I will happily concede that I have no idea who the target audience is for the whitepapers and probably don't know the accepted industry definition of a whitepaper. However, I'm very happy to see Rust flourish either way.
Affinity is pay once, use as long as you want.
It looks like Figma is a SaaS. Might as well pay for Adobe if one wants to give up even the meagre amount of ownership they had over their software.
Is there anything you miss from sketch? what do you think sketch does better?
Reading between the lines here, they didn't go with a more mature language like Java because they were worried GC tuning would be a problem?
Given all the other issues they noted with using a less mature language like Rust in production, that's a pretty heavy load to take on in exchange for not having to tune GC. Isn't GC tuning a fairly well understood problem? Is there something about encoding large documents that makes it a significantly greater obstacle?
Or are there other unstated considerations at play here? For example, I mean this completely earnestly and not cynically, but there is a lot more PR and recruitment value in blogging about a hot new cutting-edge language than "how we rewrote our TypeScript server in Java".
To name just two ways in which I consider Rust more mature than Java:
* It has a lot of fundamentals that are based on more academic languages which have explored some specific PLT space for a long time and let it mature. The fruits of that are now in Rust (many aspects of its type system for example).
* Rust's community has an almost absurd ability (contrasted with most other languages) to focus on specific core libraries and tooling. For example serde[1], the Rust serialization library, is well-understood, simple, mature and supported in basically every Rust library out there.
The second point is extremely impressive when contrasted with the state of things in Java-land, where you often have many different solutions for the same problems. Sometimes even the well-known and used ones are of very questionable quality[2].
[1]: https://serde.rs/ [2]: https://github.com/FasterXML/jackson-databind/pull/1423
[1] https://blog.takipi.com/the-ultimate-json-library-json-simpl...
Your second point is invalidated by the blog itself: "Many libraries are still early".
GCs also tend to have cliffs where you hit some unknown threshold when they start thrashing where deterministic release of memory tends to be much more linear.
I've experienced this, having previously worked on Java services.
It can be quite a pain and can throw a wrench into your understanding of how you intended to scale.
The post says GC was already a problem in their old system.
FWIW, I'm excited for Rust and also Swift for exactly these reasons. I just want to see hard justifications for using them in production this early.
I think the question was profound (if accidentally so), and I agree with him: when performance and footprint matter, it is reasonable to expect software engineers to manage their resources in a predictable manner -- which excludes garbage collected environments.
Rust is a very interesting candidate for such environments, and I think the experience described here is quite helpful for those of us seriously contemplating the language.
Then, that is what "tuning" usually means in engineering. See also "PID loop tuning".
Potential solution 1: Replace with code that doesn't have garbage collector overhead in the first place
Potential solution 2: Replace with code that has different, robust garbage collector
I get where you're coming from, but even I would be inclined to choose solution 1 in this case. That they chose Rust instead of C++ is their own business, but that they chose either of those over Java just seems sensible to me.
They can benefit further by choosing something more performance-oriented than Node.js, sure. But they weren't having issues with their whole application failing - just a specific piece. When you want to expand your driveway your first thought shouldn't be what materials you want to build your new house in.
Wonder what the article would have said in a parallel universe where the problem was migrated to something like the Akka framework on the JVM?
That being said, they mention being turned on to low resource usage specifically because of garbage collection issues and resource overhead. While the JVM is lighter than the Node.js runtime when considered across multiple cores I think that's just going to give them more runway to tackle the underlying problem they were still having.
"Instead of going all-in on Rust, we decided to keep the network handling in node.js for now. The node.js process creates a separate Rust child process per document and communicates with it using a message-based protocol over stdin and stdout. All network traffic is passed between processes using these messages."
For example, I ran into the futures issues. Now I use nightly's async/await, and it's a massive improvement while still being early days.
I also use the Non Lexical Lifetimes feature, and see a lot of ergonomic wins there.
The other stuff is being worked through as well - Failure is coming along, and I hear good things. Libraries are always improving.
I would be a lot more terrified if these issues were surprising, or not being dealt with, or were extremely hard to fix without breaking the language, etc. Instead, it's a list of things I've run into and can even solve today with a few features.
> While we hit some speed bumps, I want to emphasize that our experience with Rust was very positive overall. It’s an incredibly promising project with a solid core and a healthy community. I’m confident these issues will end up being solved over time.
Rust is still a relatively new language, and most of the cons list fits in with that. They also talked about how they worked around these issues, and what we're doing to address them, which is pretty wonderful.
I'm also happy to talk more about any of the specific points here to add more context; for example, the comment about error-chain and failure is spot on. error-chain is the older, more battle tested library, failure is the newer one that's try to address some issues with it. The ecosystem is still shaking out. (I personally love failure.)
One needs to have a robust instinct for separating the excited statements of early adopters and fans from the more mundane reality: developing a language ecosystem takes a long time and other languages have had decades to iterate.
Would decoupling the workers and the documents they work on not solve this problem? Granted this might be non trivial, but it might have solved the fundamental issue thats arguably more interesting.
Admittedly, the blog is light on details here and I am unfamiliar with the product. With the re-write, they also just moved the problem to rust. So I think that your suggestion to move the problem to the storage layer could have been another viable solution.