Rewriting the heart of our sync engine
dropbox.tech
dropbox.tech
Really good blog post, imo.
Rust was adopted at Dropbox for some serving infrastructure use cases more than a year before the sync rewrite was started, which was about four years ago. I'd say we solidly predated the "rewrite it in Rust" meme.
I believe that this rewrite was only successful because of Rust's ability to both interact safely/efficiently with underlying OS APIs (they're pretty much all C-like) and to encode complex concepts into the type system and the compiler. Rust isn't the only language with these properties, but it is one of the few -- and it's one that we really enjoyed using.
Just curious, what are the other languages that competes with Rust in terms of correctness, safety, ergonomics and efficiency?
Go is more on the ergonomics side and Haskell on the correctness side, but are there any serious alternatives for Rust that checks all the boxes?
Go is tricky if you run on diverse platforms and want to do a lot of FFI, since cgo overhead is significant. The type system in Go is also not very powerful, which is both good and bad.
Haskell tends to hit performance walls that are very difficult to debug, and has a pretty similar learning curve to Rust (most people you hire onto the team won't know the language already).
The predominant competitor in this space is probably a high-level dynamic language combined with C/C++ library code. With good tooling and good practices to mitigate footguns, the extensive library support in C++ has a lot to offer.
Of course, I think Rust makes a better trade-off there, but early in the project it was not at all obvious that the good parts outweighed the fact that we would probably have been the biggest user of any library we depended on. We had some fun adventures in stress-testing HTTP/2 support here :)
> As soon as I paid for the Plus Plan, I started having to wait several minutes for a file to upload. Even from one computer to the other with LAN Sync enabled. The day before that, my files would sync faster than I could hit refresh in the browser I was developing in.
that definitely seems like a bug, could you report it? click on the dropbox tray, click on the dropdown in the top right, and then click "Report Bug." this will collect some information about your client to help us debug.
also, a lot of the standard library APIs are really well thought out, and it's great to then use the type system to build great internal APIs as well. the ability to design APIs to make correct use easy and incorrect use difficult feels like good "ergonomics" to me :)
Contrast this with other languages, which are often clunky when they don't need to be and/or "easy" when they shouldn't be.
Have you read https://fasterthanli.me/blog/2020/i-want-off-mr-golangs-wild...?
OT: I really really hate the top bar that drops down and covers what I'm reading when I'm trying to scroll up. It's such a frequent pattern on the modern web and I can't wait for people to get rid of it. (It violates the symmetry of scrolling up and down)
Also, after the startup phase, IIRC Dropbox used the `FindFirstChangeNotification` Windows API function, which can miss changes under heavy load - do you still use this method, or have you moved to something else, such as using the NTFS USN journal, or a file system minifilter driver?
Also, have you previously explored the possibility of using the NTFS change journal, and if so, what challenges did you face? (I don't mean to come across negative; I'm genuinely interested to learn why you didn't use this a decade ago, since there must be a reason).
The biggest challenges with this kind of work (not specific to the USN journal) are in the lack of reliable cross-platform support for features which can be shimmed to look similar. The more unique the code path, the harder it is to test en masse.
It's also the case that the execution environment on Windows machines tends to be very diverse due to a long history of backwards compatibility and the ability to set complex domain policies -- Dropbox strives for a good user experience, and "go talk to IT to have them change this setting" is rarely one of those.
disclaimer: worked on sync at Dropbox; don't work on it anymore; don't have current context
Any ideas? To help with recruitment?
Hint, open the file directly, zoom out a bit, and squint.
it's easy to have "nested" futures in rust, where one top-level future (our control thread) owns its subcomponents (like a protocol component, scheduler component, and so on), and the `poll` method for the top-level future then `poll`s all of its subcomponents. in this case we manually write our own `poll` implementations, but it's also easy to express these patterns with combinators like `FuturesUnordered`.
then, we have really good cancellation properties in our system. if we decide to cancel a file, we just drop its future, and we can be sure that all of the resources held by its "subtree" of futures will be released.
I'd be curious to hear your take on the "structured concurrency in Rust" conversation here https://trio.discourse.group/t/structured-concurrency-in-rus...
in my understanding, the pure futures approach in rust goes even further than the nursery approach in trio, where there isn't even a "local" notion of forking off a task in the background. if a future wants to do an operation in the background, it's responsible for `join`ing on it itself or finding some other way to ensure it gets `poll`ed. this setup is a good fit for rust since the parent future maintains ownership over its child, and dropping the parent future immediately cancels the background tasks.
but, for what it's worth, I think there's lots of tradeoffs in this space, and we have yet to find really good patterns for structuring async code. so it's good to see so much experimentation.
The code goes into this state because no one appreciates piece size enhancement. And those enhancement has no "impact". Everyone is waiting for the point that everyone's suffering from the bitrotten code is beyond most people's tolerance. And everyone knows well that some really bad things happen during the process, and they'll wisely refrain from doing anything meaningful.
That's basically one fact of how software engineering is done in practice.
Poor humans.
And this phenomenon is well known even in ancient time in other areas of society. Like even in Qin dynasty cicra 300BC, a Chinese physician Bian Que once claimed that his brothers are better than him [1], because they prevent illness before they occur.
And of course, everyone only knows Bian Que.