How we designed Dropbox’s ATF – an async task framework
dropbox.tech
dropbox.tech
Neither this or the first method guarantees a lack of concurrent execution. A long GC pause or VM migration after the second check could allow the job to get rescheduled due to timeout. The first worker could resume thinking it still had one heartbeat left to execute before giving up on the job and it could've already been handed out to another worker in the meantime.
But often with technical blogs like this, you get a “dumbed-down” version that is inaccurate but summarizes in a few minutes what is essentially many person-years of work.
The sad thing is that Dropbox Product has so heavily dropped the ball that users like myself (from back in 2009) have switched away in droves over the past few years.
I understand that Dropbox core functionality wouldn't have been enough to multiply the valuation of the company to what investors expected. But it would have been nice to not jam collaboration features into the product and mess up the simple, platform-native UI with it's current abomination. I'd pay $10/mo forever if I could get the 2010-esque Dropbox Mac client and sync service back since it's way better than anything else (especially iCloud).
I do wish you could PAY for a basic version (maybe make the collab stuff free as part of some trial or something).
I use Dropbox personally to keep documents synced between my computer and my wife's and also to grab documents I need from the web if I'm on another computer. I occasionally share a folder if I need to give a large number of files to someone.
I recently had a notification come up on the dropbox taskbar icon and it popped up this huge window that looked like a massive electron app. In the old days, there wasn't even a UI, just a context menu that also showed the state of the sync.
For me, Dropbox provides the most benefit when it's not visible, running invisibly in the background doing it's thing.
The client's UI is a bit odd, but at the end of the day it's really good at what it's supposed to do: Syncing files.
Performance is also great: I'm using multiple machines to write code on, and I keep my local git repo on Dropbox. I can literally save a change on my notebook and run it on some other machine 3 seconds later.
On Mac and Linux you might want to check out maestral (https://github.com/SamSchott/maestral), a third-party client that works really well.
So yes, you can put a git repo in Dropbox if you want. But that’s still an altogether different thing than “hosting” a git repo like you might do with GitHub or Gitlab.
I mentioned above that I switched away; iCloud Drive is pretty mediocre since syncing sometimes doesn’t happen instantly and there’s no way to force-sync. I’ll probably move my Mac synced preferences back to Dropbox if this works well. Though I’m also worried about them deprecating their APIs since that seems to be a popular move these days.
The Linux Dropbox client is essentially abandoned (it's broken both functionally and aesthetically). On this platform, it definitely does not even acceptably do what it is supposed to do, so an alternative is good for the company.
I wonder though, why Dropbox limits the syncing functionality at the API level (by not offering partial file sync).
They've really gone downhill by adding unwanted bloat, and it just seems to be accelerating. Meanwhile their core product is degrading. Abandonment of the Public folder in spite of a huge outcry from customers was disappointing. The user experience is plastered with advertising to try their other products, even if you turn off all the relevant notification settings. And lately I've been running into subtle functionality bugs in the client.
Would happily give my money to a competitor focused on a lean, reliable product.
Dropbox is in a bad spot, I think in part because they haven’t focused deeply enough on real collaboration tools or other related product spaces.
Honestly, it works. Never had a bit of trouble. I might not be a mega power user, but all my stuff is there and I just don’t think about it. I moved away from Dropbox about 2 years ago – because I got sick of the shitty invasive macOS client software – and I regret nothing.
Bonus: invoking Spotlight on an iPad and having your files just show up is kinda magic.
I switched _sync back because Alfred and other apps warned that using iCloud Drive for sync doesn't work great because it doesn't always update properly, and I've found that to be the case.
I like these third party clients for things that are showing up (E.g. Apollo for Reddit).
I do like how iCloud and Onedrive have become so tightly integrated/transparent. (Benefits and banes).
Disclaimer/claim: I worked on this system and on gmail delivery.
Using email as an analogy, you have to commit the message to durable storage before you respond 250 to DATA.
To achieve at-least-once, you need only track which tasks have been successfully retired, and persist that knowledge in the database by either deleting or mutating the task. During a cold start you scan the persistent store to find tasks that were still pending/live at the time your process began.
I understand companies have different requirements, but if you look at the history even on Hacker News this problem is basically being resolved by different companies at least once a quarter.
It looks simple on the surface. So almost any company ends up creating an implementation similar to the one described in the article. Then it learns that it is much harder than looks, but it is usually too late. So they end up maintaining it for a long time with the original team long gone.
BTW I believe that temporal.io (I'm tech lead of the project) is so far the best open source solution to this problem.
temporal.io our fork of Cadence will have PHP support very soon. Support for other languages is coming. Python and Typescript are the highest priority.
i worked on this site btw - would be happy to receive feedback on the site, particularly if any wording was confusing or unclear!
I get it's something related to workflows, beyond that I have no clue what problem it aims to solve or where or why I'd use it.
It is hard to describe as it is a new way to build distributed applications that doesn't have a commonly agreed name yet.
public void execute(String customerId) {
activities.sendWelcomeEmail(customerId);
try {
boolean trialPeriod = true;
while (true) {
Workflow.sleep(Duration.ofDays(30));
activities.chargeMonthlyFee(customerId);
if (trialPeriod) {
activities.sendEndOfTrialEmail(customerId);
trialPeriod = false;
} else {
activities.sendMonthlyChargeEmail(customerId);
}
}
} catch (CancellationException e) {
activities.processSubscriptionCancellation(customerId);
activities.sendSorryToSeeYouGoEmail(customerId);
}
}Seems to be a fairly straightforward, minimal (which can be good; that's not a criticism) workflow engine.
edit: seems like it is up to the callback owners/authors to deploy and run their own task workers, which presumably implements the individual task logic:
The design is very intentional in driving an ownership model where lambda owners own all aspects of their lambdas’ operations. To promote this, all lambda worker clusters are owned by the lambda owners. They have full control over operations on these clusters, including code deployments and capacity management. Each executor process is bound to one lambda
We have “RemoteJob”a that run entirely on a separate server. The name of the job, and a message (including information about how to run the job) are passed the remote queue which will schedule the job based on priority and availability. Ideal for sending emails, making requests to external services, processing web hooks, updating records in the search engine, bulk deletion, etc.
We have “LocalJob” that run in the same process but are deferred until after the response to the client is finished sending. These can be used for simple things like preparing a payload for a remote job, small bulk deletion operations, sending in app notification to a large pool of users, etc.
Additionally we have a “CallbackJob” that may schedule out a bit of work and callback an API endpoint in the app when there is load available, either to return the value of the job or to trigger additional processing. Our current use essentially been to come back the app when load is available to recalculate some expensive cached values.
B2B/SaaS providers, take note: If a company gets really big, they may be less likely to afford your enterprise offering. This is because the part of the business that is paying for your thing still has a tiny budget. Yet the number of users they have to support is 10x larger than their budget. So to catch bigger fish, you should actually dial down your cost for bigger orgs. It won't seem fair to the smaller ones, though, so probably this should remain confidential.
Companies such as google, Microsoft and aws rely on user’s data. It’s difficult for them to offer privacy.
But with C. Rice on board, I have doubts about its direction.
Getting this when I click the link.
Really? Such an important step to understand even a little bit about what real companies are using and you stopped at Google.com? Its not hard to network and interview with former employees to understand some of the bigger picture stuff in use at large companies.
Dropbox using Python, the real question is what Celery didn't solve for them. My guess would be scalability.
"Scalability" is a great scapegoat for making dubious decisions, but my guess here would be the "task priority" requirement.
[0] https://engblog.nextdoor.com/nextdoor-taskworker-simple-effi...
Today you can containerize and autoscale Celery workers, so having task specific queues and worker groups will solve your resource utilization problems. For even better utilization, you can have gevent and process based workers depending if the task is resource or I/O heavy.
In years I have never seen Celery workers hang, we process millions of tasks (hundreds of task types) every day. Since the post is very old, I'm sure that they had to deal with a few nasty bugs back then.
However, Celery does have real scalability issues: If you run over 1000 worker nodes Celery gossip alone will be well over 1000 messages per second. Having 1000s of gossip queues isn't ideal either given that RabbitMQ is not easily scalable. You might end up with multiple clusters which can be a bit of a pain in a single codebase. I'm not sure if running Celery on SQS solves some of these issues.
Also I am not sure how scalable RabbitMQ is, redis as broker scales quite well but you lose durability when you use redis.