Futures for C++11 at Facebook
code.facebook.com
code.facebook.com
There's been a push to standardize "Promise" to be the read-only side and "Future" to mean the resolver/promise pair, but Folly Futures use the opposite convention.
Agreeing on terminology is hard. :)
Edit: I think it's fine that Folly adopted the C++11 convention. I'm just whining in the open that that promise vs. future may never be agreed upon.
The following is my favorite interview response of all time.
Q: what's a quaternion? A: I don't know... But I bet it has four parts.
Genius! An honest answer followed by a wise educated guess. Now, what's a future and what's a promise? How can you remember the difference? Good luck!
Those terms should be retired forever.
Except of course it isn't. Oops.
The important thing is to always treat the name the same. Generally, I treat types as nouns. So an emitter is something that emits something. It's passed to the routine that consumes the data. On the other hand, a consumer consumes, so it's passed to a routine that emits something for consumption.
Learning these terms isn't that tough, either. I don't find "promise" any less intuitive than "statement", "expression", "parameter", or "argument". I'm just used to the latter ones.
I agree the Scala/C++11 way might make a little more sense intuitively, but it really doesn't help flipping the definitions around.
https://en.wikipedia.org/wiki/Futures_and_promises
The distinction between futures and promises in the wikipedia article matches Scala's terminology exactly.
It's always bothered me that Baker and Hewitt felt the need to coin the term "future" when they were clearly aware of "promise" and "eventual". Their main distinction seems to be that a future is a 3-tuple (process, cell, queue). One can argue that their future meets the "read-only view" definition because cell is only to be written to by the same tuple's process. But they never explicitly distinguish a future as "a read-only placeholder view of a variable," and E was the first publicly-available language to make that very distinction (please correct me if I'm wrong), only they happened to use "promise" in that regard.
In any case, I trust your judgement on the matter more than my own, just wish everyone had picked a nomenclature and stuck with it.
It was posted as https://news.ycombinator.com/item?id=9746522, but we merged that thread with this one so as not to have two on the front page.
Very similar to C#'s TPL.
And VS2015 has experimental support for async/await in C++: http://blogs.msdn.com/b/vcblog/archive/2014/11/12/resumable-...
https://code.facebook.com/posts/225525624316574/building-and...
- Standard blocking code in OS threads.
- Future-based code.
- Coroutines with explicit annotation.
The latter is probably only viable for very large organizations in terms of engineering cost. The other two are doable for smaller organizations.>we provide a flexible mechanism for explicitly controlling the execution context of callbacks, using a simple Executor interface
Why can't we have this instead:
fooAsync(input) {
//...
USING_B_THREAD_POOL
//...
USING_A_THREAD_POOL
//...
USING_B_THREAD_POOL
//...
}
What are the disadvantages?The std::future interface is good, but it has no truly suitable implementation. I'm not familiar with FB's Folly but at a glance I don't see it solving the fundamental problems with both std::future and callback-hell; .then() seems to just move callback-hell into one place rather than having it spread out -- it's still hell. std::future on the other hand is tied directly to the OS thread interface. Are you able to .wait() on say, 1,000,000 items at once? You'd either need 1,000,000 threads or you're bound by the slowest waiting entity in the current thread. It's definitely flawed.
The solution is to once again break down the execution context. Like threads did with processes, a new division is necessary in userspace based on stackful context switching (think boost::context crossed with boost::asio using the yield_context feature). Userspace contexts can "go to sleep" and "wake up" while being agnostic to the bounds of the thread-pool, and working fine on even a single thread. This model allows synchronous-looking programming while really acting in an asynchronous way -- which is what the obsoleted OS process/thread system offered.
"Async" and "callbacks" are really just a model of no-model -- it hasn't inverted the stack, it's basically eliminated it -- and I treat it as nothing more than temporary.
Anyway - great project. I'd love to use it sometime maybe when I get a bigger dev machine.
This is very similar to HPX, the general purpose C++ runtime system for parallel and distributed applications by the Stellar Group - https://github.com/STEllAR-GROUP/hpx.
I recently saw the excellent video presentation by Hartmut Kaiser on this https://www.youtube.com/watch?v=5xyztU__yys and a lot of concepts in folly futures are quite similar. However the most striking thing in HPX was that all the building blocks are serializable, and the presenter mentioned that it is so because you could serialize and move a thread to a different machine and run it there.
Same talk maybe?
I didn't realize Posix threads were that inefficient?
First, I can't see why the memory overhead would be anything more than having a separate stack and entry in the TCB?
Second, can you not just save the stack pointer, program counter and registers into the TCB, then do the reverse when restoring a thread?
Finally, I wouldn't have thought you'd have too many TLB misses either, thus leaving the only expense being trapping to the kernel to switch between threads?
Can anyone explain?
Relatively speaking, I don't think context switching is _that_ expensive compared to other areas you can focus on like doing smarter memory management.
I don't know exactly what they mean by threads having heavy memory overhead, but possibly they mean cache interference mentioned in that link? I'd be curious if there is an actual large memory chunk other than stack/scheduling details in play here too.
On modern cores too, there are a decent chunk of registers, including floating point registers. I've looked into timings before for embedded applications and the performance hit here isn't trivial when looking at interrupts and the like; but I'd be surprised if it was that overwhelming on servers.
For example, let's say I type the string "123". If there's a future / promise / async whatever anywhere, we risk processing the characters out of order. We may not catch that in testing, because it's usually fast enough that it comes out in the right order.
> It is more efficient to make service A asynchronous, meaning that while B is busy computing its answer, A has moved on to service other requests. When the answer from B becomes available, A will use it to finish the request.
What if the first request is "set sharing permissions to private" and the second request is "upload this compromising photo?" Obviously it would be bad to process these out of order. How is this handled?
Say we are making a Todo list app. We have a list of Pending and a list of Completed todos, and when the user completes one, we move it from Pending to Completed. But how do we ensure other threads can't see the transient in-between state? Traditional concurrent programming might solve this with a lock:
lock()
Pending.remove(todo)
Completed.add(todo)
unlock()
This problem is very well known, and any discussion of threads will spend a lot of time on locks, queues, serialization techniques, etc. for avoiding races. Now, with futures: Pending.remove(todo).then({Completed.add(todo)})
We've got the analogous race condition, even if we're single threaded. But articles on Futures never seem to discuss techniques for mitigating this. Why not? Is there a Futures equivalent for a lock?If you're writing in C++ with shared data structures accessible from multiple threads, you need locks anyway; futures don't change that. If you're writing in Javascript, your code is inherently single-threaded and no locks are necessary.
This is usually up to the API you are using. You may not have a choice.
> your code is inherently single-threaded and no locks are necessary
Locks are necessary in the single threaded case. See my example:
Pending.remove(todo).then({Completed.add(todo)})
Nothing prevents another operation from executing between the remove() and add() calls, and seeing the transient state. You need the analog of a lock to prevent that. What is that with Futures?In the browser world, futures or promises are an abstraction over some operation that doesn't block the UI thread, and therefore can allow some other work to happen in the meantime. In your example, if Pending and Completed provide an interface to some remote API, then a lock provides no benefit because making separate RPCs can't possibly be atomic anyway. If they're operating on local data or the DOM, then making them return futures is pointless because the work will happen on the same thread, and no other code could possibly observe the intermediate state anyway. (This is a core principle of the browser event loop: Javascript code is never preempted, it can only yield control by running to completion. Apologies if I'm repeating stuff you already know.)
In other languages, futures are more flexible because they can contain CPU-bound work that operates on shared memory. In that case, futures don't magically absolve you of the need to protect that shared memory. But if you have multiple operations on the same shared data, it once again doesn't normally make sense to decouple them with futures in the first place.
Futures and promises make more sense in functional languages because they don't have a shared memory model, which avoids having to share resources between parallel tasks using locks. Google dataflow programming for more insight into how lazy evaluation and referential transparency parallelize work when there's no mutable state.
Upgrade the user wetware. Have them wait until confirmation that sharing permissions are now private before beginning upload.
val upload = for {
p <- Future(setPerms(private))
u <- Future(uploadPhoto(photo))
} yield u
upload onComplete {
case Success(uploadStatus) => // do stuff
case Failure(why) => // do stuff
}
These will not happen concurrently unless the futures are declared outside of the for comprehension. u <- Future(uploadPhoto(p))
right?> The SU service fetches candidate accounts from various sources, such as your friends, accounts that you may be interested in, and popular accounts in your area. Then, a machine learning model blends them to produce a list of personalized account suggestions.
> This enabled us to reduce the number of instances of the Suggested Users service from 720 to 38.
Technically all tasks can be collected in a list and then you just cancel everything in the list but I have yet to find a scenario where I've needed this.
https://github.com/twitter/util/blob/master/util-core/src/ma...
I'm sure there is a good performance reason they haven't done this yet. I don't know enough about the design of programming languages to know all the complexities involved. I'm just idly wishing.
To my knowledge, our C++ Thrift code is here: https://github.com/facebook/fbthrift
(I work on Presto and occasionally Swift at Facebook, but am not at all familiar with modern C++)
void run_and_report()
{
something_long_running()
.then([=]{ report_results() });
} // blocks here until the then() block executesI suppose they also have caching built in, but that is easy to build on top of FRP.
"One question you may ask yourself, is why RxJS? What about Promises? Promises are good for solving asynchronous operations such as querying a service with an XMLHttpRequest, where the expected behavior is one value and then completion. The Reactive Extensions for JavaScript unifies both the world of Promises, callbacks as well as evented data such as DOM Input, Web Workers, Web Sockets. Once we have unified these concepts, this enables rich composition."
When people are trying to do fancy streaming stuff with futures and come to me for help, I generally recommend they do the streaming with Rx instead. (eg rxcpp from Microsoft)