Async might be a fad
cs.oswego.edu
cs.oswego.edu
C#, F# and coffeescript have excellent syntax that remove the line noise caused by writing async code. Actors and channels also nicely remove line noise from async programming.
If anything as computers become more powerful and distributed you'll see threads and locks disappear rather than async computation. Async computation in an imperative style is what most want, threads and locks are what we have.
Of course, C# has great async support; but it is still a complicate thing that programmers must be very cautious about when applying.
Sync:
let file = File.Open("foo.txt")
let data = file.Read(8192)
// Do some compute stuff with data
ASync: let! file = File.OpenAsync("foo.txt")
let! data = file.ReadAsync(8192)
// Do some compute stuff with dataBut yeah multithreading is easy.
https://idea.popcount.org/2013-09-05-it-aint-about-the-callb...
Basically, it's easy to show that callbacks are a much harder paradigm to work with considering flow control.
That's it. There is a place to use callbacks, but if you need anything that is not trivial and won't blow up at some point, you should use threads. Greenlets, processes or whatever you call it, things with stack that take time to context switch.
I strongly believe threads are better, if not anything else is due to the fact that when you do "spawn", you make an explicit statement, saying: here we demultiplex - programmer beware of flow control here!
I just change
serve request
to forkIO (serve request)
And magically I have multithreaded aynch IO that seems to be extremely performant.Also, the discussion is in the context of imperative programming.
Multicore processors and the fact that single-threaded performance basically hit a wall explains why an N:1 threading models in something with the use cases of Java fell out of favor. M:N, while still "green threads", has a somewhat different set of trade-offs versus 1:1 native threads than N:1 does, though.
So I think that's the future we're looking at: most code async, with the likes of LAPACK and some ultra-low-latency I/O libraries (hedge funds and the like) being the only users of traditional heavy threads.
There are TCP buffers. If it gets full ya maybe you'll wait a bit but it's thread vs rope relative to blocking a whole big roundtrip data exchange while blocking a thread. Nobody really tunes the write buffers because it doesn't slow down our apps. That is not the case with calls to remote services.
Note that TCP buffers are pretty big so the chance of the thread blocking while writing are very low.
The Greenhouse framework in Python is one of them (http://teepark.github.io/greenhouse/master/). The docs (linked) have a concise discussion of the various approaches to parallelizing IO operations with pros and cons of each, and, for most database applications, an obvious winner.
Perhaps the callback mechanism of Node is a fad—continuation based techniques can make async code look like non-async code. Perl's Coro and Ruby's new(ish) fibers are examples of how that could look.
Aren't computers actually synchronous? The abstraction on the OS layer is what makes computers async.
At the programming-language layer, async means you can write programs based on how the user and your program interact with each other. Other applications (and the OS itself) behave asynchronously too. Why make it harder?
Of course it shouldn't have to be that way (it didn't use to be), but it's very convenient to think around "what is actually happening" instead of "what the synchronous execution of the program is doing".
EDIT: but now I've seen you only meant a specific subset of async. I'm leaving the comment here anyways.
Proxy deals with slow client delays. App server serves app requests at speed. Tune number of proxy connections to keep app servers reasonably busy. Scale each layer individually.
IMO Whether the proxy is async or not is a matter of taste.
I'm looking at the port of Quasar to Clojure as my sole reason for looking at Clojure over Erlang.
I think if done right, this will bring down costs for scaling. I recently looked at Linode and AWS. My first thought was holy crap! To get a small 2 GB Ram box is rather expensive. It would be cheaper to just pay for a business internet connection and host my own boxes. If I had a spike in connection with normal threading on a small box (my laptop with 16 GB), Java would start to sputter around 200 threads due to context switching. Connections would be rejected. Quasar server would accept the connections, just have a high latency. I could live with that.
The message was "You are all super late to the party lol".
And sorry anyone who wants to claim that Java's green threads are somehow a better programming model than async IO ala node.js + promises is pretending to write code. Yes async IO is not easy, certainly nowhere near as simple to manage as process / fork but with consistent coding style your can still end up with a system that behaves predicably and most importantly can be reasoned about.
Meanwhile I bitterly regret the days and weeks of my life lost to debugging threaded code. Never again!
Node has a very beginner-friendly primitive for doing this.
People in general are really bad at grasping the concept that $programming_paradigm are not all equally suited for a given problem.
Object Oriented, functional, async, strict type system, etc.
I've noticed an almost evangelical nature of people in trying to push their prefered environment on others. Guess what? Node.js is not the answer to everything (and I do 90% of my work in node!). But you know what? Neither is OO, or functional programming. There may be many advantages to your language of choice, but that doesn't mean it's always better.
Fetching resources from many different areas - especially external or unpredictable (3rd party) services - that has async written all over it, and node makes that easy.
Web request maybe, database definitely not. Async means you put a high load on the resource you connect to. In the context of a database that makes absolutely no sense. In fact, it's much more common that if you have an async app that you have a connection pool that gives out limited number of connections. At that point you might just use threads.
For example, in Go, when you read a value from a channel, it's just like a good old blocking call, as far as the programmer is concerned.
On the other hand, an "async" read would involve callback, promise, or some other constructs.
heart-beat - yes we'll need concurrent threads for handling concurrent requests in the blocking world; the question is whether this will result in too many threads, which depends on the application.
Sure. But how do you handle that? IO results need to be communicated back to the UI thread somehow.
Throwing together a whole new thread+stack against a named method that communicates back via explicit messages (or however you want to get the results of IO back to the UI thread) seems like a bit much if all you want to do is download a motd.txt and stuff it in a label.
The hard part is, if the modifications do not form a single, predictable, serialized chain, how can the programmer reason about them? This problem is independent of async/sync. If you use C# async for a UI action, you still need to worry about it.
It reminded me of the "I know, I'll use regular expressions" joke.
Here we have an inherently concurrent problem - user actions and some IO actions occur concurrently. That problem cannot be reduced by some API or language trick.
Definitely not my experience. jQuery is very popular among unskilled programmers because its async model is very easy to understand. I use it to teach async!
> That problem cannot be reduced by some API or language trick.
Why? Reducing problem by API or language tricks is exactly what abstraction is for.
Some UI APIs aren't thread safe and can't reasonably be made thread safe. They may even be kind enough to check the calling thread and intentionally throw exceptions (or otherwise abort) if invoked from a non-UI thread.
Hilariously, I've had to work around race conditions in such an API. Just because my access was forced to be on one thread, doesn't mean it's implementation was!
> The hard part is, if the modifications do not form a single, predictable, serialized chain, how can the programmer reason about them?"
Reusable async patterns can help you knock off the "single" and "serialized" bits of the microcosm you care about, leaving you to grapple with only the "predictable" one without the pointless "easy" boilerplate distracting you with additional complexity, line count, and bugs.
> If you use C# async for a UI action, you still need to worry about it.
Agreed. If you just spam the async and await keywords without understanding what you're doing, you're not magically going to get the benefits of a single, serialized chain of events. And they're not going to turn the inherently unpredictable response timings and contents of a series of web queries into a predictable one.
It is a powerful programming model that unfortunately quickly devolves into spaghetti code if not carefully maintained, but properly done it is quite nice and alleviates a ton of worries about synchronizing threads.
For example if I have a honey pot that collects random events, the clients can just send the data to it without expecting a result (IE: I don't care if it's successful or not) and honey pot is not expected to write any sort of response.
Clients send data and move on.
I don't think people who know what they are doing, use async calls for mission critical operations.
If async call returns response and client is required to read the response it's no longer considered non-blocking from technical perspective.
What I do write is lots of C/C++ client/server applications for HPC/HFT/DC workloads where speed both in req/sec (throughput) and speed in min/avg/max(secs/req) (latency) matters. In these environments I almost exclusively use non-blocking I/O. There are several reasons:
1) Threads are not free. Even if you use a thread-pool to avoid spin up costs, context switching overhead matters. Every time you call blocking I/O, you make sure that the kernel will wake up, schedule another thread, and do anything else that it decides to do. Waste time that you could have used to do useful work. Non-blocking I/O puts you in charge of your own "thread scheduler". Your "threads" are functions, they are "cooperatively scheduled" and you can make full use of every cycle that you get.
2) Programming with threads is hard. Trust me. If you think it's easy, or I'm soft, you haven't done it enough. At some point you will need shared state across those threads. And then you'll need locking and unlocking. (also Mutexs are slooooowww) And then you'll need to handle error cases, and you'll need to make sure that all the unlocking is done right in all of the right places. And then you'll need signaling between your threads. And you'll need semaphores or similar. And 3 months down the line, you're thinking to yourself, when a foo exception causes a bar signal, will a baz handler deadlock? Will it make progress? Humans just aren't designed to reason about this sort of thing.
With a single threaded, non-blocking design, it's really easy to reason about exactly what is happening with all of your state. Debugging is obvious and straightforward. This is necessary if you're like me and don't write perfect code first time. There's only ever one function accessing shared state at one time. The "scheduler" is working for you, not against you. If you write your code simply, cleanly and efficiently, you'd be amazed how much work a modern CPU can really do. Honestly, once you've saturated a 10G NIC what more do you want to do?
3) If you buy into the non-blocking design, then, as long as you only use 1 process/thread per core, almost anything a thread can do, a process can do better. Threads have no memory protection, anything you touch probably belongs to some other thread and you're inviting subtle bugs. Processes have memory protection by default if you want to share things you can do it explicitly via safe mechanisms (shared memory rings, pipes, IPC etc). Shared memory rings are (can be) so fast that data is more or less local so if you want to use shared state from a TCP connection or whatever, you can always "dispatch" work to another process to do it for you. You get the benefits of many cores working for you as well as a clean and obvious programming model.
Ultimately, if the question is one of syntax, then I'd happily believe that JS has some ugly syntax for doing these things, but if the question is one of design, then you should think really really hard before deciding that a threaded model is the correct one for you.
Encapsulation is your friends. If each thread deals with a non-overlapping (set of) object(s), then many of the issues are gone. What is left is messaging between a thread and the process, which can be done using a thread-safe queue.
I think the view that processes are "heavy weight" is a dated one. From an OS point of view, processes are pretty much the same amount of "work" as thread. Each has a context, each needs to be scheduled. Processes do have some extra state (notably the TLB context) but modern machines are very good switching these. Spinning up processes is somewhat more expensive than spinning up threads, but you really shouldn't be doing either on the critical path.
I agree that everything is context specific, although my main application area is really low latency scenarios (handfuls of microseconds) and techniques that work well there tend to port well to slower situations pretty easily (at least in my experience).
I disagree that encapsulation and non-overlapping objects will save you from threading nightmares. Every design starts out with clean boundaries and beautiful abstractions. Every design ends up in spaghetti soup. It's just a matter of how long it takes to get there.
Encapsulation saves you from threading nightmares. Your spaghetti notwithstanding. I've done this over 20 years and several startups, writing entire application environments on everything from embedded to desktop to server, and it works fine. Try it.
I'm very surprised that many people here argue that async code is much better to understand than sync code. Ok, so that part is subjective, and let's file it under personal preference. For people who love synchronous coding but fear the cost of threads, I'm trying to make an argument that the fear is probably not justified.
My reaction is due to my experience which is that threaded programming is something that's very hard to get right and especially to maintain. Async programming cleans up the threading and makes it kind of implicitly cooperative.
I was involved in a big move of some core infrastructure from a multi-threaded design over to a pipeline of async style apps. The result was a huge boost in productivity and debugability which worked out really well for the company.
It might be an idea to look at this 'empirical' data and figure out which webservers use forking/threads and which use events/async, then you may realize why the high concurrency webservers took the 'silly' route of avoiding threads.