Pitfalls of Callback-Based APIs
thelig.ht
thelig.ht
That is why programming is hard. Programming takes thought, design, planning, and why programmers aren't a factory line where you can just add a bunch in. Programming is solving tough problems and making them easier using good balance for time, budget, maintenance, integration and more.
Careful nuances, balances, structures, flow and many other things can be disrupted if programming is treated incorrectly or with lack of care from too many cooks, too much pressure, or not enough prototyping. This problem will always exist, good individuals and teams will always know there is no all encompassing fix from some new tech, there will always be the balance that determines success using all advancements at your disposal that help you and consistently improving that. The design has to continually keep up with the evolution of the project.
> This article isn't really about callbacks and it isn't even really about APIs.
Well then, perhaps you should have named the article something else...?
The essential problem is that any program which can incorporate arbitrary recursion is subject to the halting problem. The seemingly simple initial fixes, like a lock or asynchronous callback queue, can improve things. The author proposes third-party code analysis tools to help; I'm less sure of that. The Gödel dragon is not easily pushed back, and we often accidentally create Turing machines.
I think the title's fine. It's better than "Callbacks and the Halting Problem", since most of the essay covers, after all, "Pitfalls of Callback-Based APIs".
Yes, further automated analysis can find some of these problem. That's what the author of the essay proposes, and it's good to know that there are some solutions.
Let's go down the road further and expand the system so there are multiple loggers, and where subsystem logger configuration via an external configuration file occurs at run-time. Some configurations are acyclic, others are not.
This could be prohibited by edict - if a program cannot be analyzed then it's not valid. As the essay points out, there are also formal verification techniques - and that very few people use them. As I pointed out, it's surprisingly easy to make a Turing machine by accident. (See http://beza1e1.tuxen.de/articles/accidentally_turing_complet... and its HN comments https://news.ycombinator.com/item?id=6577671 .)
In practice then, it's very unlikely that most complex software will be able to use these tools, because it's hard to restrict the solution space to what those tools can analyze.
Here there be dragons.
Logging is weird and special, running at unusual times in unusual states, and therefore has unusual requirements.
So, yeah, having a pile of arbitrary functions that your logger can call is not really going to work. But if this were a "notify when a comment is posted" example, the author would really struggle to find problems with that same approach, because that's a much safer operation that happens in a more predictable way than logging is.
The only real takeaway from this article is "be careful with your logging implementation," but hopefully you already knew that.
I don't think the "Pure Solution" mentioned in the article works. The only effect of a pure function is to compute its return value, so not only is it safe to invoke pure functions as callbacks, it's also perfectly useless — you can't tell whether it even got invoked or not. (Rian in another thread argues that qsort's comparison function is a "callback", but I don't think that's what people usually mean by "callbacks".)
I've always considered function arguments to functions like qsort() as callbacks, and from my interpretation of Wikipedia it seems to agree. Now I wonder if my interpretation has been overly general or if a more specific interpretation has become more commonplace.
If qsort's comparison function is a callback, is that true of any function argument to any higher-order function? If so, why the additional term? If not, what makes a function passed as argument a "callback"?
The message pumping example the author gives has the exact same issue. The core here is that two items (database & logger) are mutually dependant. That will always be an issue. It doesn't matter if your promises are circularly dependant, your context objects, your callbacks, or your globals.
Static analysis won't help - fully resolving circular dependencies is equivalent to solving the halting problem, IIRC.
And so the only way to prevent those dependencies is to constrain the design in a way that prevents circular dependencies. That's where the ad-hoc rules mentioned in the article stem from. They artificially restrain the problem space to prevent (classes of) circular dependencies.
The point of the article is that callback-based APIs like this obscure the actual problems (like circular dependencies).
You're right that constrained designs prevents issues like these but callback-based APIs aren't inherently constrained. They allow anything to happen, which is why they conversely encourage errors like these (and are subject to pitfalls).
typedef enum {
CONSOLE_LOGGER,
DATABASE_LOGGER,
/* etc. */
} LoggerType;
void add_logger(LoggerType);
In this API you're constrained by what loggers you can add. It's not the totally unchecked free-for-all that a callback-based API provides. The user is strictly unable to shoot themselves in the foot.But this limits the expressive power that callbacks provide. Sorry I don't have any other API recommendations. The only way I can think of to stay expressive while safeguarding against unintended abuse is to include code analysis.
> fully resolving circular dependencies is equivalent to solving the halting problem, IIRC.
I don't think so. Any good dependency injection framework would detect this kind of circular dependency.
Just analyzing the dependency graph of injected components is indeed just a topological sort.
Is it realistic to think that any app starting with a bunch of assumptions about being single-threaded can ever just "become multi-threaded" without bloodshed?
It also makes them UNIVERSALLY USELESS. Think about it: a function that definitionally has no side-effects is being called solely for its side-effects.
The article using that to wrap up makes it the equivalent of a complicated category theory proof that successfully proves that the empty set has TONS of amazing properties...
The article doesn't advocate always requiring pure callbacks. It's just offered as one possible way to make it easy to reason about the correctness of using an callback API.
The point that the article is trying to make is that because callback APIs don't expose their correctness requirements and because their correctness requirements are externally defined, they encourage programming error like this.
I won't shed a tear if callback-hells are replaced by proper async APIs (like in Stream and Future-based dart:async, or the Thenable in the new JS standards), but this article seems to be misalinged.
Downvoters: care to elaborate?
The essay is about callback-based APIs, so types and encapsulation are out of scope from the start. It then expands the scope and observes that OO principles don't address the points covered. For example, it says that one of the "magic" requirements would be a way to state: "Do not add a callback that calls log() or acquires any locks held while log() is called." There is no type for that.
Other than pure functional code, there's no way to do that.
The essential equivalent for a multi-actor system is deadlock prevention. OO principles don't help there either.
You can get the problem with a stream. Consider a logger stream, where the listener opens a database connection to save the value then closes it, and the database adds 'open' and 'close' events to the logger stream. This will lead to an geometric explosion of events on the stream, because nothing at the API level says you shouldn't put those pieces together that way, other than the documentation.
I could go into the OO details of how this stream might be implemented in Dart, but really the OO nature of the stream API obscures the essential self-referential nature of the problem.
You can propose an equivalent counter-example if you want to demonstrate that those APIs really do solve the problem. I happen to agree with the well-written essay, and OO principles or "proper async APIs" solve nothing.
We do some crazy async HPC code, and during on-boarding, new members invariably get some initialization / error handling wrong with raw callbacks, and convert pretty quick to FRP after that. Unstructured async is crazy.
I don't understand the "change the logger" comment. The example was the logger sends a message to the listener, which opens a database handle, which sends a message to the logger, and repeat. There's nothing to changing. I don't see how FRP helps eliminate that cycle.
//try making the logger first
var logger = require('myLogger')(backends);
var listener1 = require('mySync')(logger);
var backends = Rx.Observable.fromArray([listener1, listener2]);
//try making the backends first
var listener1 = require('mySync')(logger);
var backends = Rx.Observable.fromArray([listener1, listener2]);
var logger = require('myLogger')(backends);
From there, you'd switch to unsafe FRP methods (imperatively injecting into event streams, e.g., Subjects in Rx), which is the warning sign of weird cyclic behavior.
Or just SSH to a server and run tcpdump without arguments :)