Local State Is Poison (2012)
awelonblue.wordpress.com
awelonblue.wordpress.com
What, and why would you want to upgrade during a loop?
>Those concepts are replaced by discovery – potentially in an infinite graph of stateful resources
What?
>To achieve large scale, robust, resilient, maintainable, extensible, eternal systems,
So many buzzwords...
Could someone give an example with before/after-pseudocode. I understand neither what the problem is nor the how the proposed solution is expected to solve it.
The entire tree concept he is talking about seems like a misplaced faith that you can achieve that kind of separation.
My gut says that in moderate to large sized programs so much state would end up under the root, just out of laziness, time constraints, over-complexity or lack of programmer skill, that you'd have the worst of all worlds.
I just caught the end of the global variables era and have worked on a few programs where the state of things can get modified anywhere.
It is not pretty.
Not pretty is an understatement. It's ugly at best, and rage inducing most the time.
See Erlang for a trivial example.
One of my personal policies is to not introduce any new acronyms in the work I do, and especially don't redefine ones that are already familiar in the programming field. Just come up with a proper descriptive name, even if it's slightly longer. Let the acronym follow once the name is already known.
Seeing as how your friend worked for the EPA, maybe I shouldn't be surprised that this phenomenon often seems to surface within organizations whose name is typically given in the form of an acronym...
VisualWorks Smalltalk VMs had this ability, provided you were a bit careful about how you modified the system. The system would update atomically, and if you were doing things like adding instance variables that safely defaulted values, the system could update in the middle of transactions, then complete them with the new code base. There was the caveat the adding an instance variable to a class meant that every instance of the class basically had to be recopied with the new slot, which could potentially pause the VM for a significant span of time while it performed this operation.
There are embedded RTOS that have similar capabilities. Sometimes, you just want to get very close to zero downtime.
One of the reasons global state gets a bad rap is because we're always trying to minimize side effects, and that's a worthy goal. The idea not only makes sure the program's components test well individually, it also enables better component re-use (although I'm still convinced that re-use without refactoring is mostly a myth in practice). However, in striving to ban explicit global state we have developed an astonishing array of cruft and complexity which counteracts the effects we wanted to achieve in the first place.
I agree with the article that well-defined and clean global data makes a lot of code easier to handle and it also eliminates unnecessary work, both on the human and the machine side. Of course, this idea breaks down again when the global substrate becomes muddled and structurally broken. At which point we've come full circle.
In my opinion a mixed approach is probably advisable for most projects and in fact, that's the way we already do things in many cases, even if we're forbidden to use the actual phrase "global variables". Judicious use of both paradigms yields the best results in my opinion. Maybe it's time to actually start calling global state by its real name, without expecting to be stigmatized for it.
A common design pattern in clojure is to have a single massive data-structure that represents all application state, but to take it apart recursively whilst updating so that each function is only passed the pieces of state that are relevant. That way you can look at a function and immediately know what it can and can't read/write out of the global state.
A nice way of doing this in an imperative language might be to pass bits of the global data structure by reference (ala lenses) so that a given function can only read/write state contained in its arguments. The data is still global in the sense that it is all accessible from the root object and is not encapsulated, but access can be restricted on a per-function basis.
That's true. The equivalent effect in the "all local" paradigm is local objects the state of which can't meaningfully be understood or manipulated by neither the programmer nor other objects, leading again to unexpected behavior that is painful to track down.
A lot of the precautions that are common sense when working with global data are already instinctively followed (or at least understood) by most programmers I think. In effect the lenses are already working when there's an understanding when and where manipulation occurs. This mostly coincides with the notion that complex data should be manipulated by a well-defined model and wherever sensible there should be only one mechanism for doing it.
def foo(x):
x_plus_one = add_one(x)
x_plus_one_times_two = times_two(x_plus_one)
return x_plus_one_times_two
Here is a function with some local state. Which part do you feel "can't meaningfully be understood or manipulated by [...] the programmer"? I'm not trying to bait you, but I'm struggling to grasp the argument you are making.Abstractly speaking, when I say "A often has the property B", you can't meaningfully counter that with "you're wrong, here's an example where !B".
Fortunately the confusion is cleared up now.
[1] https://en.wikipedia.org/wiki/Referential_transparency_%28co...
def foo():
a = []
append_foo(a)
return a
def append_foo(a):
a.append('foo')
But then I wouldn't call this "local mutability", in the sense that you modify your list in a different part of the code. Not a problem in a trivial example like this, definitively an issue when the actual action performed is ten levels deep.You can't return locally created variables [beyond primitives].
Beyond simple primitive reading/returning functions, you often create some arrays or structs, then pass a pointer to the necessary arrays and structs to your functions, so it can read and write them.
This way the functions can stay pure.
This is just free vs. bound variables. You're absolutely right that it would be a pain to write things such that there are no free variables from parent environments. However, those functions aren't necessarily from the top-level (global) environment.
>However, in striving to ban explicit global state we have developed an astonishing array of cruft and complexity which counteracts the effects we wanted to achieve in the first place.
I think it's important to avoid global, mutable state in most cases. Relying on functions and constants from the top-level environment is fine, but storing a program's state in the top-level environment is just asking for a whole lot of painful debugging.
Our brains struggle to reason about how state evolves over time. Add in concurrency, and the problem easily becomes intractable. On top of this, testing stateful components is burdensome.
So, the state needs to be kept as separate as possible from the complex algorithmic logic of the application, so that the state-handling-parts can be kept as simple as possible, and the complex parts can be kept as easily-testable as possible. If this means that the state is handled globally, then fine, but it is not really about where the state is held, but rather about how easy is it to reason about and test.
My rule of thumb is this: we should be able to test our complex mathematical and algorithmic components as stateless (pure) functions, independently of any stateful parts of the application. The remaining stateful parts of the application should have a simple and well understood lifecycle, preferably well away from any concurrency, and with tightly controlled and documented state transitions. (OOP is handy for this, although it must be kept on a tight leash).
A first stab at the "tree-shaped resource space" referred to in the article would be the abstract syntax tree of the program itself. Each node would have a unique URI, which can be a physical directory path on a filesystem or can be stored in a database structure. Every local variable would be defined by a path in a flat namespace. The "parent" directory of the local variable would be the function it is defined in. Security rules can be created that simulate many of the features of variable scoping rules. The most basic rule, that variables are only visible within the scope of their parent function simply means that the only variables that can be referenced within a function are those within the same directory. Again, none of this seems too difficult to implement especially if your language uses a LISP syntax.
I'm tempted to implement a toy version of this, if for no other reason, because I've been wondering about good ways to serialize the program state of a DSL that I've been working on. Performance seems to be the big problem with using a database or filesystem. A global map of URI's (that is easy to serialize) with some sensible access/permission strategy doesn't seem too bad and could be transparent to the developer.
http://pl.atyp.us/wordpress/index.php/2005/07/explicit-state...
Briefly, the kind of global state that's needed for debugging should be easy to find "from outside" - which precludes local variables along with other common idioms. It might still be distributed, and that can still be problematic, but the key point is that many ways of avoiding global state are worse than what they avoid. Global state itself is not the problem; inadequately contained or constrained changes to it are, and there are other solutions besides elimination.
“local state is good, global shared state is bad” and all those (very easily over)simplistic kind of thoughts are like pain killers. They might alleviate you in a moment of affliction but they can be also be addictive beyond the point of benefit. In that regard, yes, something could be poison.
Your line of thought here will make you converge to invigorate some kind of proceduralism. Sorry I don’t know what your domain problem is but you seem to be experiencing an object oriented overhead that you feel like starting to hurt.
You can go ahead and proceduralize things (functions against a remote datastore) but I wouldn’t be so fast in questioning the object design fundamentals. I'd try harder* to remove the original painful overheads or whatever real pain is in your design.
*by harder I don’t mean to be muscular or that you aren’t paying effort to it. Harder could mean to do something as easy as asking to hacker friends to use their fresh unbiased view for a problem/code review.
Listen all, pay attention to some, then ignore everybody (including me)
The data in the store can still be protected (with access tokens or existential types or whatever) so that an item can be only accessible by parts of the code that have the "key".
http://c2.com/cgi/wiki?GlobalVariablesAreBad
State, in general, at any scope, can make things difficult no doubt. But global state is classically bad because it's hard to reason about across large chunks of distributed code, pollutes namespaces, creates concurrency nightmares, etc. Reducing the scope of state to a manageable range of, say, less than a dozen lines of code, into short-lived references, is clearly far better than the alternative and a reasonable approach in the vast majority of cases. Calling it 'poison' ratchets the rhetoric way beyond the gravity of the problem. No, local state is not considered harmful.
What it seems OP is really talking about in practical terms is pure stateless programming, where the application has no implicit or explicit references to a value whose authoritative data resolves in main memory. If you were to tell me the only state you have in your application is Filesystem or database data, "just beyond the edges of our program logic," I'd say you'd basically achieved the fabled 'stateless' programming ideal, long held as a kind of Holy Grail of functional application development, and as OP points out, that's not often achieved even in the strictest functional environments.
I don't want to diminish the points made, the article was instructive to me as yet another anecdote about the perils of shared mutable state at any scope. But the fundamental principle, that one should avoid shared mutable state as much as possible -- which is the upshot of the essay -- has been axiomatic for quite some time.
It's actually a pretty common design pattern in clojure to keep all application state in a single datastructure (eg http://www.chris-granger.com/2013/01/24/the-ide-as-data/ http://thinkrelevance.com/blog/2013/06/04/clojure-workflow-r... http://channel9.msdn.com/posts/Rich-Hickey-The-Database-as-a...).
> But the fundamental principle, that one should avoid shared mutable state as much as possible -- which is the upshot of the essay...
I think you missed the point. The OP is arguing that even non-shared mutable state should not be encapsulated away but should be accessible from some root data-structure. That way you can eg serialise the whole state of your program and restart it elsewhere or traverse the state with debugging and monitoring tools.
He points out that the traditional evils of global state (unrestrained mutation, non-reentrant code) have been solved in filesystems and databases and that those solutions could equally be applied to keeping state in-memory.
In other words, separate data from logic and keep all of your data in one place (whether that be a database, file-system or some well-controlled in-memory structure).
The contribution isn't the non-novel insight that mutable state is bad, it is that Awelon is attempting to actually solve the problem of how to remove "program state" from a program.
Maybe not to eliminate the local state itself, but definitely for the organisation / state sharing / layered security / persistence and many other things he listed.
What's interesting about the OP is the idea of making that the only source of state, so that every other language construct is a pure function of its inputs.