It's hard to beat closures for elegance. You say: generate a link, and if the user clicks on it, run the following code. There's so much less to think about. In a closure, state is implicit: if you refer to a variable defined outside its scope, its value stays around. When you switch to putting state in urls, you have to explicitly encode and decode each bit of state you'll need. You also have to worry about bad things someone might do by sending requests with different arguments. It's sort of like the difference between dynamic memory allocation and explicit malloc, times 10.
[edited for grammar / clarity]
First make sure that the lambda doesn't have free variables by adding the free variables to the argument list. (lambda (x) (+ n x)) => (lambda (x n) (+ n x)) Because the resulting lambda doesn't have any free variables you can turn it into a global function. (lambda (x n) (+ n x)) => function#345 and somewhere else you (define (function#345 x n) (+ n x)).
Next you store the values of the free variables in a hash table (use the memory location of the data as the key for the hash). Now you have a string representation for this closure. The representation contains the function's name (function#345) and the keys of the free variables in the hash table.
An example. Assume that users are represented by a cons cell (cons user-name karma). We're writing a page that has a link called "increment karma". When you click this link your karma is incremented.
(let ((user the-current-user))
(display-link "increment karma"
(lambda () (set-cdr user (+ 1 (cdr user))))))
First transform this to: (define (function#345 user)
(set-cdr user (+ 1 (cdr user))))
(let ((user the-current-user))
(display-link "increment karma"
(make-closure function#345 user)))
make-closure in pseudo code (assuming there is one free variable/argument): (define (make-closure fn arg)
(add-to-hash *free-variables* (memory-location arg) arg)
(add-to-hash *functions* (memory-location fn) fn)
(concat (to-string (memory-location fn)) "/" (to-string (memory-location arg))))
When the server gets a request it parses the string representation of the closure, looks up the free variables and calls the function. (define (invoke-closure closure)
(let ((fn <lookup thing before "/" in *functions*>)
(arg <lookup thing after "/" in *free-variables*>))
(fn arg)))
This solves the problem of allocating closures but the global hash table keeps growing. This may not be a problem in the average case. Consider the usual data in free variables: front page links, users, etc. These things have to be stored on the server anyway. For temporary data that isn't stored permanently you can use a seperate hash table that is cleaned up every x minutes or after it gets too big (just like the closures you're storing now).What do you think? I'm sorry if this explanation isn't clear. I'm not good at explaining these things, especially in english.
You could write a whole PhD dissertation on this topic. My gut tells me it probably wouldn't be worth it. Closures don't feel like something one should expect to be representable. But I'm not 100% certain.
A way of serializing closures such that they kept bindings to symbols, numbers, strings and conses in the local environment would probably be pretty useful. (On the web, I think you usually want to preserve fairly simple values - usernames or whatever.)
map-closure would make it easy to do at the language level, too.
It may actually be possible to include as an axiom a few functions that allow a closured function to be destructured.
For example, it may be possible to define a set of functions which can be applied on a function to determine if it encloses any variables or not, what number of enclosed variables, and the variable values (if not the names).
It would also be possible to add as an axiom a function to extract the "code part" of a closure, at least as an opaque (but uniquely-comparable) object. Speaking as an implementeer hacking on the low-level side, generally the implementation of a closured function is simply two pointers: a pointer to the enclosed environment, and a pointer to the code (of course, in arc2c it's just an array where the first entry is a reference to the code and the succeeding array entries are the enclosed variables).
Only the pointer to the enclosed environment needs to actually be considered: the code is constant (except for code updates, but one would expect updates to be rarer than actual invocations of the code).
If we can extract the data of each closured variable, we can determine if it's trivially serializable (e.g. strings, numbers, proper lists of strings and numbers), and if so, it is potentially possible to encrypt the values into the URL.
We can then also add some functions to reconstruct a closure, given only an opaque code reference and a bunch of data.
Since my Arc implementation, SNAP, will require the ability to serialize and deserialize closured functions (including code and data) anyway, I'll have to extend Arc that way.
The advantage of this is that you don't need to explicitly state what you want to close over: you just close over variables, and the base system will destructure your closure, and just encode the trivially serializable variables into the generated URL link.
What I was trying to explain is a system that stores the data on the server and passes pointers to this data to the client. Web applications that use databases already work this way. The id/primary key column of a database table is the pointer.
To serialize a closure you generate ids for the values of the free variables. Then you save the values in a hash table and send the ids to the client (instead of the actual values). To deserialize a closure you fetch the free variables from the hash table.
This system doesn't work yet because this hash table keeps growing. Every time you serialize a closure you put the free variables in the hash table. This is wasteful because some of the values might already be in the hash table. This is simple to fix; just look if the value is in the hash table, and if it is you return the existing id (this can be done efficiently by having two hash tables: ids->values and values->ids).
What I was trying to explain was to make this seamless to the programmer, i.e. the programmer doesn't have to use a let-like form listing the variables he or she wants to use, he or she just uses them. Basically the base system scans through the enclosed variables, checks if they're trivial, and encodes them directly if so; if not, it puts the data for retrieval somewhere.
Mostly I'm concentrating on keeping Arc succint, and modifying the base system without, if possible, modifying higher-level code.
In any case your suggested implementation does look good.
> if it is you return the existing id (this can be done efficiently by having two hash tables: ids->values and values->ids).
Maybe I'm being dense here, but how is:
(user state + resource state) -> link -> code
different than:
(resource state -> link) + user state -> code
other than that in the first instance you're encoding user state in the control layer?
In both cases you're mapping (user state + resource state) to a specific block of code. How does the order you do it in affect anything else?
I find that both are extremely useful.