33 karma · joined August 20, 2020
Publishing your actions on the Internet is a little different. If people were affected by the action, they are affected (likely unknowingly) by the publication too - and the audience that you grant right of reply has at best an ideological horse in the race, not true skin in the game. And not much courage is required to engage with an opposing position.
So "living publicly" on the internet leaves a permanent door open to ideological conflict, mob behaviour, and creates a disconnect between action and reaction - in both time and space.
Kinda alien for a monkey brain to wrap banana powered neurons around.
Interesting, what's being visualized there is actually a failure mode for an unidentifiable equation - the valley where the error is zero and therefore all solutions are acceptable. Introduce noise into the measurements of error and that valley being too flat causes odd behaviour
You and the cup are objects, and physically send messages as you interact. That leads to changes in the physical world as each actor decides what to do with the incoming information, by physics or by conscious action.
So far so good. Except software is just information, and so the software version of that interaction includes the "person put hot cup down on table" event. That interests somebody, so they rapidly express their displeasure and rush to put a coaster underneath...
And that is valid a model of computing. Direct messaging between interacting objects, a stream of events of the produced changes, and actors that consume that stream for things and optionally chose to initiate a new interaction
1. Gradient descent is path-dependent and doesn't forget the initial conditions. Intuitively reasonable - the method can only make local decisions, and figures out 'correct' by looking at the size of its steps. There's no 'right answer' to discover, and each initial condition follows a subtly different path to 'slow enough'...
because...
2. With enough simplification the path taken by each optimization process can be modeled using a matrix (their covariance matrix, K) with defined properties. This acts as a curvature of the mathematical space, and has some side-effects like being able to use eigen-magic to justify why the optimization process locks some parameters in place quickly, but others take a long time to settle.
which is fine, but doesn't help explain why wild over-fitting doesn't plague high-dimensional models (would you even notice if it did?). Enter implicit regularization, stage left. And mostly passing me by on the way in, but:
3. Because they decided to use random noise to generate the functions they combined to solve their optimization problem there is an additional layer of interpretation that they put on the properties of the aforementioned matrix that imply the result will only use each constituent function 'as necessary' (i.e. regularized, rather than wildly amplifying pairs of coefficients)
And then something something baysian, which I'm happy to admit I'm not across
Kinda like shadows on the wall. A nice, steady light on a clear surface makes it easy to pick out the shapes your shadow makes.
So form an explicitly disloyal "party" from within the existing system. If it catches on...
The link tries to explain the story, and I'm hoping a few of you might see what I'm trying to achieve. If acorn could find a community of people wanting to use it to solve problems I believe it could turn into something quite beneficial! Apologies for the detail, however. A snappy summary seems to have defeated me.