1,055 karma · joined April 1, 2012
Find me at http://tullo.ch and github.com/ajtulloch.
Email is andrew@tullo.ch
[1]: https://www.intel.com/content/www/us/en/docs/intrinsics-guid...
Essentially if you have a Beta(a, b) prior then your prior mean is a/(a+b) and after observing n samples from a Bernoulli distribution that are all positive, your posterior is Beta(a+n, b) with posterior mean (a+n)/(a+n+b). So in your example you effectively have a Beta(0, x) prior and x (“suspicious”/“gullible”) is directly interpreted as the strength of your prior!
It’s worth internalizing almost every single detail if you’re an engineer interested in writing high performance numerical codes on modern hardware.
> The idealized market was supposed to deliver ‘friction free’ exchanges, in which the desires of consumers would be met directly, without the need for intervention or mediation by regulatory agencies. Yet the drive to assess the performance of workers and to measure forms of labor which, by their nature, are resistant to quantification, has inevitably required additional layers of management and bureaucracy. What we have is not a direct comparison of workers’ performance or output, but a comparison between the audited representation of that performance and output. Inevitably, a short-circuiting occurs, and work becomes geared towards the generation and massaging of representations rather than to the official goals of the work itself. Indeed, an anthropological study of local government in Britain argues that ‘More effort goes into ensuring that a local authority’s services are represented correctly than goes into actually improving those services’. This reversal of priorities is one of the hallmarks of a system which can be characterized without hyperbole as ‘market Stalinism’. What late capitalism repeats from Stalinism is just this valuing of symbols of achievement over actual achievement.
The "kernel trick" from kernel SVMs only works because of the existence and uniqueness result from the RRT on the underlying Hilbert space.
defn softmax(t) do
Nx.exp(t) / Nx.sum(Nx.exp(t))
end
See https://ogunlao.github.io/2020/04/26/you_dont_really_know_so... etc.The math here is pretty much first-year undergraduate level calculus, and it's worth going through section 2 since it's quite clearly written (thanks Prof Domingos).
Essentially, what the author does is show that any model trained with "infinitely-small step size full-batch gradient descent" (i.e. a model following a gradient flow) can be written in a "kernelized" form
y(x) = \sum_{(x_i, y_i) \in L} a_i K(x, x_i) + b.
The intuition most people have for SVMs is that the constants a_i are, well, constant, that the a_i are sparse, and that the kernel function K(x, x_i) is cheap to compute (partly why it's called the 'kernel trick').However, none of those properties apply here, which is why this isn't particularly exciting computationally. The "trick" is that both a_i and K(x, x_i) involve path integrals along the gradient flow for the input x, and so doing a single inference is approximately as expensive as training the model from scratch (if I understand this correctly).
– Corey Robin in "The Reactionary Mind"
- Mountains of the Mind (a history and first-person account of mountain climbing)
- The Wild Places (a history and exploration of the 'wild' landscapes of the British Isles)
- The Old Ways (a history and exploration of the ancient paths of the world)
are all really excellent (and can be read in any order). He's a fabulous writer, kind of like a Kapuściński for the natural world.