Writing an LLM from scratch, part 10 – dropout
gilesthomas.com
gilesthomas.com
The Llama thing is interesting, though!
If the weights are effectively set to zero by the dropout, shouldn't the propagated error in the backward pass be zero too, automatically?
(I.e., as I understand it, OP's intuitive notion of "fairness" is literally how the error propagation works: Neurons are adjusted by the degree by which they contributed to the output)
I could well be misunderstanding you, though!
That's a good example. Read up backpropagation on wikipedia again, and I think you're right there and I had some misunderstandings.
Eq. 4 of [1] says:
> However, if [neuron] j is in an arbitrary inner layer of the network, finding the derivative of [loss for one specific target value] E with respect to [output of j] o_j is less obvious. [...]
Considering E as a function with the inputs being all neurons L = { u , v , … , w } receiving input from neuron j, [...] and taking the total derivative with respect to o_j, a recursive expression for the derivative is obtained:
[derivative of E with respect to o_j] = sum ℓ in L ( [derivative of E with respect to o_ℓ] [derivative of o_ℓ with respect to net_ℓ] * [weight of neuron ℓ for o_j])
Therefore, the derivative with respect to o_j can be calculated if all the derivatives with respect to the outputs o_ℓ of the next layer – the ones closer to the output neuron – are known.*
So if a neuron is disabled through dropout, this would affect all neurons in the layer "before" it (i.e. closer to the input layer).
I think you could also argue that a dropped out neuron has its set L being artificially set to empty, so the sum in the formula would reduce to zero. But that would indeed be something different than setting the weight to zero.
[1] https://en.wikipedia.org/wiki/Backpropagation#Derivation