> The gradient computation is only ill-defined at places where a spike would get added or deleted.
I feel like I'm missing something here. Like if you do it naively, the gradient is zero when a spike isn't added or deleted. And infinite when it is. Which is completely unhelpful.
Now the "natural" solution is to invent some differentiable approximation of the spiking network, and compute derivates of that, and hope that the approximation is close enough that optimizing it leads to the spiking network learning something useful.
A more principled version might be to inject some noise into the network. This would mean that you have a probability of spiking in a certain pattern (or better, a class of patterns that all have the same semantic). You could differentiate the probability of correct output with respect to the weights and try to drive it towards 1.
Is your approach in either of these classes?