I don't think the answer to a high quality SNN model is going to require purpose-built hardware or a supercomputer to run. I think you will see emergence even with extremely rudimentary single core CPU-bound models if everything else is done well.
Event-driven is a superpower when architected correctly in software. This means you can pick clever algorithms and lazily-evaluate the network. Implications being that networks that would ordinarily be impossible to simulate in real time can now be simulated this way. You could have neuron state in offline storage that is brought online as relevant action potentials are enqueued for future execution.
I feel like I'm missing something here. Like if you do it naively, the gradient is zero when a spike isn't added or deleted. And infinite when it is. Which is completely unhelpful.
Now the "natural" solution is to invent some differentiable approximation of the spiking network, and compute derivates of that, and hope that the approximation is close enough that optimizing it leads to the spiking network learning something useful.
A more principled version might be to inject some noise into the network. This would mean that you have a probability of spiking in a certain pattern (or better, a class of patterns that all have the same semantic). You could differentiate the probability of correct output with respect to the weights and try to drive it towards 1.
Is your approach in either of these classes?