To be clear, all hardware branch predictors are "relatively simple state machines"; they need storage in the branch prediction tables which must super-fast to access, which means they can only store a few (sometimes dozens, but certainly not hundreds) bits per branch to reach the access latency goal. With input to the predictor encoded as binary, the weights quantized and small and encoded into binary, and the history small, even perceptrons are "relatively simple state machines". After all, their implementation is just going to become some combinatorial logic in the end.
Some updates to it were big point of "what we did in Zen" presentations when first Ryzen and EPYC CPUs landed.
bot> A 5% miss rate in branch prediction leads to a 14.7% improvement in misprediction rates on a trace of SPEC2000 benchmarks compared to the gshare predictor. The use of machine learning-based predictors has the potential to improve these results further.
So the circuitry is complicated despite superficial simplicity of the model.
All branch predictors need some way of storing their state and selection logic, and the way a perceptron branch predictor stores its data is just a big table indexed by some hash of the program counter of the branch, which is pretty standard for branch predictors. Also, all branch predictors have a sort of "backpropogation" in that pipelined processors produce the actual result of the branch (possibly many) cycles later, so this also is not as much of a factor. Since the training is a function of the weights you do not need to store extra data beyond the threshold, but that is already being computed as the prediction anyways.