I think a big question here is, "what is the goal?" A key distinction is whether the NN is intended to handle the signal directly (the NN
is the filter) or if it's just a technique for finding filter coefficients (the NN
designs the filter for the signal).
For some tasks, like system identification (used in echo cancellation), the NN is the filter - aformentioned adaptive filters are used for this case right now. It can also be used for black box modeling, which has numerous application in real time or otherwise.
For others like Butterworth (and other classic designs) there's not really a good reason to use a NN. Butterworth (and Chebychev I/II, elliptical, optimum-L, and others) are filter design formulae with a closed form (for a given filter order) that yield roots of transfer functions that have desirable properties - I'm not sure how a learning approach can beat a formulae that are derived from the properties of the filter they design.
There are iterative design algorithms that do not have a closed form, like Parks-McClellan. It is however quite good at what it does - it would be interesting to compare these methods against some design tricks to reduce filter order.
There are some applications of filter design that NNs can do that we don't have good solutions for with traditional algorithms, like system identification, another is in optimization (in terms of filter order) and iteratively designing stable IIR filters to fit arbitrary frequency curves (for FIR it's a solved problem, at significant extra cost compared to IIR).
As for topologies of the filter itself that is an interesting angle. Topologies affect quantization effects (if you wanted an FPGA to realize the filter with the fewest number of bits, topologies matter) and for time variant filters (like an adaptive IIR) there are drastic differences between the behavior of various topologies. I'm not sure how it would manifest with a NN designing the filter.