I would like to augment just a couple of your observations with a few explanations --- I think the reason the architectures that you mention work so well is because they exploit hierarchical structure inherent in the data. Lots of things in the universe have hierarchical structure, e.g. vision has spatial hierarchy, video has spatial and temporal hierarchy, audio has temporal and spatial hierarchy.
(side note, thats how a baby learns vision "unsupervised" -- the spatial patterns on the retina have temporal proximity, so the supervisor is "time")
RELU is good, but you have to remember how you're vectorizing your data - if you're vectorising is normalizing to 0-1 then RELU fits nicely, but if you're scaling to mean 0 std dev. 1, then half your samples are negative and you'll get information loss when RELU does this: http://7pn4yt.com1.z0.glb.clouddn.com/blog-relu-perf.png so remember to keep your vectorizer in mind.
> My gut feeling is that I'm talking from my rear;
Nope not at all, you have good insight - my email address is in my profile, come join the slack group too (email me, let's go from there)