> When I hear superscalar, I generally expect to see a processor dynamically checking data dependencies at runtime
No. They use brute force, they have as many state machines as their superscalarness, and these SMs working simultaneous, saving results in OoO buffer, and periodically synchronize via special algorithm.
Each SM have full set of shadow registers and state of CPU.
To be precise, exists examples of non-orthogonal or "partially" OoO machines - Pentium MMX and IBM POWER. In them, CPU core divided to few parts, MMX parts dividing registers file, POWER have separate pipeline for memory operations, and these parts could do only in their limited space, but other mechanics is same as ideal orthogonal OoO design.
For example, when happen branch and unknown which side will run and have spare resources, SS running each sides (in separate OoO pipeline), and when will know finally effective side, results of other side just discarding, and results of effective side merging with other calculated data.
If happen branch, but resources limited, works predictor, which learn on previous runs and usually could predict with >80% accuracy, which side will run, so most probable side chosen.
When just happen spare resources on run, got last updated PC+1 and run other SM in parallel from there, than on some place (mostly on buffers fill), all data from OoO exec merging with main flow.
All this mean, carefully crafted program could see difference of OoO vs non-OoO execution (not exactly in memory, but by measure jitter, as merge OoO pipeline takes some time), but if not care, mostly have excellent compatibility and if fortunate, OoO will run as many times faster as number of OoO pipelines, which could be significantly large, ie in Pentium-3/4 was 4 or more pipelines.
And as side effects, this all gives high pressure on cache subsystem, and sometimes optimized in unsafe way (but cheap), so happen vulnerabilities, like Meltdown and Spectre.