SMT, i.e. simultaneous multi-threading, means that in every clock cycle many instructions are simultaneously initiated from all threads and then those instructions compete for the multiple execution units of a superscalar CPU, in order to keep busy as many of those execution units as possible. For each of the concurrent execution units, e.g. for each of the 6 integer adders available in the latest CPUs, a decision is made separately from the others about which instruction to be executed from the queues that hold instructions belonging to all simultaneous threads.
It is true that once fetched an OoO CPU does a significant amount of scheduling and it is possible that in a given clock cycle instructions from both threads are getting fed to an execution unit. But I don't think that's the essence of SMT.
For example the original larrabee is described as 4-way SMT, but as P5-derived it was a simple in-order design with very limited superscalar capabilities. I very much doubt that at any time instructions from more than one thread were at the execution stage.
While the shared Intel decoders alternate between the threads and the queue that stores micro-operations before they are dispatched is also partitioned between threads, this front-end is decoupled from the schedulers that select micro-operations for execution, which may choose in any clock cycle as many uops as there are execution units and in any combination between the SMT threads.
Even in the first Intel CPU with SMT, Pentium 4, up to 3 instructions were fetched and decoded in each clock cycle and there were places where up to 7 instructions in any combination between the 2 SMT threads were executed during the same clock cycle.
In modern CPUs the concurrency is much greater.
Except the newer *mont cores that have truly separate decoders and fetchers and could indeed decode for two hypertreads separately.