The answer for why it wasn't productized might just be pretty straightforward.
LLMs still are better than Jev at the task, just across the board slower.
Anyone who had a reason to try this already tried it (ads/recommendations) - back in 2023/2024 during the first fine tuning wave and it was accurately determined that it was not worth the effort, the results were more bogus than just using CoT, so frankly parallelism meant nothing if bogus * parallel = bogus.
So thrown into the dumpster and nobody really cared to revisit because it was already tried.
Pretty much sometime between then and now it somehow became the state where the tradeoff makes sense now.