Full threadboltzmann-brain·> Our implementation is up to 2x faster than optimized speculative decoding baselines and up to 5x faster than autoregressive decoding with open source inference engineswhat about per-FLOP?View on HN