Full threadBharath1234·Since the Transformer decoder cannot be parallelized during inference, how can it be cost effective at scale?!View on HN