Full threadvisarga·This is how language modeling evolved - it used to be able to only output 5-10 words that made sense at a time. Now we get 5-10 seconds of video at a time.View on HN