HN
Hacker News
Top
New
Best
Ask
Show
Jobs
Comment by marmadukester39 | Hacker News Reader
Parent
Full thread
marmadukester39
·
Is it? Videos are just sequences of frames
View on HN
rdedev
·
Each frame of the image would have to be divided into many sequences. Atleast that's how transformer based image models work. Then you have to account for audio data too in the same way. It just blows up the compute required
Reply on news.ycombinator.com