HN
Hacker News
Top
New
Best
Ask
Show
Jobs
Comment by sadhorse | Hacker News Reader
Parent
Full thread
sadhorse
·
Does every token requires a full model computation?
View on HN
onedognight
·
No, you can cache some of the work you did when processing the previous tokens. This is one of the key optimization ideas designed into the architecture.
Reply on news.ycombinator.com