This is why they went with the “laptop” cpu. While it’s slightly slower than dedicated memory, it allows you to run the big models, at decent token speeds.
This is why they went with the “laptop” cpu. While it’s slightly slower than dedicated memory, it allows you to run the big models, at decent token speeds.
Biggest limitations are the memory bandwidth which limits token generation and the fact it's not a CUDA chip, meaning longer time until first token for theoretically similar hardware specifications.
Any model bigger than what fits in 32 GB VRAM is - in my opinion - currently unusable on "consumer" hardware. Perhaps a tinybox with 144 GB of VRAM and close to 6 TB/s memory bandwidth will get you a nice experience for consumer grade hardware but it's quite the investment (and power draw)
I understand it's faster but still...
Did they at least do an internal PSU if they went the Apple way or does it come with a power brick twice the size of the case?
Edit: wait. They do have an internal PSU! Goodness!
https://community.frame.work/t/framework-desktop-deep-dive-p...
Currently avoid machines with soldered memory, but if memory can be replaced and still have similar performance, that would change things.
You absolutely can (and should) build your own for slightly cheaper. Just find the fastest DDR5 CUDIMMs you can paired with the fastest memory bus mobo.
Without CUDA, being an AMD GPU. Big warning depending on the tools you want to use.
https://docs.scale-lang.com/stable/
https://github.com/vosen/ZLUDA
It's not perfect but it's a start towards unification. In the end though, we're at the same crossroads that graphics drivers were in 2013 with the sunsetting of OpenGL by Apple and the announcement of Vulkan by Khronos Group. CUDA has been around for a while and only recently has it gotten attention from the other chip makers. Thank goodness for open source and the collective minds that participate.