Because 17B active parameters should reach enough performance on 256bit LPDDR5x.
Because 17B active parameters should reach enough performance on 256bit LPDDR5x.
Tenstorrent is on fire, though. For small businesses this is what matters. If 10M context is not a scam, I think we'll see SmartNIC adoption real soon. I would literally long AMD now because their Xilinx people are probably going to own the space real soon. Infiniband is cool and all, but it's also stupid and their scale-out strategy is non-existent. This is why https://github.com/deepseek-ai/3FS came out but of course nobody had figured it out because they still think LLM's is like, chatbots, or something. I think we're getting to a point where it's a scheduling problem, basically. So you get like like lots of GDDR6 (HBM doesnn't matter anymore) as L0, DDR5 as L1, and NVMe-oF is L2. Most of the time the agents will be running the code anyway...
This is also why Google never really subscribed to "function calling" apis
RTX 3090: 24GB RAM, 936.2GB/s bandwidth
Tenstorrent p150a: 32GB RAM, 512GB/s bandwidth
an extra 8GB of ram isn't worth nearly halving memory bandwidth.
Tenstorrent p300 is coming at 64 GB and 1 Tbps but that's not the point; even p150a with plenty of bandwidth (512 GB/s is fine for inference) and four 800G ports. But hardware is not the problem: even if they had the hardware, they wouldn't know what to do with it. Privacy is a hobby to most people, making you feel good.
The Tenstorrent cards exist, but are low in availability and the software is comparatively nonexistent. I'm excited for them too, but at the end of the day, I can buy a used 3090 today and do useful work with it, while the same is not true of TT yet.
god I love this website.