653 karma · joined May 9, 2023
I'm not sure your argument stands when it comes to OP's idea of a single card with 128GB VRAM. This would be enough to run ~180B models with reasonable quantization --we're not near maxing out the capability of 180B yet (see the latest 32B models performing near public SOTA).
This indeed would push rapid and wide adoption and be quite disruptive. But sure, it wouldn't instantly enable competitive training of 405B models.
No one ever talks about just how easy it is to add GPIO to anything.
I know it sounds ridiculous, but ~1-2 nights of absolute terror per week is totally worth it compared to how it used to be (getting maybe ~1 good night of sleep every couple weeks.)
Oops now you gotta hire a new one.
Regarding the standalone servers, I suspect they’re aiming for usability over elegance in the short term. It’s a classic trade-off: get the protocol in people’s hands to build momentum, then refine the developer experience later.
That's disappointing as the devtools approach always has limitations.
Kura agents, Runner H, and scrapybara will all end up more reliable than you.
Further, it's now easier than ever to decouple one's online life from google and so anyone not doing so is taking unnecessary risk.
There's so much more to it than money.
With recent DDR generations and many core CPUs, perhaps CPUs will give GPUs a run for their money.
Beg to differ.
We have to consider the fact that human prompting + selection of outputs is essentially RLHF, so the models can and will continue to get better over time.
It's not the end of the Internet, it's the beginning of a new era.