Frontier models are far exceeding even the most hardcore consumer hobbyist requirements. This is even further
Frontier models are far exceeding even the most hardcore consumer hobbyist requirements. This is even further
IIRC the 512GB mac studio is about $10k
(commentary: things are really moving too fast for the layperson to keep up)
I'm not saying it is a perfect analogy, but it is by far the most familiar one for people to describe what sparse activation means. I'm no big fan of over-reliance on biological metaphor in this field, but I think this is skewing a bit on the pedantic side.
re: your second comment about pruning, not to get in the weeds but I think there have been a few unique cases where people did lose some of their brain and the brain essentially routed around it.
Typically, input gets routed to a number of of experts eg. top 2, leaving the others inactive. This reduces number of activation / processing requirements.
Mistral is an example of a model that's designed like this. Clever people created converters to transform dense models to MOE models. These days many popular models are also available in MOE configuration
MOE is not the holy grail, as there are drawbacks eg. less consistency, expert under/over-use
https://www.youtube.com/watch?v=zwHqO1mnMsA
I wonder how well the aftermarket memory surgery business on consumer GPUs is doing.
I want one. Hot air blows.
This will absolutely scar, if not char, your cornea faster than you can blink.
There is nothing special about "lasing power." It amounts to a 45-watt light bulb, nothing more and nothing less.
Of course the laser is tightly focused. That's pretty much one of the defining properties of laser devices. How else do you think the laser is heating the microprocessors in the video?
They shouldn't be focusing it to a point under any conditions. Whether it's as safe as it could be is a different question, of course. For instance, you'd like to think that the act of configuring it for a smaller beam footprint would reduce the power at the same time, as opposed to requiring a separate adjustment that might be overlooked by the operator. Would have been nice if the video had addressed that and other safety considerations, for sure.
A lot depends on the exact wavelength. 1400 nm and longer is much less worrisome than near-visible IR.
The laser is collimated but not focused so by your logic it will be fine.
This is advice on par with eating tide pods.
About all we can agree on, I think, is that neither of us knows enough about the product to argue about it usefully.
Unlike you I do know what I'm talking about.
I feel like because you didn't actually talk about prompt processing speed or token/s, you aren't really giving the whole picture here. What is the prompt processing tok/s and the generation tok/s actually like?
Maybe my 6000 Pro spoiled me, but for actual usage, 6 or even 9 tok/sec is too slow for a reasoning/thinking model. To be honest, kind of expected on CPU though. I guess it's cool that it can run on Apple hardware, but it isn't exactly a pleasant experience at least today.
But again, not if you're using thinking/reasoning, which if you want to use this specific model properly, you are. Then you have a huge delay before the actual response comes through.
> MacStudio is the simplest solution to run it locally.
Obviously, that's Apple's core value proposition after all :) One does not acquire a state-of-the-art GPU and then expect simple stuff, especially when it's a fairly uncommon and new one. You cannot really be afraid of diving into CUDA code and similar fun rabbit holes. Simply two very different audiences for the two alternatives, and the Apple way is the simpler one, no doubt about it.
There are consumer-ish hardware that can run large models like DeepSeek 3.x slowly. If you're using LLMs for a specific purpose that is well-served by a particular model, you don't want to risk AI companies deprecating it in a couple months and push you to a newer model (that may or may not work better in your situation).
And even if the AI service providers nominally use the same model, you might have cases where reproducibility requires you use the same inference software or even hardware to maintain high reproducibility of the results.
If you're just using OpenAI or Anthropic you just don't get that level of control.
Which ones? I wanted to try a large base model for automated literature (fine-tuned models are a lot worse at it) but I couldn't find a provider which makes this easy.
https://docs.cloud.google.com/vertex-ai/generative-ai/docs/m...
Lambda.ai used to offer per-token pricing but they have moved up market. You can still rent a B200 instance for sub $5/hr which is reasonable for experimenting with models.
https://app.hyperbolic.ai/models Hyperbolic offers both GPU hosting and token pricing for popular OSS models. It’s easy with token based options because usually are a drop-in replacement for OpenAI API endpoints.
You have you rent a GPU instance if you want to run the latest or custom stuff, but if you just want to play around for a few hours it’s not unreasonable.
> https://docs.cloud.google.com/vertex-ai/generative-ai/docs/m...
I don't see any large base models there. A base model is a pretrained foundation model without fine tuning. It just predicts text.
> Lambda.ai used to offer per-token pricing but they have moved up market. You can still rent a B200 instance for sub $5/hr which is reasonable for experimenting with models.
A B200 is probably not enough: it has just 192 GB RAM while DeepSeek-V3.2-Exp-Base, the base model for DeepSeek-V3.2, has 685 billion BF16 parameters. Though I assume they have larger options. The problem is that all the configuration work is then left to the user, which I'm not experienced in.
> https://app.hyperbolic.ai/models Hyperbolic offers both GPU hosting and token pricing for popular OSS models
Thanks. They do indeed have a single base model: Llama 3.1 405B BASE. This one is a bit older (July 2024) and probably not as good as the base model for the new DeepSeek release. But that might the the best one can do, as there don't seem to be any inference providers which have deployed a DeepSeek or even Kimi base model.
I feel like private cloud instances that run on demand is still in the spirit of consumer hobbyist. It's not as good as having it all local, but the bootstrapping cost plus electricity to run seems prohibitive.
I'm really interested to see if there's a space for consumer TPUs that satisfy usecases like this.
https://openrouter.ai/deepseek/deepseek-v3.2
This only bolsters your point. Will be interesting to see if this changes as the model is adopted more widely.