H100 is not an everyday product. Laptop is
edit: fixed typo
Maybe an analogy could be made to espresso, nice espresso machines get costlier. But, you can still get quite good results out of a manual machine like a Flair.
I think this is why the suggestion to rent a machine is not to helpful. In this analogy we’re on BaristaNews, we all know about the industrial machines, lots of folks use them at work. But, the topic of what sort of things you can do on your manual machine at home has come up.
No, reasonably-priced coffee machines is an enabling factor for many people.
If coffee machines weren't reasonably priced, they would not be "very widespread".
Anyway, I was assuming personal use, like the messing-around experimenting that the article is about. (Or who knows, maybe it was part of the author’s job.)
"Pull out credit card, sign up for some thing and pay a bit of money" is a non-trivial bit of friction! Extremely non-trivial!
Especially in a corporate context - you have to get the expense approved. It's not clear if you can put company data onto the machine. Whereas generally running local things on corporate laptops is far less controversial.
"Download this tool and run it." is still an extremely powerful pitch. Pretty much the only thing that beats it is "go to this website which you can use without any signup or payment".
Anyway, I thought the context was doing stuff for personal use/fun, not work.
In my personal life, when its time for fun, I close the laptop and go do some gardening.
Cloud H100 don't count because you need lawyer to review ToS and other agreements.
There's also no overages or hidden charges with a laptop. Past simply breaking it. You know the replacement cost ahead of time, though.
On that note you can rent an H100 for an hour for under $10 which might make for a slightly more interesting test, whats the best model outcome you can train in under an hour.
Far cheaper these days. More like $2-3 for a consumer to do this. For bulk deals, pricing is often < $2.
In terms of conpute efficiency though, Nvidia still has Apple beat. Nvidia wouldn't have the datacenter market on a leash if Apple was putting up a real fight.
What exactly is your point? That instead of expressing workloads in terms of what a laptop could do, you prefer to express them in terms of what a MacBook Pro could do?
"Best model you can train with X joules" is a fairer contest that multiple people could take part in even if they have different hardware available. It's not completely fair, but it's fair enough to be interesting.
Training models with an energy limit is an interesting constraint that might lead to advances. Currently LLMs implement online learning by having increasingly large contexts that we then jam "memories" into. So there is a strict demarcation between information learned during pre-training and during use. New more efficient approaches to training could perhaps inform new approaches to memory that are less heterogenous.
tl;dr: more dimensionally correct
We can / should benchmark and optimize this to death on all axes