1,550 karma · joined December 25, 2021
Jos Stem style?
But i think this 'being wrong' is kind of confusing when talking about LLMs (in contrast to systems/scientific modelling). In what they model (language), the current LLMs are really good and acurate, except for example the occasional chinese character in the middle of a sentence.
But what we mean by LLMs 'being wrong' most of the time is being factually wrong in answering a question, that is expressed as language. That's a layer on top of what the model is designed to model.
EDITS:
So saying 'the model is wrong' when it's factually wrong above the language level isn't fair.
I guess this is essentially the same thought as 'all they do is hallucinate'.
SaaS is a business model while malleable vs. rigid is a property of the software itself.
> looking like really good product managers.
Exactly and that's a different field with a different skillset than developer/programmer.
And that's the purpose of technology in the first place tbh, to make the hard/tedious work easier.
> They are young and inexperienced today, but won't stay that way for long.
I doubt that. For me this is the real dilemma with a generation of LLM-native developers. Does a worker in a fully automated watch factory become better at the craft of watchmaking with time?
For me, the core feature of Netlify is building and deploying static websites quickly, with minimal configuration and triggered by git commits.
Does any of these really resemble that experience (except for the CDN Netlify uses, of course)?
Dokploy vs. CapRover, Dokku, Coolify
Cool, awesome job, as far as i can tell as a fan of the movie!
So you did what was best for yourself... and the group.
While it is seemingly hard to calculate it, maybe one should just make a database website that tracks specific setups (model, exact variant / quantisation, runner, hardware) where users can report, which combination they got running (or not) along with metrics like tokens/s.
Visitors could then specify their runner and hardware and filter for a list of models that would run on that.
Do you guys know a website that clearly shows which OS LLM models run on / fit into a specific GPU(setup)?
The best heuristic i could find for the necessary VRAM is Number of Parameters × (Precision / 8) × 1.2 from here [0].
[0] https://medium.com/@lmpo/a-guide-to-estimating-vram-for-llms...
It forcefully tries to hack you and steal your scarcest resource.
Also recommend this course by Gilbert Strang @ MIT:
https://ocw.mit.edu/courses/res-18-009-learn-differential-eq...
https://www.youtube.com/playlist?list=PLUl4u3cNGP63oTpyxCMLK...
I would think LeCun was aware of that. Also prior sequence to sequence models like RNNs have already incorporated information about the further past.