15,833 karma · joined March 11, 2013
Website: https://kay.is
Blog: https://fllstck.dev
GitHub: https://github.com/kay-is
I think, it's a bit much to call this "no hallucinations".
Technically true, but in practice you could still choose the wrong result or the probabilities can be off.
Like, yeah, you don't hallucinate, but only because you force the user to decide in the end.
None of my grand parents made it over 65, all of them heavy drinkers.
So, I'd probably have that going for me.
All these people writing about how they turned their life around, sleeping better, thinking clearer, getting more healthy, etc.
I stopped drinking 3 or 4 years ago and I don't feel any better for it.
However, I only drank like one or two nights a month, when going on a party, so most of my life I was sober anyway.
The only thing I noticed is that I'm now more awkward in social situations. I mean, sure, I might be healthier than before, but I don't have any numbers for that. It's just a thing I assume to be true.
People are constantly complaining about GPT/Claude constantly changing under their apps without notice.
The lock-in is less pronounced as it is with AWS or MS.
So if anything moves, the interactions don't change their speed, they stay constant, but the delta to the movement changes, and thus time dilates.
Would be interesting to know if there is some absolute/fundamental global time that has nothing to do with the time we perceive.
Is Jev a decoder (e.g., BERT) or is it some kind of encoder (e.g., GPT) that just happens to be trimmed down to only outputting a handful of tokens for the answers and their probability?
However, it might have fewer restrictions than a BERT and/or is smarter (whatever that means).
Somehow I expected inference engines are generic LLM runtimes that can execute any weight.
So, to get this right.
Someone trains a model.
They release the weights and a reference implementation of the model architecture.
Then a provider has to host this model either by running inference via the reference implementation, an open source implementation, or build their own.
Does this mean, providers don't just differ in quantisation and configuration, but also in inference engine implementation?
I'm using it right now and it's noticeably faster.
I'd also say, it seems smarter, but I think that's because of some harness updates I installed. (I haven't used pi for almost a month)
I was hoping for a bit more, but it's still 100% faster for a very good price, so I won't complain.
I have a Samsung Fold 4. The front screen is too narrow and the folding screen doesn't open all the way anymore.
Not what I expected from a 1400€ phone.
https://www.geeky-gadgets.com/deepseek-v4-1-flash-review/
I hope some of those speed increases will make it to production.