I was watching some youtube hype demo of it to pick items (each an emoji) from a pile according to questions. One of the questions was: "what can a magnet attract?" and it picked like 8 metal items, while a literal magnet was left unpicked. Ofc the video was too busy praising it to notice. Was all I needed to see
That it can make mistakes? That’s expected, it deal with plausibility like other models no? You would need to look at the actual response and the assigned probabilities to evaluate. Generally demos are not a great way to evaluate a technology, it’s a way to get hype but the next step is to actually look at the details
But isnt the whole point of Jev that "we trained it so you dont have to"? Yet it make such a simple mistakes still.
Yes, I don’t think hallucinations will ever go away. But because it doesn’t work with language the type of hallucinations aren’t too comparable to LLMs. But it is still a risk of course, whatever system you design should take that in consideration