Don’t get me wrong, I’ve used LLMs and been amazed by their output, but the p-zombie statistical model has no idea what it is saying back to you and the idea that we should trust these things at all just seems way premature
Don’t get me wrong, I’ve used LLMs and been amazed by their output, but the p-zombie statistical model has no idea what it is saying back to you and the idea that we should trust these things at all just seems way premature
Your argument that this is just a statistical trick sort of gives away that you do not fully accept the usefulness of this new technology. Unless you are trolling, I'd suggest you try a few queries.
Sure, but my baseline expectation is far above primary school level.
But why are these use cases different? It appears to me that code is at least subject to sustained logic which (evidently) translates quite well to LLMs.
And when you ask an LLM to be creative/generative, it’s also pretty amazing - j mean it’s just doing the Pascal’s Marble run enmasse.
But to ask it for something about the world and expect a good and reliable answer? Aren’t we just setting ourselves up for failure if we think this is a fine thing to do at our current point in time? We already have enough trouble with mis- and dis- information. It’s not like asking it about a certain period in Japanese history is getting it to crawl and summarise the Wikipedia page (although I appreciate it would be more than capable of this) I understand the awe some have at the concept of totally personalised and individualised learning on topics, but fuck me dead we are literally asking a system that has had as much of a corpus of humanity’s textual information as possible dumped into it and then asking it to GENERATE responses between things that the associations it holds may be so weak as to reliably produce gibberish, and the person on the other side has no real way of knowing that
Do you trust Wikipedia (based on volunteer data), do you trust news outlets (heavily influenced by politics, lobby groups, and commercial companies), do you trust blogs or forum posts (random people on the internet)?
Any local models can go off the rail very easily and more importantly, they're very bad at following very specific instructions.
I have trouble convincing colleagues (technical people) that the same question is not guaranteed to result in the same answer and there's no rhyme or reason for any divergence from what they were expecting. Imagine relying on the output of an LLM for some important task and then you get a different output that breaks things. What would be in the RCA (root cause analysis)? Would it be "the LLM chose different words and we don't know why"? Not much use in that.