Are they really understanding, or putting out a stream of probabilities?
Are they really understanding, or putting out a stream of probabilities?
The "lie detector" is used to misguide people, the polygraph is used to measure autonomic arousal.
I think these misnomers can cause real issues like thinking the LLM is "reasoning".
prefillContext()Probabilities have nothing to do with it; by any appropriate definition, there exist statistical models that exhibit "understanding" and "reasoning".
Lays out pretty well what our current knowledge on understanding is
The idea is: if you have a substantive point, make it thoughtfully; if not, please don't comment until you do.
The previous truncation ("From GPT-4 to GPT-5: Measuring Progress in Medical Language Understanding") was baity in the sense that the word 'understanding' was provoking objections and taking us down a generic tangent about whether LLMs really understand anything or not. Since that wasn't about the specific work (and since generic tangents are basically always less interesting*), it was a good idea to find an alternate truncation.
So I took out the bit that was snagging people ("understanding") and instead swapped in "MedHELM". Whatever that is, it's clearly something in the medical domain and has no sharp edge of offtopicness. Seemed fine, and it stopped the generic tangent from spreading further.
* https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
Generic Tangents is my new band's name.