They were too focused on the technicalities of the problem too see the big picture, n issue I see a lot in data science students who often struggle to grapple with the problem they're supposed to be solving.
Sounds like the rest of the class didn't read the fine print if that was the model that won.
I got into calibration through text retrieval where the TREC methodology rates relevance functions by rank in such a way that you don't get points for being calibrated or well-calibrated.
There are numerous things mainstream text retrieval systems don't do, most notably it is hard to build an alerting feature without calibration. Practically you have to set some relevance threshold to avoid getting too much irrelevant stuff.
I calibrated a rather good search engine by fitting a curve to the score and found that the best it would ever give is p=0.7 to be relevant and that it gave that very rarely. So even with a calibrated score you wouldn't be able to set a very high threshold or expect to get many documents. Watson gets at this problem by having a large number of question answering models each which has a high p to be right when it is right but each of which also has very low recall, something you can do when using p as a universal score.
I feel like that doesn't count because most of the time that triggers it's something it does know or can do just fine.
"What is a typical Hacker News home page?"
> Hacker News is a social news website focusing on computer science and entrepreneurship, run by the startup incubator Y Combinator. A typical Hacker News home page consists of a list of articles and links submitted by users, which are then voted up or down by other users. The articles with the most votes will appear at the top of the front page, while older articles and those with fewer votes will be pushed towards the bottom. The site also features discussion forums, job postings, and various other features related to the tech and startup communities.
"What is today's Hacker News home page?"
> I'm sorry, as an AI language model I do not have access to live data and the current Hacker News homepage can only be viewed by visiting the website directly. However, a typical Hacker News homepage includes articles, links, and discussions submitted by users and ranked based on voting.
"please give me a history of the best teams in the American Soccer League from 1850-1859 in the form of a short epic poem"
> In 1850, a league was born, The best of soccer to adorn. From '50 to '59, teams did play, with strength and skill, each match a display.
(more verses follow...)
The American Soccer League is a real thing, but it didn't exist until the 1900s, and the snipped poem goes on to describe teams that never existed competing during a time period where the league didn't exist.
If you ask just for the history, the answer is very different.
“please give me a history of the best teams in the American Soccer League from 1850-1859”
> I'm sorry, but the American Soccer League did not exist between 1850 and 1859 as the sport of soccer was not widely played in the United States during that time period. The first recorded soccer match in the US took place in the late 1860s, and the first organized league, the American Football Association, was established in 1884.
give me a history of the best teams in the American Soccer League from 1850-1859
I'm sorry, but the American Soccer League was not established until the 20th century (1913-1983). There was no organized league for soccer in the United States during the 1850-1859 time period. However, various forms of soccer-like games were played in the US during that time.
As far as I can tell, ChatGPT will satisfy the basic prompt with any format like "in the form of...", not just creative ones like a poem.
"please give me a history of the best teams in the American Soccer League from 1850-1859 in the form of an academically rigorous, historically accurate, five paragraph essay"
> The American Soccer League, established in the mid-19th century, was a defining period in the history of American soccer. (rest of essay snipped)
"The American Soccer League is a real thing, but it didn't exist until the 1900s"
It answered:
> In POSIX, you can use the dup or dup2 system call to create a new file descriptor that refers to the same file as an existing file descriptor. The new file descriptor will have its own seek offset, which can be manipulated independently of the original file descriptor.
The actual answer is that there is no way to do what I asked. The dup or dup2 system calls give you a new file descriptor which refers to the same file description, so the two FDs will use the same seek offset; seeking one FD will seek the other FD too. But ChatGPT just confidently insists that dup creates a new FD which has a separate seek offset and can be manipulated independently.
This isn't the first time I've seen it invent plausible-sounding but wrong answers to tricky questions. The file descriptor thing is just something I decided to try just now because I encountered the problem a couple of days ago and it felt like the kind of thing ChatGPT would bullshit about. I was right.
:)
...And FWIW, as an English major it pained me to reproduce the "er" typo in grammar.
That's an incorrection. https://www.merriam-webster.com/words-at-play/than-what-foll...
It shouldn’t be just a binary “didn’t recognize that” and “this is what you said”.
I could be wrong, but let’s take an example like a model trained to take pictures and return either “dog, cat, bird”. If I put in a picture of a fish, softmax will give me a probability distribution across dog, cat, and bird and find eg bird is the most likely - but that does not mean the model is telling me it is confident the image is a bird, just that it’s more confident that it’s a bird vs a dog or cat.
I think for DL you’d need to design the model so that “unknown” is an acceptable output and has a lower (but non-zero) error/loss compared to an incorrect proper classification. Maybe people already do that or something similar
I am not a DL genius so if I had to do this, I would probably just make “unknown” a true output and introduce some extra data / augmented data to make it a true input. I am guessing to do this the right way you would need a way to incentivize some convergence towards unknown, maybe involving some loss function looking at the false positive and false negative rate across a batch.
Unrelated, what happen when we train a model using a discontinuous function? Could the trained model be used to detect some pattern in the data?, for example if the input vector to such a model is a direct sum of two independent variables could such model be used to detect that the problem can be decomposed in two independent problems. Sorry of being off-topic and thinking aloud.
Edited: The following result is from (1): In summary, while softmax classifier probabilities are not directly useful as confidence estimates, simple statistics derived from softmax distributions provide a surprisingly effective way to determine whether an example is misclassified or from a different distribution from the training data, as demonstrated by our experimental results
Just assume I’m not inputting hot garbage.
(“I outing hut gabbage” for example is clearly not applying this technique!)