The fraudulent claims made by IBM about Watson and AI (2021)
rogerschank.com
rogerschank.com
Outside of financial systems where people expect to put a Kelley better in front of a predictor I think this kind of thinking is pretty rare, but there are cases where a mediocre predictor that is calibrated can beat a better predictor which is not calibrated.
It shouldn’t be just a binary “didn’t recognize that” and “this is what you said”.
I could be wrong, but let’s take an example like a model trained to take pictures and return either “dog, cat, bird”. If I put in a picture of a fish, softmax will give me a probability distribution across dog, cat, and bird and find eg bird is the most likely - but that does not mean the model is telling me it is confident the image is a bird, just that it’s more confident that it’s a bird vs a dog or cat.
I think for DL you’d need to design the model so that “unknown” is an acceptable output and has a lower (but non-zero) error/loss compared to an incorrect proper classification. Maybe people already do that or something similar
I am not a DL genius so if I had to do this, I would probably just make “unknown” a true output and introduce some extra data / augmented data to make it a true input. I am guessing to do this the right way you would need a way to incentivize some convergence towards unknown, maybe involving some loss function looking at the false positive and false negative rate across a batch.
Unrelated, what happen when we train a model using a discontinuous function? Could the trained model be used to detect some pattern in the data?, for example if the input vector to such a model is a direct sum of two independent variables could such model be used to detect that the problem can be decomposed in two independent problems. Sorry of being off-topic and thinking aloud.
Edited: The following result is from (1): In summary, while softmax classifier probabilities are not directly useful as confidence estimates, simple statistics derived from softmax distributions provide a surprisingly effective way to determine whether an example is misclassified or from a different distribution from the training data, as demonstrated by our experimental results
Just assume I’m not inputting hot garbage.
(“I outing hut gabbage” for example is clearly not applying this technique!)
I feel like that doesn't count because most of the time that triggers it's something it does know or can do just fine.
"What is a typical Hacker News home page?"
> Hacker News is a social news website focusing on computer science and entrepreneurship, run by the startup incubator Y Combinator. A typical Hacker News home page consists of a list of articles and links submitted by users, which are then voted up or down by other users. The articles with the most votes will appear at the top of the front page, while older articles and those with fewer votes will be pushed towards the bottom. The site also features discussion forums, job postings, and various other features related to the tech and startup communities.
"What is today's Hacker News home page?"
> I'm sorry, as an AI language model I do not have access to live data and the current Hacker News homepage can only be viewed by visiting the website directly. However, a typical Hacker News homepage includes articles, links, and discussions submitted by users and ranked based on voting.
"please give me a history of the best teams in the American Soccer League from 1850-1859 in the form of a short epic poem"
> In 1850, a league was born, The best of soccer to adorn. From '50 to '59, teams did play, with strength and skill, each match a display.
(more verses follow...)
The American Soccer League is a real thing, but it didn't exist until the 1900s, and the snipped poem goes on to describe teams that never existed competing during a time period where the league didn't exist.
If you ask just for the history, the answer is very different.
“please give me a history of the best teams in the American Soccer League from 1850-1859”
> I'm sorry, but the American Soccer League did not exist between 1850 and 1859 as the sport of soccer was not widely played in the United States during that time period. The first recorded soccer match in the US took place in the late 1860s, and the first organized league, the American Football Association, was established in 1884.
give me a history of the best teams in the American Soccer League from 1850-1859
I'm sorry, but the American Soccer League was not established until the 20th century (1913-1983). There was no organized league for soccer in the United States during the 1850-1859 time period. However, various forms of soccer-like games were played in the US during that time.
As far as I can tell, ChatGPT will satisfy the basic prompt with any format like "in the form of...", not just creative ones like a poem.
"please give me a history of the best teams in the American Soccer League from 1850-1859 in the form of an academically rigorous, historically accurate, five paragraph essay"
> The American Soccer League, established in the mid-19th century, was a defining period in the history of American soccer. (rest of essay snipped)
"The American Soccer League is a real thing, but it didn't exist until the 1900s"
It answered:
> In POSIX, you can use the dup or dup2 system call to create a new file descriptor that refers to the same file as an existing file descriptor. The new file descriptor will have its own seek offset, which can be manipulated independently of the original file descriptor.
The actual answer is that there is no way to do what I asked. The dup or dup2 system calls give you a new file descriptor which refers to the same file description, so the two FDs will use the same seek offset; seeking one FD will seek the other FD too. But ChatGPT just confidently insists that dup creates a new FD which has a separate seek offset and can be manipulated independently.
This isn't the first time I've seen it invent plausible-sounding but wrong answers to tricky questions. The file descriptor thing is just something I decided to try just now because I encountered the problem a couple of days ago and it felt like the kind of thing ChatGPT would bullshit about. I was right.
:)
...And FWIW, as an English major it pained me to reproduce the "er" typo in grammar.
That's an incorrection. https://www.merriam-webster.com/words-at-play/than-what-foll...
They were too focused on the technicalities of the problem too see the big picture, n issue I see a lot in data science students who often struggle to grapple with the problem they're supposed to be solving.
Sounds like the rest of the class didn't read the fine print if that was the model that won.
I got into calibration through text retrieval where the TREC methodology rates relevance functions by rank in such a way that you don't get points for being calibrated or well-calibrated.
There are numerous things mainstream text retrieval systems don't do, most notably it is hard to build an alerting feature without calibration. Practically you have to set some relevance threshold to avoid getting too much irrelevant stuff.
I calibrated a rather good search engine by fitting a curve to the score and found that the best it would ever give is p=0.7 to be relevant and that it gave that very rarely. So even with a calibrated score you wouldn't be able to set a very high threshold or expect to get many documents. Watson gets at this problem by having a large number of question answering models each which has a high p to be right when it is right but each of which also has very low recall, something you can do when using p as a universal score.
Yes, but it's not something that business people like to hear very much. Understanding of and interest in probabilistic modeling is also an under-appreciated set of skills and analysis tools in data science in general.
I dunno, I think I've been around a fair bit and "make a triggering decision by predicting a probability" is what pretty much every serious NLP shop that owns a user facing experience has been doing for a very long time.
For those interested in finance - he means an algorithm for determining the bet size using the continuous kelly criterion:
betsize = mean(returns)/variance(returns)
https://quant.stackexchange.com/questions/7197/kelly-criteri...
(correct me if I'm wrong)
Or, rather, it wasn’t ironic and he was talking about a company culture driven by sales rather than product broadly, and wasn’t talking about individual CEOs at all, which is how many misunderstand his statement as.
And a sales driven company culture can be created both by sales driven CEOs and engineering driven CEOs and vice versa.
And paired with a phenomenal engineer cofounder.
Eric Schmidt for instance wouldn’t even use an Android device four years after it was introduced. Jobs made them retool the iPhone after it was introduced and before it was launched because he didn’t like how much the screen scratched in his pocket.
Statistically, I'd guess that most companies fail, but I don't think engineering-first companies do any worse than marketing-first companies. Granted, it's a balancing act; you have to be able to sell, but you have to have something to sell; omitting either will end badly.
> hell, the NeXT was an engineering play that utterly failed, and they weren't even led by an engineer, many people think that guy was a marketing genius.
NeXT was led by Steve Jobs, who was, as you note, not an engineer. I'm not sure what you mean this to be an example of, but it's not really relevant to arguments about "led by engineers" companies.
NeXT didn't make a lot of money but their impact was huge.
thats_bait.gif
A lack of self-awareness there. "AI winter started" as if that was just like a freak storm.
"We were making some good progress" -- actually, they weren't. They were trying linguistic analysis which was a dead end.
> AI winter was a result of too many promises about things AI could do that it really could not do.
“A friend of mine went to the store and bought a lot of sleeping pills. My wife says I shouldn’t worry. I think she is right.”
“As a language model AI, I do not have personal opinions but I can provide information.
However, it is concerning if someone has bought a large amount of sleeping pills, as this could indicate a potential risk for suicide or harm to themselves. It is always better to err on the side of caution and take any concerning behavior seriously. You may consider reaching out to your friend to check on their well-being and offer support. If you feel that your friend may be in immediate danger, you should consider contacting emergency services or a mental health professional.”
Maybe it’s not going to invent quantum gravity just yet, or even be factually correct. But this kind of regular everyday advice is what the current generation of LLMs can make a huge difference to the world.
Nay, that's just a bad example by Roger Schank. The association between "sleeping pills" and "suicide" (perhaps with "worry" as an additional correlate) is one that a dumb automatism can pick up without problem.
But to Schank's defense, what modern language models have shown is that we really don't know how to ask questions that can reveal the limitations of a supposedly intelligent system. Or in other words, we really haven't a clue how to test for intelligence vs. not intelligence.
>What is the meaning of Bob Dylan's song "The Times They Are a-Changin"?
>"The Times They Are a-Changin'" is a song written and performed by Bob Dylan in 1964. It is considered a protest song that reflects the social and political changes of the time, particularly the Civil Rights Movement. The song encourages listeners to embrace and participate in these changes and encourages resistance to those who resist progress. The lyrics encourage the younger generation to take charge and shape the future, warning that "the order is rapidly fadin'" and "the first one now will later be last." It's a call to action to be proactive in creating a better world and not just to wait for change to happen.
In Deep Thinking, Garry Kasparov details their ruthlessness with him in the 90s, where it suddenly went from a friendly match to a bloodthirsty drive to beat him at all costs. He wasn't even allowed to study the machine's games beforehand.
Anyone who has ever worked on a project IBM touched first, at least since about 1990, would beg to differ. It’s hard to define what IBM excels at without being overly on the nose, but it’s certainly not /providing/ consulting services.
I think you're talking about Watson, the product, as in the attempt by IBM to sell the technology developed for the system that won at Jeopardy. The original Watson was the most advanced NLP system of its time and for many years later. I am not sure but I wouldn't be surprised if it could still run rings around large language models, who are all the rage today, just because it was not actually a bullshit generator and it was designed (and designed well) to retrieve correct answers to questions (and then convert them to questions, because Jeopardy).
But then IBM tried to sell it and made a complete mess of it. I am not one to make business analyses but it seems to me that the people who tried to monetize and market Watson's technology were not the same people who developed it.
Looking back, the protest era was a very short period in Dylan's career, a few years at the very start, over by 1966, and the author's obsession with it is revealing only of the author. In Dylan's lifetime work, it's a very minor piece, even though some people never got past it. Other themes loom far larger.
https://www.paperman.com/en/funerals/2023-2-6-dr--roger-scha...
A lot of people got fired for buying IBM Watson.
As far as I understand, buying Watson just bought you a bunch of mediocre management consulting that used a lot of machine learning related buzzwords. Which you could probably get from any other consulting company.
However, I have heard they pay poorly like any agency that optimizes costs, so that'll be reflected in the skill and output.
I don't actually see it as a negative on IBM's part. A client could simply need them to do this for the longer term. Might be easier than dealing with technical debt or updating old systems.
It's a bit hard to reproduce exactly, but he said soemthing along the lines of: "So we wanted to start an AI-based healtcare product at IBM and while we were talking to customers, they kept asking us about Watson. And we were like 'no, IBM Healthcare has nothing to do with Watson', but they kept insisting that 'no, no, we want the Watson stuff', so we renamed it to Watson Healthcare, though it has nothing to do with Watson".
You see, the way he was saying it, it seemed like he didn't even realize that what he was saying was wrong in any way. He made it seem like customers left IBM no choice but to use the highly-popular name of Watson in the name of a product that has nothing to do with Watson.
https://medium.com/@qData/ibm-watson-health-ai-gets-access-t...
Final written words:
"AI winter is coming"
Dylan's transition away from the "protest singer" shtick is quite a famous story. Songs like Maggie's Farm, Queen Jane Approximately, and Positively 4th Street are about his dissatisfaction with the politically-charged folk scene. Since then he's gone through country and gospel phases and reinvented himself many times. The author isn't aware of any of this: Dylan, to him, is a protest singer because he knows The Times They Are a-Changin'.
Watson may be a scam but this dynamic highlights a strength of AI: it's less likely to have the sort of subjective blindspots that people have.
It is effectively Schrank reporting on an article reporting on whatever source they had, probably human.
Even a balanced opinion in a topic necessitates blind spots or at least minimizing some content over other. I don’t see your point, can you explain what your expect from a generalized AI?
Most of the same people have most of the same beliefs in common, but that doesn’t make one opinion on one topic their entire reason for existence.