Princeton ‘AI Snake Oil’ authors say GenAI hype has ‘spiraled out of control’
venturebeat.com
venturebeat.com
The claims of "liberal bias" being just a completely bunk study: https://www.aisnakeoil.com/p/does-chatgpt-have-a-liberal-bia...
No evidence of GPT-4 getting worse over time: https://www.aisnakeoil.com/p/is-gpt-4-getting-worse-over-tim...
The analyses are measured and well written. I always enjoy when a new one comes out.
I read the the article you linked to, and either you're putting words in their mouth and/or making claims that go way beyond the conclusions presented in the article.
The authors choose to distinguish between "capability degradation" and "drastic behavior change." Now, almost no one I've witnessed complaining about past and present results of ChatGPT care one bit about making such a distinction. They often simply complain that their previously good result producing prompts are now spewing garbage.
The authors explicitly admit that "the kind of fine tuning that LLMs regularly undergo can have unintended effects, including drastic behavior changes on some tasks." i.e. Fine tuning may in fact be causing the worse results.
Furthermore, the authors merely make an additional point that such a drastic behavior change shouldn't be assumed to be a "capability degradation." Meaning that, theoretically, if you toy around with your past prompt enough you should be able to eventually find an alternative version of the original prompt that outputs the past result you were looking for. Again, none of the complainers that I've come across really give a shit about this theoretical conclusion. They only care about past prompts now giving different results. In the article, the authors freely admit that this "drastic behavior change" could be objectively occurring due to fine-tuning.
To further prevent readers from making sweeping generalizations from their comments, the author's explicitly state that OpenAI makes it pretty damn difficult to do reproducible research on the LLMs in question, so you should not come to any hasty conclusions regardless ("As we have written before, this underscores how hard it is to do reproducible research that uses these APIs, or to build reliable products on top of them.")
Well, "building reliable products/processes on top of ChatGPT" is all the complainers care about anyway, so we're back to square one. Far from disproving the claim that ChatGPT is becoming more unreliable and producing worse results for the complainers use cases, the authors are in fact saying that their grievances may be justified for the only practical criteria the naysayers actually give a shit about.
Take blockchain for example; people spent years and millions of dollars looking for additional usages for the technology beside coins/coin-contracts; and aside from the forementioned usages it didn't really gain routine adoption. That was why the non-coin-blockchain always struct me as a "hype train" or "snake oil" because it was a solution in search of a problem. NFTs are even worse since they haven't found any purpose to exist yet (money laundering?).
Contrast that with the "GPT" offerings (ChatGPT/Bing/Copilot/Bard/etc). Multiple of my colleagues are actively using them routinely every workday and when you demo it to another person, they understand it, and they too start utilizing it in their workflow. Heck, my mom discovered it herself on Bing's homepage and was telling me I should check it out.
That's the opposite of "snake oil" or a hype train, it is arguably a competitive threat to search engines.
PS - Small disclaimer: "AI" is a nebulous term. I'm talking specifically about LLMs. Other types of "AI" have already been here for a long time making our lives better (e.g. computer vision).
edit: I'm prepared for the enshittification of GenAI, but given competitive edge, many people will pay top $ for faster/higher parameter inference, I'm hoping ad support is minimal on the super deluxe god tier AI subscription which I'll have to pile all of my spare cash into.
How conceptually could a paid search engine solve the problem of SEO and spam? (Curation seems economically infeasible?)
Is what you're actually proposing ad-free search results? (I'd be willing to pay Google approximately $0.01 per search for this. I suspect this is probably more than they get from anything but the most expensive ads (pharma ads that confused old people will click on).)
What I would really love is a browsing-assistance agent based on something like an LLM that can summarize and de-cruft search results for me and show me content extracted from web pages in a standard format without any Javascript. (I don't need an LLM to be the arbiter of truth, I just want it to be good at summarizing documents.)
(Because Google is so wedded to the ad ecosystem I'm not sure I trust them to be the steward of technology that is basically about removing ads from my life.)
OTOH, I am a paying ChatGPT Plus user and there are some days where I feel like I get my monthly subscription price's worth in a single day.
In terms of summarization, this study Anyscale just published is interesting - llama2-70b gets within a hairsbreadth of gpt-4/human scores for summarization, so it's conceivable that sometime soon, you might be able to run a local LLM on your browser and get decent results (4-bit quants of llama2-70b currently take about 40GB of memory, you need at least 2 x 24GB or 1 x 48GB GPU to run them at reasonable speeds atm): https://www.anyscale.com/blog/llama-2-is-about-as-factually-...
1. How expensive are these to operate when compared to search engines? If search engines are about as good at information retrieval, but are much cheaper, what does that mean for the future of these services?
2. How good are these systems actually when compared to search engines? I get that for things like code, you can immediately verify results through testing, but GPT was trained on what looks good to humans thanks to HFRL, not necessarily for correctness.
I really wonder what's going to happen with this comparison in the near future. I think there's two possibilities:
1. Running inference on models and keeping the models relevant by introducing new data remains extremely expensive and companies like OpenAI operate at massive losses. Eventually this leads to their products getting worse.
2. Models get cheaper to run and efficient enough to run on mobile hardware without losing much fidelity. OpenAI still might be in trouble in this scenario if "open source" models like Llama2 get popular.
I think in this aspect the difference between GPT vs Google not in level of tech, but because GPT took text from copyright owners and distributes it without their consent and without showing their ads, which is being challenged legally in my understanding.
As far as I can tell, it performs a search and then feeds the full text of the search results to an LLM and present a summary with citations.
I think the review was basically, "I’ve had worse cocktails. I’ll be honest."
So...meh?
No amount of Blockchain "Webinars" would have convinced my kids to use the blockchain.
- My mom cant figure out Blockchain either
- My sister (doctor) and brother (3 degrees) cant figure out Blockchain either, because they are too afraid of scanning the wrong QR code and having their balance drained
My mom, sister, and brother all use ChatGPT. My brother uses it actively for work.
In any case, defining public use or accessibility as the metric of relevance of something is quite illogical. You end up with countless absurdities like relativity being worthless and CocaCola being the most relevant invention ever.
Which 'blockchain'? Blockchain is a singular noun, that's like saying "they can't figure out phone".
I think this is an important detail. A lot of "forced into it" numbers will get presented as "happy users" in various unfair metrics comparisons.
Can you define what a metaverse is in this case and how is it different from what we used to call an MMORPG?
Roblox involves no block chain and no VR, so what exactly is "metaverse-y" about it?
I disagree. Kids will turn anything not too abstract into toys. An improvised radioactive sample (actually some neon indicator bulbs in a box) + Geiger counter brought by my parents = a version of a hide-and-seek game from my own past.
Also, a model maglev train is something so obvious that I am sure such toys will appear immediately once the technology becomes viable.
I’d wager the biggest non-speculative use case is pig butchering scams and ransomware.
Probably, but drugs are also a big use-case (and I don't just mean retail, its used in wholesale for transferring sums of money internationally to buy precursors, for example)
Sounds like grade A Microsoft marketing speak. But wait, there's more!
In this instance, I don't see how this offers any advantage over the usual registrar-based DNS: at the end of the day I'm just using the ENS rootnode keyholders as my registrar, and I don't see how the blockchain is solving any problem that couldn't be more efficiently solved with a traditional database.
Not too long ago the US government put pressure on all financial institutions to stop processing payments for Wikileaks without them ever having been convicted of committing any crimes. That's just a single example plucked out of the sky, regardless of your feelings about that organization.
The great advantage that we get from a decentralized, opt-in trust infrastructure is that we have an avenue around these unjust and extralegal encroachments.
Having said all that, even if you don't care about politics, there are countless examples of centralized entities changing ownership, or changing strategies, and their users pay the price. See: Twitter taking people's usernames.
You might be fine with dealing with living at the whims of billionaires but I want my digital life to be more durable than that.
2. The second part argues about centralized entities changing ownership. I don't think the example of Twitter usernames is a good use-case for blockchain: there's nothing here that isn't already solved by public key crypto alone (which is how blockchain solves the problem anyway).
Maybe Twitter here is just an example and you want to talk about corporate control over payment services? I don't think this is a compelling argument. There's no shortage of payment services and even if one of them is bad you can always find another. Yeah: Stripe or your bank go out of business tomorrow, but switching to Apple Pay or a different bank is not a huge problem.
I agree with you that public key crypto solves many of these problems, but you still need to publish your public key somewhere. And you need an infrastructure where public keys are first class citizens with the platform.
I gave ENS as an example precisely because it is open source digital public infrastructure where such things can be published forever, and it is not corporately controlled. It happens to use a blockchain to do consensus and create incentives for people to opt into running the public infrastructure. I personally think these game theoretical incentives are integral for the functioning of the platform but am not married to the idea if there are better ones.
Something like Ethereum is an anti-authoritarian platform. You don't use it unless you desire the qualities of anti-authoritarianism. For everything else there's centralized solutions.
I think we've seen some people striving for a competitive advantage, and willing to see one exist, but I don't know if we've seen any "AI company" turn a meaningful and durable profit yet (open to being wrong on this).
To me (so far), it seems like GPT use cases that apply it as a table stakes feature of some greater application (rather than an ecosystem or a platform), are the ones that are actually showing promise in terms of expected utility + value capture.
If GPT4+ level engines become as cheap to execute as a SQLite query, I think that's where things get interesting (you can start executing this stuff at the edge).
But I still can't see new companies (a la OpenAI, MidJourney, etc.) making a lot of money in this scenario, it seems to overwhelmingly favor companies that already have distribution.
Yea the sql queries, string or array manipulation are better and fast than google 8/10 times because of the errors it generates.
Granted, if you are spending 10x the future value of the product your are offering ... then even a 10x decline in compute costs won't get you where you need to be.
they quantized model from 16 bits to 4 bits which was low hanging fruit, and looks like they can't quantize it anymore to 2 bits..
Pixel 6+ phones have TPUs (in addition to CPUs and an iGPU/dGPU).
Tensor Processing Unit > Products > Google Tensor https://en.wikipedia.org/wiki/Tensor_Processing_Unit
TensorFlow lite; tflite: https://www.tensorflow.org/lite
From https://github.com/hollance/neural-engine :
> The Apple Neural Engine (or ANE) is a type of NPU, which stands for Neural Processing Unit.
From https://github.com/basicmi/AI-Chip :
> A list of ICs and IPs for AI, Machine Learning and Deep Learning
In my opinion, Microsoft's proposed 365 Co-pilot pricetag of $30 per month per user will probably make a profit. High usage users being offset by low usage users when corporate 365 group memberships come into the mix. Many corporates will take a chance on it, willing to throw plenty of money at anything that is perceived as having a chance of improving hard to define competitive edge, creativity and productivity.
On the other hand, I’m not sure that you can legitimately argue that it is not over-hyped. The only way that it isn’t is if we achieve “singularity” imminently - because that it is what a sizable number of people, including possibly some people at OpenAI, are actually expecting.
That was no reflection of the actual potential it just indicated that people at all points of the value chain had not yet grasped how or what the internet was or how it would mesh with humanity.
That’s where we are now. It’ll burst but not because it’s overhyped; because we don’t collectively understand the implications.
Also, even if some people in some usages get benefit out of it, does not mean that fuels hype and snake oil to try to sell it to uninformed people to use in inappropriate places. Baking soda is good for making bread, lots of people use it for that with great success. It is snake oil for curing cancer, however.
And for art NFTs, much of the interest is from the collectors space or collector sentiment as these are an improvement in supply transparency and provenance in comparison to other mediums of collections.
An important snake oil trait is making grand claims which have little basis as fact, particularly with things which are difficult to understand or difficult to verify (if not impossible.) You can say anything about a product if the prospective buyers can't verify the claims. And so, the claims tend to get broader and flashier.
Therefore, a snake oil sales pitch could still be applied to a highly useful product.
Nobody knew what it meant then, but it was still used to hype up products
We're in a housing bubble, at the peak of a multi-century economic mega-cycle, we're overdue for earthquakes, tsunamis, asteroids, the caldera under Yellowstone, flu pandemics, etc etc. It's super easy to say all that and demand people should listen, what's hard is committing to any actionable prediction.
If you’re wrong in your predictions nobody cares because it’s better than if you’re right.
On the other hand, if you’re an optimist and you are wrong you look like an idiot and the social penalty is higher.
So people continue to predict doomsday and most of the time it doesn’t happen.
Well, the same is true for any prediction of anything unlikely. People continue to predict world-changing innovations, and most of the time they are wrong. That's the reality of any low-probability/high-impact prediction, in the positive or negative direction.
As a specific case of that: most startups fail, so predicting a startup will fail is safe.
I'm pretty prone to exactly that prediction. But I like this quote from Erik Davis (https://www.google.ch/books/edition/High_Weirdness/Rcq2DwAAQ...): "In the court of the mind, skepticism makes a great grand vizier, but a lousy lord."
The point of most critiques/polemics on AI Snake oil isn't on effectiveness of AI in context. It's about mis-application, belief its AGI, belief the answers are right, and belief it has no downsides. It's often mis-applied, It's absolutely not, and not even on the road to AGI, the answers are not always right, and it has massive social downsides, employment included.
It has upsides, sure. It's not the job of snake oil warners to catalog the good, they're pointing out the abysmally bad, in the current AI hype-wave.
It's absolutely not AGI, but it's not "not the road to AGI". Every step forward could be the road to AGI.
What people are doing now is amazing. I love looking at pictures of the pope in puffer jackets. I love "this person does not exist"
It could be on the road to AGI, but then, I could win lotto. and I do enter lotto, even knowing the odds. The thing is I don't plan as if I WILL win lotto and a lot of "this AI could be the road to AGI" is about planning as if it will, to secure a slice of the future.
It could? yea. But, it isn't.
I don't think you are focusing on what matters most. It's really the language itself that contains the intelligence, not the AI models. The models are just vessels for absorbing all that knowledge encoded in text. Just like human brains - we can have different neural wiring but learn the same things through education.
So the huge datasets these AI systems are trained on are key. That's how models like GPT-4 gain such language understanding. The architecture matters less than all that linguistic data it ingests. And language has been evolving for millennia, long before AI. It replicates through culture, speech, writing. Now with AI it has a whole new medium.
Fascinating question - how will AI affect language evolution going forward? As models produce more human-like text, which can further train better models, it's like a Lamarckian evolution. Acquired linguistic intelligence gets absorbed by AI then improved and propagated back into the corpus.
So while AI tech will keep advancing, it's language evolution that's most profound. By creating this new way for language to evolve, AI could really reshape cultural evolution. Since language enables intelligence, it'll have big impacts on where AI is headed next.
Maybe individually but not collectively. Let's not forget that we are the authors of the data AI consumed, although very few of us actually made a difference - myself included. If something like AGI is truly possible it will replace our process of discovery and creativity, arguably one of our best features. We need to know if it's better than ALL of us rather than any of us because arguably all of us will make that sacrifice.
This is happening, and the reddit thread is just a single instance of a wide phenomenon. I wonder how one can square "snakeoil" theories against reality like this.
Another instance is the 'hey pi' chatbot. For some people it's a superior experience to BetterHelp, or even a lot of so-called professional therapists out there.
The lack of actual human presence is a bit deflating for me, but I still got a good session out of it once. It feels like it has the potential to be at least better than nothing for some of the lonely people out there. The option for different voices is a nice touch that adds to the illusion.
There are so many possibilities. I'm looking forward to seeing what happens with everything from NPCs in video games to robotics (among other things, to actually explain what they're doing & why, and converse). Not to mention the applications in education and health care. Anyone who thinks this is a flash in the pan has not observed enough of what's going on.
Contrast that to the people on https://old.reddit.com/r/StableDiffusion, where people willingly experiment with multiple approaches to media generation. At least some of those people will be interested in 3D modeling and will willingly take OP's job, because they genuinely enjoy it and OP now does not.
In the end, adapt, or others who do will take over. That's been the historical precedent for millennia, and it will not stop now.
The camera is the best metaphor. In the 19th century, being an artist was a real thing; it was really a technical field. You were a portrait painter and the best artists would study for years to be as accurate to reality as you could - look at how the light moves and all of that. All of a sudden, the camera comes out and artists fear this is the end of art because this new thing can capture reality better than anyone can.
But then very quickly, people realized that in some ways this liberates the artists. It's not a coincidence that the impressionist movement coincided with the advent of the camera. People realized it's not about capturing reality but the expression and feelings conjured. This led to an explosion about what art is and come out of the trap about painting, nobility, and grand scenes into things that really evoke and challenge us.
People are now saying AI can write pretty well. It can code pretty well. It can create movies, images pretty well. What that tells me is that it liberates the creator to move beyond that. Someone can elevate and integrate and manage these tools.
Most criticism of AI has all the hallmarks of coping mechanisms. They seem to shift quickly between declaring it an ineffective parlor trick, and then say it's so good and effective that it poses a risk to the human spirit. There are a lot of legitimate criticisms I've seen, especially of OpenAI's potential regulation capture and misuse of models, but calling it all snake oil is hilariously naïve and seems more like a fear of change.
My favorite part with crap predicting doom is the sheer unawareness of the rate of change of the rate of change.
If you had asked me 18 months ago how long after GPT-3 until something existed with the capabilities of GPT-4, I'd probably have guessed about 5 years.
If you asked me 5 years ago how long until an AI could explain why a joke was funny, I'd have maybe guessed at least a decade or two, if it was even possible.
If you asked me 10 years ago if I'd see AI replacing artists or copywriters in my lifetime, I'd have guessed maybe when I was in a retirement home (I'm still a fair ways away from that).
No one thought what exists today was even possible within our lifetimes a decade back.
I'm reminded of the Louis C.K. routine about everything being amazing and no one is happy.
Unthinkable AI has already become so normalized that people are predicting its doom based on shortcomings that are less than two years old because AI doing fucking IMPOSSIBLE things (or so everyone thought) is only that old.
The rate is outrageous, and while I do think there's currently a significant setback with obsolete alignment approaches being carried forward to models that probably need new techniques, that's going to be a temporary step back in parallel to significant strides forward in the underlying technology from hardware to model design to improved knowledge in how to squeeze the most water from the rock.
I just hope when all these folks turn out to have been dead wrong that we don't collectively forget. Futurists should live and die by their record, but too many have goldfish memory and continue to listen to false futurists well after they've shown their own snake oil hand.
Is this actually the case? Are there examples of professional writers/academics/etc who have lost their jobs to LLMs?
If AI meant the human was now free, I'd be happy, but ATM AI seems to mean the human will be jobless and soon homeless and starving.
"In the last few months, there has been this increasing so-called rift between the AI ethics and AI safety communities. There is a lot of talk about how this is an academic rift that needs to be resolved, how these communities are basically aiming for the same purpose. I think the thing that annoys me most about the discourse around this is that people don’t recognize this as a power struggle.
It is not really about intellectual merit of these ideas. Of course, there are lots of bad intellectual and academic claims that have been made on both sides. But that isn’t what this is really about. It’s about who gets funding, which concerns are prioritized. So looking at it as if it is like a clash of individuals or a clash of personalities just really undersells the whole thing, makes it sound like people are out there bickering, whereas in fact, it’s about something much deeper."
Oils derived from various snakes in the desert were found to be effective treatments and remedies for various things by the Native American population, and therefore widely used by certain tribes.
Once the White Man caught wind of this, he capitalized on its reputation, and a booming business of scams bloomed anywhere there were gullible buyers. Of course, the snake oil was rarely authentic or effective for its advertised uses.
Therefore, snake oil was given a very bad and rather undeserved reputation for the rest of history, and medical science continues to flounder.