Its value is mostly, and paradoxically, for training LLMs, to get "unspoiled" models.
Its value is mostly, and paradoxically, for training LLMs, to get "unspoiled" models.
• Only photos taken before October 1, 2023, will be accepted
(I chuckled.)[1] https://web.archive.org/web/20240322134531/https%3A%2F%2Fwww...
It goes even deeper. If OpenAI has 100M users, and they consume let's say 10K tokens per month on average, that means 1 trillion tokens are read by the user base, who then go out and act in the world, at the very least they will read the text. The impact of AI onto the world is huge because we flock to it by the millions.
Such a massive new source of text will surely influence how language is spoken and potentially support many activities, including scientific research. One of the consequences will be much wider circulation of information, as users only need to point it to the right direction and it will readily combine information across multiple fields. Recently updated AIs might be even more informed than domain specialists because nobody got time to read everything new, like LLMs, we still need most of our time for work.
Tangentially related: I really loved the 'shibboleth' that went like
complex and multifaceted [things] that encompass [other things]
from November 2023 [1]. Even though it was clear this will soon disappear from newer models being too obvious, I was still quite surprised, that "multifaceted" didn't even get into more recent study from March 2024 [2] that pointed out skyrocketing of terms such as commendable, intricate, and meticulous.[1] https://news.ycombinator.com/item?id=38501589 / https://blog.j11y.io/2023-11-22_multifaceted/ / https://twitter.com/padolsey/status/1727555475573440996
[2] https://news.ycombinator.com/item?id=39909692 / https://arxiv.org/abs/2403.07183 / https://twitter.com/myfonj/status/1773371415149891871
Articles, social media posts and photos are now often AI generated. There is going to be a certain percentage of today's students that will graduate "knowing" blatantly wrong things an LLM told them. These things will eventually make their way to books and society.
One thing the scientific community never got around to accepting themselves and admitting to the public is how much of the learning is trust-based rather than fact-based. Nobody can prove everything from the ground up, you need to put your trust into your predecessors. Eventually you'll have the capacity to doubt and check a percentage of what they said, but not everything. When this circle breaks and you don't know if your lecturer is for real or just parroting LLMs we'll be in trouble. The "financialized" ultra-competitive environment of higher education doesn't help with that either.
In the years I spent working with research scientists, I observed the exact opposite of this. They were acutely aware of this and it was a common topic of discussion. It's important to them because they need to be as aware of and to call out as many assumptions as possible.
Having spent time in academia myself though, my reading is that while everyone professes an epistemology that accepts its own limits, in practice everybody behaves as though they'd grasped absolute truth through divine revelation.
Before chatGPT I'd see the words "it's worth noting.." maybe five times in my entire life (39 years old).
I now see that stupid phrase dozens of times per day, every single day, everywhere I go.
Including HN, which tosses it around even more than Reddit.
This particular phrase indicates that what comes next is not part of the main thrust of the argument but a side observation in support of it. It prepares the reader or listener to receive it in that spirit rather than have to figure its status out for themselves. Used superfluously or incorrectly it can be pompous and empty, but used well it can genuinely improve comprehensibility, even if a word processor's grammar checker might choose to put a squiggly line under it.
When I've encountered it, it usually comes off as a way to just make the writer sound smarter rather than a way to clarify communication.
Following code performs the task: <some code>
Please note that * the code can only work when using 64 bit architecture. * the variables a and b are to be provided and can't exceed 100.
It's one of the standard phrases that people should double-check when they find themselves using it, to make sure it's really called for.
Also chatgpt has a habit of weakly editorializating almost anything as a way of covering its a*. For example, I asked it how much Jira costs, and after it told me it added:
"It's important to note that pricing and plans may change over time, so I reco.."
ok, sure pricing and plans may change. we kind of already guessed that. and it's not important to note it.
I asked how much Flat costs and it added: "It's important to note that Flat may offer discounts or promotions from time to time, "
Again, not important. And if it was important, I would just note it. I would actively utilize that in my decision-making process.
"However, it's important to note that the endianess is dependent on the system architecture"
Again, completely unnecessary. just say, "Endianess is dependent on system architecture."
"It's important to note that the value returned by time() is in UTC." ok, yeah that's kind of important. but just say, "the value is in UTC."
Like if I said to you, "It's important to note that you must press submit on a web page to return data to the server," you'd say something like, "yeah is that important though? that's just how it works?!"
Like I don't need to say, "It's important to turn the key in a car." It is "important" in the sense that it's also vital or CRITICAL, or utterly essential. It's just how it's done. How it works. It's not particularly important.
"Overall, the ovipositor is an important reproductive structure that enables female insects to lay their eggs "
is it important? is it more important than other reproductive structures? how much more? why? aren't they all important??!?
Generally, "it is important" is a convenient phrase for abdicating responsibility. It removes people or entities from the picture as well as objections. It conjures a world where question-asking isn't present, without exceptions, or context. It also hides that invisible authoritative relationship between two people. So if you want to be vaguely manipulative to people who don't think critically, you can obtain some basic compliance without overtly threatening them this way. The fact that it won't work on most people, but chatgpt batters us with it is just tiring.
Furthermore, as my friend pointed out, ChatGPT works by stroking your ego and fanning its own flames. It tries to make you feel clever for using it. Telling you that what it is telling you is important is a low-grade sleazy way of inflating itself in your eyes, and also hinting to you that this information is worthy of your attention; that perhaps you can pat yourself on the back for choosing chatgpt... because it's not just random information from a google search... ChatGpt is giving you important information. don't you know.
> It is worth noting that before chatGPT I'd see the words "it's worth noting.." maybe five times in my entire life (39 years old).
:-D
This is the tell that present-generation LLM "AI" is not actually intelligent. It doesn't truly understand anything. It's just a general purpose lossy compression algorithm for text that is queryable through replaying the model to complete prompts.
I don't necessarily mean "just" dismissively. That's obviously incredibly useful. My laptop now has local LLMs on its hard drive that contain a conversationally queryable digest of the sum total of a large fraction of human knowledge, and we've basically solved natural language interfacing to computers and natural language translation.
But it's not "AI" in the sci-fi sense. If it was, it would be generating truly novel insightful content that enlarged the base data set not padding the data set with regurgitations of what is already there.
I actually go back and forth about how much this represents a step toward true thinking AI. It's hard to judge because the novelty of these things being able to work so well with language causes a "wow" reaction that might cause me to overestimate just how large a step we have made. Are these really just nothing more than linguistic JPEGs or is there a lot more going on here?
Agreed, I like to refer to it as collective intelligence because it's only as "smart" as what people have already published
The way you can "talk to" these machines is certainly clever, and the way they generate their replies is fascinating technology and certainly beats the shit out of a traditional chatbot, all of these things are true and I think they have a bright future as such. Natural language UI is a cool concept and it's cool to see some serious progress there. But that's all it is, and all it really can be. I don't know for certain what the path is to true emergent intelligence from the machine, but I assure you that shoving billions of words scraped from the internet through an LLM is not that.
Now, if for whatever reason you want ream upon ream of pretty dull and repetitive content churned out on a given subject, that's far longer than it needs to be to convey the limited information within as apparently tons of online businesses do, written damn near to perfection? ChatGPT's your hookup, no question. However you will never get blinding insights, you will never get new ideas, and it will never be your friend. Sorry.
It also isn't a computer as we were used to.
It's a new thing. Being able to query the latent space of a large subset of human knowledge using natural language is the closest to 'intelligent computer' as we've ever come to.
Yes it is lossy compression, it confabulates the bits that it lost, but it's still useful.
Just because it can't do everything any human could do or conceive of being done, doesn't mean it isn't doing things that humans do but rocks and even dogs can't.
I would agree with this statement, but it's still laughably far from it. As the other comment said, it lacks the ability to rank the veracity of the information it presents: it will present it in a grammatically correct way, but it has no way of knowing if the information it's presenting is accurate which explains ChatGPT's tendency to just make some shit up when you ask it things.
And in the odd event it is correct, it's simply digesting information from other sources, likely search engines either it's own or not. Again, this is not useless and not not-clever: however in that way it does remain pretty useless because of the above issues of accuracy. The best it can do is get you the most highly ranked information on a topic based on it's pool of knowledge, and that may or may not be the most correct one.
Also I'd add: these two issues compound one another, because the LLM is obscuring the source of it's information. You can have it add the link to it's citations, that would be helpful, but you don't know where it got those citations or why it picked them, and the LLM can't explain that for you. And, those choices were based on factors that are definitively not the accuracy or reliability of those sources, merely their ranking in whatever search algorithm is being used. Even if you presume this AI was trained on data relevant to the topic you're asking about, that doesn't mean it (and in fact, as I understand LLM, it's basically incapable of) understands that topic in a way that makes sense to a subject matter expert.
If you asked an engineer what size lumber would be required to construct a floor that could support a small vehicle, that engineer would probably have a rough guess based on their previous experiences, and further, could then do the work: research the weight of the vehicle, the situation the structure will be installed in, the environment it will exist in, etc. and come back with a solid answer to your question. An LLM, by comparison, would consult the writings it was trained on for similar questions, and create an answer it thinks matches yours. But what does "similar" mean here? Is the wood and the dimensions the same? Is what it's looking at based on a small vehicle or a cement truck? Is it looking for articles about structures in similar climates or did it pick some that were closer matched to the structure but in completely different environments? On and on.
Using an LLM this way takes the problematic aspects of googling things and basically factors them out, because now you can't see what it's googling, which results it's picking or why, and can't verify their relevance. How many times have you searched specific things and left a google result completely befuddled as to why on earth it included it? Now imagine that, except there's essentially zero chance the AI caught that, and it incorporated that irrelevant data anyway.
However, it does have verisimilitude.
IMO, that quality (of appearing to be correct) is so attractive to us humans, that all other considerations are being discarded.
Eventually someone will claim the goal posts are being moved and they will probably be right because there were never any to start with.
Oh nonsense. AI: Artificial Intelligence.
Artifical - made or produced by human beings rather than occurring naturally, especially as a copy of something natural.
Intelligence - the ability to acquire and apply knowledge and skills.
ChatGPT is certainly artificial, but it lacks the ability to apply knowledge and skills. It will do it, KIND OF, on the knowledge front, if you ASK it to, but we wouldn't call a car that moves when the accelerator is pressed a creature, nor would we call google intelligent because it can respond to queries when prompted. It's a machine. Firmly in the realm of a machine that responds to input from other actors to accomplish a task presented it.
You can leave a ChatGPT instance running for a thousand years and it will never do anything until prompted. That's not intelligence.
There's simply no debate here to be had. It's a CLEVER machine, and it does things a lot of previous machines could not, but a machine it remains nonetheless.
I do not see how this requirement you have stated is an inherent property of how you have defined "artificial" or "intelligence". I think what you have described fits more with the traditional depiction of AI, but I don't agree that it is inherent.
Which I why I agree with the person you responded to: I think there is no consensus view on what is and isn't AI because people seem to have different tests, weights, etc. in their mind for what "counts". I think it's a bit of a pointless exercise, but a mildly interesting one at least.
Because we've had software since the inception of computers that only does things as it's told and when it's told, and that's called... software, programs, apps, etc. Intelligence belies something entirely different in the minds of laypeople and many tech people alike. To assert otherwise is to upend an entire understood element of our shared culture for the purposes of marketing. It's ridiculous.
Hell, ELIZA still beat ChatGPT 3.5 in terms of % of people that believed they were interacting with a human in a recent study! [1] Notice that we only properly identified humans about 60% of the time too.
[1] https://arstechnica.com/information-technology/2023/12/real-...
That's what I mean by autonomous.
I've been 100% in the "exciting tech being painfully misrepresented" camp all along too, and I think more people are saying that as time passes, but I think you might surprised what little people need for insights, ideas, and friends.
Rubber duck debugging, Brian Eno's Oblique Strategies, John Cage's I Ching use, journaling, etc already delivered a useful degree of those things with far less sophistication, and these generative tools really can take many of those to another level.
This is not an argument for it not being intelligent nor is it a coincidence. It has been shown mathematically that there is a deep, fundamental connection between compression and intelligence (an intelligent agent that behaves optimally effectively ends up computing the Kolmogorov complexity of the program that best fits and predicts the environment i.e. optimal compression). I do not know what you mean by "truly understand" since it obviously has an understanding of the various words it uses even if it's knowledge of the "meaning" of a word is simply a vector in a very high-dimensional space. I suppose you think this representation of meaning is profoundly different from meaning as it is represented in the human brain, that the former is a mere facsimile of the meaning while the latter must be the meaning itself, assuming such a thing exists. Even among humans I have conversations with people who use words or phrases for which they clearly have a poor sense of the meaning and make contradictory or nonsensical statements as a result.
All this "tells" is that it cannot (effectively?) bootstrap itself to higher performance levels with it's own output. This is not that surprising since humans have the same problem (for similar fundamental reasons). You do not teach children to write better by making them read their own literary output.
There are many areas where that estimation might be difficult without human intervention: arts of all kinds.
However, in a more fact-based area like programming, a NN training routine could be modified so that what a LLM thinks it's learning can be tested. That could apply to CS and math, at least.
Eventually, NNs integrated with robotics could learn physical sciences without human intervention or pre-selected training material, and would still converge to superhuman knowledge even starting from random ideas. Giving a NN access to robotics without a human in the loop is probably too dangerous, though... it might decide to unsafely test its understanding of nuclear reactions, or the r-number of ebola...
I know that I regard anything I see on the web that was created after the advent of LLMs to be suspect and requiring additional vetting before giving it much weight.
Maybe we'll get stickers like non-GMO plants for organic code.
One big problem is that Google also lies about dates when you do a 'before' search, so AI spam still shows up even if you filter it out.
Scraping a new dataset without contaminated content is impossible now. There's no authoritative source on the age of any given website. (You can try places like the Internet Archive, but most such archivers aren't interested in handing their data over to AI firms because of the bad PR.)
The only way to ensure you don't have AI content is to use a dataset created before ~2021-2022. This too is hard because all the big ones are illegal to posses as they contain CSAM.
The real answer is that AI developers played themselves. Humanity was creating exponentially more data each year, and by rushing these AI tools out the door, all that data is useless now. By the end of this year there'll be more new "AI-polluted" data on the web made after 2021 than the total amount of "clean" data made up to 2021.
By the time the actual content comes back up, it won’t matter.
It's the same here, we have many trusted sources that we don't believe would lie about when content is from - Internet Archive, newspaper archives, widely used web crawls, Wikipedia exports, etc. Obviously you don't use an AI (or a black box search engine) to get non-AI content, you get it from people and systems you trust.
We've been yelling about Web of Trust for decades by now, and never built it and now it's too late.
I think you're being realistic about the underlying adversarial dynamics. I.e., it's an arms race regarding recognizing AI authorship.
>*Google Books is indexing low quality, AI-generated books that will turn up in search results*
This was never about articles, it is about books, though. You can't retroactively publish a book in 2021 if we're not in 2021.
It's bewildering to see how people here are so confident in their tone while arguing about things they think are said in the linked post, but don't know for sure because they never even opened it to read it.
You make a valid point though, which is actually more scary now that I think about it... it is becoming less and less possible for non-google (/other huge) companies to train models on verifiably human-generated content, as only Google (/other huge companies) have enough accurately timestamped data to train on. Hmm.
It's not like this is exactly a new phenomenon - Plato's famous "Allegory of the Cave" was a thought experiment in how the media we consume can create what we perceive as reality - it's just that we're discussing an uncomfortably recent intersection of philosophy with computer science. I said it a few times in the threads that blew up over Google's diverse Nazi bungle; I didn't get into computers to discuss anthropology and how humanity constructs meaning; I actually did it to avoid that and do work I considered purely practical after washing out of a liberal arts degree. That said, the whole thing ironically felt like very good art in the sense that it got people to discuss the nature of these generative AI systems. For me especially, the image of the Black Pope was very thought provoking as a (lapsed) Catholic; why is it that there has never been a black pope?
There were a few popes of African descent, though; don't know if their skin color happened to be black
Still, I don't get why it would matter to any Catholic (I'm not Catholic), except to (some) Catholics in the US of A maybe, since race seems to reign supreme over there, even more than God Himself to (some, supposed) believers
I can get into my personal feelings if it matters; I don't think they matter much to anyone except me. The quick summary is that I dug the grave of a great-great Aunt of mine by hand at a small cemetery attached to the Catholic church in the tiny, forgettable town she was born in. One quirk of this rural, very white (I assume the entire population is a fairly close cousin of one type or another to me), Catholic town is that the priest is an African immigrant. My aunt was quite racist and so I thought it was a sort of ironic that the last group of people who saw her out of this world were twenty-five percent black.
So an AI-generated image of a black pope made me stop and think for a minute about what Catholic identity means; a concept that as someone who was raised in that tradition but is unable to myself believe in the divine has always been problematic. Basically, it was good art, though I don't think that was the intent of the corporation that created it.