Nearly all of the Google images results for "baby peacock" are AI generated
twitter.com
twitter.com
I wouldn't mind a modern take on geocities sorta system. Where: (1) You can make a webpage that could be about whatever the bleep you wanted it to be about. (2) Only allowed a reduced subset of web technologies. (3) That was free from any advertising or commerce/sales. (4) Was only available to individuals or businesses no larger than closely-held corporation. (5) had clear limitation on AI uses. (6) Had a complete index, categorized and tagged, of all the sites available.
But if I am being honest, that is just the nostalgia for the old internet talking.
This is not going to work - as time progresses, there will be less and less nostalgic people who are willing to put up with that complexity. And "non-commercial" part will ensure that there _never_ be an option to say: "I am tired of fixing my homeserver once again, I am going to put up my site to (github|sourceforge|$1 hosting) and forget about it".
Compare to early web. First thing that came to my mind was Bowden's Hobby Circuits site [0]. It's designed for advanced beginners - simple projects, nice explanations. And there are no hoops to jump through - I've personally sent the links to it to many people via forums, private emails, and so on. It apparently went down in 2023, but while it was still up, I remember regularly finding it from google searches and via links from other pages.
[0] https://web.archive.org/web/20220429084959/http://www.bowden...
I'm ready to pay for a walled garden where the incentives are aligned towards me, instead of against me. I know that puts me in a minority, but I'm tired of the advertising 'net.
Bring it back. Charge me 10 or 20 a month. Give me the walled off chatrooms, forums, IM, articles, keywords, search, etc. Revamp it, make it modern. And make a mobile app.
Everyone wanted a free and open Internet, until AI and the bots ruined it all.
Well... advertising as a business model ruined it all. They get paid for getting page views, so the business model optimizes for maximum page views at minimum cost of creation. This is the end result of what Google and Facebook have spent the last 20 years building.
But I'm sure the engineers who built all this have very nice yachts, so it's all fine.
It's largely already happening in places like Discord.
I think the first company that can capture what Discord has but is not wrapped in that "gamer aesthetic" ui/ux is going to do really well.
[1] https://www.tomshardware.com/news/discord-throttles-nvidia-g...
[2] https://www.thestreet.com/technology/discord-is-making-a-maj...
And eventually someone is going to come up with a search engine for Discord, and the cycle will start all over.
It's great to read these words. People are starting to get it. The Internet is not for you, it's against you.
Patreon, forums, Discord, Facebook, Instagram and so on are all centralized.
With the Internet of the 90s, discussion happened with decentralized Usenet (owned by nobody) with more-focused discussion happening on mailing lists (literally owned by whoever was smart enough to get Majordomo compiled and running).
Email was handled by individual ISPs or other orgs instead of funneled through a handful of blessed providers like Gmail.
Real-time chat was distributed on networks of IRC servers that were individually operated by people, not corporations.
Quickly publishing a thing on the web meant putting some files in ~/public_html, not selecting a WordPress host or using imgur.
Ports 80 and 25 were not blocked by default.
Multiplayer games were self-hosted without a central corpo authority.
One could construct an argument that supports either way being better than the other, but the the Internet of today is not the same thing as it was a quarter of a century ago.
(Anyone can make a "discord server," but all that means is that they've placed some data in some corpo database that they can never actually own or control.)
The technology doesn't matter, it's changing all the time. The difference is the size of the community, and the barriers to entry. If these communities were easy to discover and join and hard to be booted out from, they wouldn't be protected from The Slop. It used to be just getting on the Internet was the barrier. Now we need new barriers: paywalls and word-of-mouth and moderation.
But AOL IM and MSN were centralized walled gardens that came rather late to the game, and neither PhotoBucket nor phpBB existed at all in the 20th century that is the context here.
Losing "dedicated servers" was a huge loss in my opinion. It was fun to play on the same handful of servers and get to know the same group of people. Dedicated servers were also free from profit-driven "matchmaking" schemes since you ended up playing with whoever was on at the time.
I also miss the chaotic multiplayer of the era, where the priority was having fun and not improving your rank on the leaderboard.
Edit: Dedicated servers also each had their own moderation. So if you wanted to play on a server that banned anyone nasty you could. But you could play with a bunch of folks dropping "gamer words" as well.
Then there were custom sprays - something no company would allow in their game today. It's sad how constrained and censored online gaming experiences are today. In many games even dropping the "f-bomb" can get you banned and typing assassin in the chat window yields "***in".
If you want to find it again, change your search engine. Use wiby.me, or search.marginalia.nu. Subscribe to RSS feeds of sites you find interesting, and go from there. Hop on gemini. Subscribe to some activitypub accounts (you don't even need an account for that). Communal IRC servers still exist, forums still exist, independent emails still exist.
Stop falling for the overall doomism HN is so quick to fall for, start doing and living what you want
1. could afford a computer back then and saw the utility of owning one. 2. had access to the internet, so either in college, a 'tech' company, or ties to some local collective that provided access.
When people say 'the old internet' they are referring to a very self selective/elite group.
And that's what made it fun I guess
I'm a little sad for anyone who didn't get to experience of pre-Internet era.
Internet is lead of our time.
I would. But I've asked a lot of people who say "no, I don't want to pay when I can read it for free. I don't mind the ads that much."
Sadly, they won't know what they were missing. It'll be the new normal
Some asshole tech apologist is probably getting ready to post that section from Plato where Socrates complains about writing any minute now.
Of course, that asshole is oblivious to the fact that most if not all of us probably just don't understand what Socrates was missing, so he's just showing his ignorance and stupidity.
> I'm ready to pay for a walled garden where the incentives are aligned towards me, instead of against me. I know that puts me in a minority, but I'm tired of the advertising 'net.
The problem is that, even if you try to do that, the incentives are probably still aligned against you, just maybe less blatantly.
Just look at how many formerly ad-free paid services are adding ads, and how hardware users literally own acts against their interests by pushing ads in their faces (e.g. smart TVs).
The guy who runs the walled garden will always be tempted to get some extra cash by adding ad revenue to your subscription feed, or cut costs by replacing human curated stuff with AI slop (maybe cleaned up a bit).
I did, and...well, let's be careful how we look back at it.
Punch the monkey? Ad supported 'free' internet that literally put an adbar at the top of your browser at all times? Dreadfully slow loads of someone's animated construction sign GIF? Waiting for dial up to connect after 20 tries? Tracking super pixels? Java web applets? Flash? Watching your favorite ISP implode or get bought up? To say nothing of the pre-Google search results (I miss the categories though).
I have plenty of good memories from those days, but it still had plenty of problems. And it wasn't exactly a bastion of research material either unless you really went digging or paid for access.
But my point is, the latter books have this has amazing post-internet, just a ravaged chaotic Wildlands filled with rabid programs & wild viruses. Packets staggering half intact across the virtualscape, hit by digital storms. Our internet isn't quite so amazing, but I see the relationship more subtly with where we have gone, with so so so many generated sites happy to regurgitate information poorly at you or to sell you a slant quietly. Bereft of real sites, real traffic. Watts is a master writer. Maelstrom.
First book Starfish is free. https://www.rifters.com/real/STARFISH.htm
Another book on my reading list!
User facing ability to whitelist and blacklist websites in search results, ability to set weights for websites you want to see higher in search results.
Spamlists for search results, so even if you don't have knowledge/experience to do it yourself, you can still protect them from spam.
It's recreation of e-mail situation, not because it's good, but because www is getting even worse than e-mail.
It is quite surreal to witness. It is certainly fueled by the commercialization of internet due to ads and centralization to user hostile platforms.
The old internet seems to be doing much better. But it lost most of its users in the last 15 years...
What do you mean by this? How do you find the old internet?
Many of e.g. the old niche forums still exist. Like, FOSS sites. GNU project sites seemed not to have aged a day in 20 years, i.e. still party like its 2004.
Also, I think non English sites are better off since Reddit mainly ate English communities and sites.
Facebook is probably what killed most of the living internet. Small community sites. Like the local Kennel club or Boat marina.
A good example of the old internet would be Matthew's Volvo site:
https://www.matthewsvolvosite.com/forums/search.php?search_i...
https://x.com/samhenrigold/status/1843040235325964549
...or a sensical comparison where it just completely misses the point.
Edit: Similar to the second one I just did Panda Bear vs Australia which informed me "Australians value authenticity, sincerity, and modesty. Giant pandas are solitary and peaceful, but will fight back if escape is impossible. "
https://www.reddit.com/robots.txt
Reddit serves a different robots.txt to Googlebot, you can see a snippet of it in Googles summary. If Kagi is getting recent Reddit results then they must be either ignoring robots.txt, using Google as a backend, or also paying Reddit for access like Google.
Bing and DuckDuckGo are certainly still locked out, I just tried searching for "reddit hurricane milton" on both and none of the results are actually from Reddit.
https://help.kagi.com/kagi/search-details/search-sources.htm...
The whole point of the API access change was to charge AI model-makers. I'd be ironic if the API change made destroyed their product and made their data unsellable.
Recently, they came around looking to recruit me. I told them fire Spez or fuck off. (17 year Reddit user here)
I think if the business model had been thought in the sense of the communities and involving mods and users in, it would have been genius, a lot of smaller companies would kill to have people genuinely recommending their products/tools that are hidden behind the biggest wall of them all...
Alas they completely ignored this as a viable avenue and went for the quick buck.
But this won't last and at some point people are going to move on from mass scraping, either because they already got what they want or because garbage goes in garbage comes out or because most of the content will be bot generated and require too much filtering to be useful.
Of course this is the opinion and rambling of a moderately educated individual and I might be totally wrong.
Change can often be for the best, pretty sure it wasn't in this case...
Aaron is probably turning once again in his grave...
I have no idea how intensive these are, as I most of these over bar drinks when I was traveling. Could've just been someone making stuff up as well. But my marketing friends have confirmed that they use Reddit very heavily, as it's a great sales funnel if you play it right.
[0] Gradually for several years already.
I guess using the correct terminology matters.
If people were actually searching up "peachick" that'd probably be SEO spammed to hell, too.
"Google sucks!"(50 upvotes)
"That's why I use Kagi!"(45 upvotes)
"Actually Kagi has the exact same problem and you have to pay for it."(2 upvotes)
As mentioned, product comparisons are a big one but another worrying area is anything medical related.
I was trying to find research about a medicine I'm taking this week and the already SEO infested results of 5 years ago have become immeasurably worse, with 100s of pages of GPT generated spam trying to attract your click.
I ended up ditching search alltogether and ended up finding a semi-relevant paper on the nih.gov and going through the citations manually to trying and find information.
For content that isn't time sensitive the best trick that I have found is to exclude the last 10-15 years from search results. I've setup a Firefox keyword searches[1] for this, and find myself using them for the majority of my searches, and only use normal search for subjects where the information must be from the last few years. It does penalize "evergreen" pages where sites continuously make minor changes to pages to bump their SEO, which sucks for some old articles at contemporary sites, but for the most part gives much better results.
[1] For example: https://www.google.com/search?q=%s&source=lnt&tbs=cdr%3A1%2C...
OMG. I'm so happy how much AI is improving our lives right now. It really is the future, and that future is bright.
Thanks guys!
Have you reported any of those issues to Kagi (support/discord/user forum)? We are pretty good at dealing with search quality issues.
I’m not affiliated with Kagi in any way and my results are not full of LLM generated content either.
Though there is an LLM generated baby peacock pretty high in the image search, and when I go to the website it literally says that it is an example of an AI generated image and not a real baby peacock.
I've been doing this for years now. The normienet as I call it is nigh worthless, and I don't even bother trying to find information on it.
I wrote some ideas up about this many years ago: https://github.com/pjlsergeant/multimedia-trust-and-certific...
People keep saying this and I keep warning them to be careful what they wish for. The most likely outcome is that "certification of human generated content" arrives in the form of remote attestation where you can't get on the internet unless you're on a device with a cryptographically sealed boot chain that prevents any untrusted code from running and also has your camera on to make sure you're a human. It won't be required by law, but no real websites will let you sign in without it, and any sites that don't use it will be overrun with junk.
I hate this future but it's looking increasingly inevitable.
Personally, I'm excited about more generative AI being added to my search results, and I'll probably switch to whichever search engine ends up with the best version of it.
And of course the first image for "baby peacock" is the same white chick thing… obviously because this story is making the rounds —_—
Search results that are full of content mills serving pre-genned content: no thanks. It's in the same category as those fake stackoverflow scrape sites.
There is plenty of space there for more volunteer editors to verify content, and likewise, WMF operates its own cloud platform where developers are automating tools that do maintenance and transformation on the human-contributed content.
Then, there is Wikidata, a machine-readable Wiki. Many other projects draw data from here, so that it can be localized and presented appropriately. Yet, its UI and SPARQL language are accessible to ordinary users, so have fun verifying the content there, too!
Not only that, we desperately need cryptographic prof that content X was produced by person Y.
Since most people don't write anything on the internet, I can pay people $5 to use their signature and operate a "sign farm".
Look at the effort these people go though to send their spam.
I’m hoping for a “this human being validated this message”. I think that alone could solve multiple problems.
If you put your name on AI spam people will flag your post and no-one will bother to see your posts for at least a few years.
And long before the Internet, people slapped random concoctions together and sold them as medicine, advertising them as cure-alls.
I'm talking about those 300-word, ad-ridden crap articles that are SEO'd right to the top, and if you're lucky you might get the 3-word answer you were looking for: "<300 words of shit>... and in conclusion, <1-step answer>.". Anyway, humans have been getting paid pennies to write those for a while.
AI just turns the throughput on that up to 11, where there's just no end in sight. I think this is like the primary failure mode of AI at this point. It's not going to kill us - we're going to use it to kill the internet. OTOH, maybe then we just go outside and play.
I just bought a home and I have been googling the best way to tackle certain home improvement projects - like how to prepare trim for painting. Virtually every result is some kind of content farm of AI-generated bullshit with advertising between every paragraph, an autoplay video (completely unrelated) that scrolls with the page, a modal popup asking me to accept cookies, a second rapid-fire modal popup asking me to join the newsletter to "become a friend"
For better or worse, Reddit is really the only place to go find legitimate information anymore.
This is frightening and, I fear, true.
But I'd also add one odd little counterpoint: some of the most useful discussions and learning experiences I've had in the last four years have happened in private Facebook groups. As soon as the incentive to build a following using growth-hacking and AI -- which private groups mitigate to a greater extent -- is taken away, you get back to the helpful stuff.
The FreeCAD group on Facebook is great, for example. And there are private photography groups, 3D printing groups, music groups etc., where people have an incentive to be authentic.
Public Facebook feeds are drowning in AI slop. But people who manage their own groups are keeping the spirit alive. It's almost at the point where I think Facebook will ultimately morph into a paid groups platform.
I prefer text content to videos by a long shot, but genuine, human text content is almost dead. Reddit might be one of the rare exceptions for now. There are also random, still active, old school forums for lots of things but they tend to become extremely hard to find.
Would I have preferred a nice 1080p, shot in good lighting on a flat white table? Yes. But those also tend to be 30 minutes along, and as you said, with a sponsorship for HurfDurfVPN in the middle.
For example, https://archive.org/details/stanleyhomerepai0000fine/page/14... links to the chapter "Painting Trim the Right Way" from the book Stanley Home Repairs, 2014.
Could also look at used book stores. Home repair hasn't changed much.
Edit: Could even fire up Wine and try the CD-ROM "Black & Decker Everyday Home Repairs" (published by Broderbund) at https://archive.org/details/BlackDeckerEverydayHomeRepairsBr... . https://www.goodreads.com/book/show/3424503-everyday-home-re... says;
> Like its predecessor in book format, the CD-ROM version offers easy-to-follow, step-by-step instructions on more than 100 common household problems, from how to fix a leaky faucet to repairing hardwood floors. What's more, the CD-ROM version incorporates animation and narration to help make the repair project even easier to understand and complete. Instructions can be viewed one step at a time or all at once, and, if desired, can be printed out and taken directly to the repair site. Included with each repair project is the projected time needed to complete the work, estimated cost, and a list of materials and tools needed.
That sounds pretty nifty, actually!
A trip to the bookstore to buy "x for dummies" can save dozens of hours of web searching.
The current iteration of the internet and AI is lacking depth, detail, and expertise.
You can find 1 million shallow answers on reddit, or echoed in AI, but anything more than the most cursory introduction is buried.
Not only is shallow information easier to generate, it is what most users want, and therefore most engines and services cater to it.
To find better content,you need to go to specialty outlets that don't cater to the lowest common denominator.
The information on working on my new car is non-existent other than Youtube videos where the majority is just a random dude who knows nothing filming himself doing a horrible self repair.
All of the flat white MDF trim you buy is primed and ready for painting, too.
We started by putting advertisements on existing content, then moved to social networking and social media, which was essentially an engine for crowdsourcing the production of greater amounts of content against which to show advertisements. Because money is up for grabs, producing content is now a significant business, and as such, technology is meeting the demand with a way to produce content that is cheaper than the money it can make.
The problem of moderating undesirable human-generated content was already starting to intrude into this business model, but now generative tools are also producing undesirable content faster than moderation can keep up. And at some point of saturation, people will become disinterested, and tools which could previously use algorithmic heuristics to determine which content is good vs bad will begin to become useless. The only way out I can see is something along the lines of human curation of human-generated content. But I'm not sure there is a business model there at the scale the industry demands.
Good fences make for good neighbors.
Its not a coincidence that the printing press brought devastating war to europe in the form of the wars of reformation [1] .
The internet is another real tool for knocking down fences for free, by anyone. Its only a matter of time when there's pushback by angry fence-owners.
We absolutely need less friction and more of minding our own business and focusing on our own back yard instead of chiming in on someone thousands of miles away.
Don't get me wrong, I'm 100% for the free flow of information, but what people (HN crowd?) don't understand is that a significant subgroup of humans cannot tolerate relentless change or challenges to their worldview for too long.
[1] https://en.wikipedia.org/wiki/European_wars_of_religion#Defi...
It’s not just the HN crowd.
For example: https://news.ycombinator.com/item?id=25667362
This is the kicker. When unfettered by regulation or leaders/workers with morals, most industries would rather avoid human curation because they want to sell you something. Amazon sellers would rather you not see or not trust the ratings because they want you to buy their stuff without knowing it's going to fall apart. Amazon makes a profit off it, so they somewhat encourage it (although they also have the dual pressure of knowing that if people distrust Amazon enough they'll leave and go somewhere else, so they have to keep customers somewhat happy).
No, curation has to come from individuals, grassroots organizations, and/or companies without a financial interest in the things being curated - and it has to revolve around a web of trust, because as Reddit has shown, anonymous curation doesn't work once the borderline criminal content marketers find the forum and exploit it.
> The only way out I can see is something along the lines of human curation of human-generated content.
...however, unfortunately, curation doesn't solve the problem of people desiring AI-generated content. That's a much harder problem. Even verifying that something was created by a human in the first place is hard. I don't want to think about that. I'm just going to focus on curation because that's easier and it's also incredibly important for the lowering quality of physical goods as well.
This distinction is important, because while AI is faster than humans, it's at best cheap gateway drugs into skilled human generations.
I see a lot of people talk nostalgically about blogs, but they were an early example of the internet changing from ever green content to churning out articles on content farms. If people remember the early internet, it was more like browsing a library. You weren’t expecting most sites to get updated on a daily - or often even a monthly - basis. Articles were almost always organized by content, not by how recent they were.
Blogging’s hyper-focus on what’s new really changed a lot of that, and many sites got noticeably worse as they switched from focusing on growing a library of evergreen content to focusing on churning out new hits. Online discussions went through a similar process when they changed from forums to Reddit/HN style upvoting. I still have discussions on old forums that are over a decade old. After a few hours on Reddit or HN, the posts drop off the page and all discussion dies.
A lot of older websites actually used to do this with a “what’s new” section or page. With blogging, “what’s new” became the entire site, with almost the entirety of the content (everything that wasn't new) now hidden.
Ironically, after mentioning that discussion dies off incredibly quickly when HN stories fall of the front page, this discussion was moved off the front page to a day old discussion. My guess is that almost no one will see it now.
I remember when forums gradually turned into q&a repos.
And with personalities you have some form of relation to, you want these more recent updates instead of sticking to topics of interest.
Reddit is at least still focused on topics instead of people. I think this is why for some it still is more interesting than platforms like Insta, Facebook or Twitter.
hasn't publishing since Gutenberg been driven by the monetization of content? Looking at the history of the Catholic Church, potentially before that too.
That's retweets.
> undesirable human-generated content was already starting to intrude into this business model, but now generative tools are also producing undesirable content faster than moderation can keep up.
> people will become disinterested, and tools which could previously use algorithmic heuristics to determine which content is good vs bad will begin to become useless.
So what these parts are saying is, tiny monoculture of bored college kids are always going to figure out the algorithm and dominate the platform with porn and spams and chew up all resources, and that both improved toolings and tie-in to monetary incentives intended to empower weaker groups to curb kids only worsens the gap, and that that's problematic because financial influencers are paying to be validated by the masses, not to be humiliated by few content market influencers.
But what is the problem with that? Those "undesirable content" producers are just optimizing for what market values most. If that's problematic, then the existence of the market itself is the problem. What are we going to do with that? Simply destroying it might make sense, perhaps.
Our entire framework of unconscious heuristics for ranking the quality of communicated information being rendered useless overnight may be a recipe for insanity and misery. Virtually nothing has made me this genuinely sad about technology in all my life.
There is only one way I can see things changing and people aren't going to like it. All content on the internet gets linked to a legal ID. Every post on facebook, every comment can be attributed to a real person.
Identity theft would go through the roof.
I don't know what the solution here is. My guess is that the "public internet" will become less and less relevant over time as it becomes overrun by low-quality content, and a lot of communication will move to smaller communities that rely heavily verifying human identity and credentials.
I have gotten used to trusting search results somewhat. Sure there would be oddball results and nearly non-sensical ones, but they would be scarce through a sea of relevant images. Now with this, I would be blind to the things I don't know and as someone who grew up it with Google "just being there", it truly scares me.
The question is what comes next, and I don’t think anyone has an answer to that.
Source: every other piece of news or social media on the planet.
The same could be said for regular search. Pretty much anything I search for yields a page of ads followed by pages of content farmers followed by pages of "almost sounds like experts but is still just a content farmer."
Financing the Internet with advertising has really made it difficult to find good quality content. The incentives are completely misaligned, unless you are a 'content creator' or Google.
Would you pay money to Google or some other company in exchange for a genuinely good search service that prioritizes well written content, and avoids AI (or human) generated crapticles?
But also, if you search the more accurate term, "peachick", you seem to get 100% real images, although half the pages call them "baby peacocks".
However, I'm uncertain whether audiences will truly enjoy this AI-generated content. Personally, I prefer content created by humans—it feels more authentic and engaging to me.
It’s crucial for AI tools to include robust detection mechanisms, such as reliable watermarks, to help other platforms identify AI-generated content. Unfortunately, current detection tools for AI-generated audio are still lacking - https://www.npr.org/2024/04/05/1241446778/deepfake-audio-det...
[Edit] We just put together a list of notebooklm generated "podcasts": https://github.com/ListenNotes/notebooklm-generated-fake-pod...
Consider whether you'd enjoy listening to AI-generated podcasts. I believe people might be okay with shows they create themselves, but are less likely to appreciate 'podcasts' ai-generated by others.
It's disappointing to see scammers and black-hat SEOs already leveraging Notebook LM to mass-produce fake podcasts and distribute them across various platforms.
I'd like to think that too, but I wonder how long - if at all - this will be true. I "want" to like human generated content more, but I suspect AI may be able to optimize for human engagement more, especially for simple dopamine inducing content (like tiktok videos). After all, we're less complicated than we like to think.
>It’s crucial for AI tools to include robust detection mechanisms, such as reliable watermarks, to help other platforms identify AI-generated content.
This will never work, unfortunately. There's no way to exclude rogue actors, and there's plenty of profit in AIs pretending to be human. If anything, we will have to watermark/sign human generated content.
It's not helpful that you're making a binary distinction here.
As an example, as much as 10 years ago, I would find Youtube videos where the narration was entirely TTS. The creators didn't want to use their own voice, and so they wrote the script, and fed it into a TTS system. As you can expect from the state of the art at the time, it sounded terrible. Yet people enjoyed the videos and they had high view counts.
Are we calling this AI-generated?
We now have better TTS (without generative AI). Way better. I presume those types of videos are now better for me to watch. You may still be able to tell it's not a human because the tone doesn't have much variance. You'd probably have to listen for a minute or longer to discern that, though.
Are we calling this AI-generated?
Now with generative AI, we have voices that perhaps you won't be able to identify as AI. But it's all good as long as a human wrote the script, right?
Are we calling this AI-generated?
Finally, take the same video. The creator writes the script, but feels he's not a good writer (or English is not his native tongue, and he likely has lots of grammatical errors). So he passes his script to GPT and asks it to rewrite it - and not just fix grammatical errors but have it improve it, with some end goal in mind ("This will be the script for a popular video...") He then reviews that the essence he was trying to capture was conveyed, and goes ahead with the voice generation.
Is this AI-generated?
To me, all of these are fine, and not in any way inferior to one with a completely human workflow. As long as the creator is a human, and he feels it is conveying what he needed to convey.
I would love to take a first draft of a blog post, send it to GPT, and have it write it for me. The reason I don't is that so far, whatever it produces doesn't have my "voice". It may capture what I meant to say, but the writing style is completely different from mine. If I could get GPT/Claude to mimic my style more, I'd absolutely run with it. Almost no one likes endless editing - especially writers!
About a year back I found that 90% of the results I was getting were AI generated, so I added a flag "No AI" which basically acts as a quick and dirty filter by limiting results to pre-2022. It's not perfect but it works as a stopgap measure.
- autogenerates URLs (tha look legit)
- autogenerates content for such URLs (that look kinda legit)
All of this would be possible if one is using Chrome (otherwise the fake URLs wouldn't lead to anywhere). Of course, full of ads.
Think about it, some people are not really looking for some web site that talks about "baby peacocks". They are looking for baby peacocks: content, images, video. If Google can autogenerate good-enough content, then these kind of users would be satisfied (may not even notice the difference).
Maybe Google ditches the URL and all: type keywords, and get content (with ads)!
Didn't they do something like that with AMP. I recall that if you were using chrome and visited an AMP site from Google the address bar would say site.com even though the content was being served from google.com.
Whether they actually do this (and whether there's any incentive to do so), is obviously not a given
The only way we can make sure the internet retains any goodness is by contributing good things to it. Passive consumption will rapidly turn into sub-mediocre drudgery. I suppose it already has.
Be the change you want to see, I guess. I’m a shitty writer, but at least I can beat the dissonant, bland, formulaic rambling of ChatGPT (here’s hoping, anyway).
I’m optimistic that a lot of us can keep something good going. We'll find ways to keep pockets of internet worth visiting, just like we did before search engines worked well.
I think eventually all digital cameras and image scanners will securely hash and sign images just as forensic cameras do to certify that an image was "captured" instead of generated.
Of course this leaves a grey area for image editing applications such as Photoshop, so there may also need to be some other level of certificate base signing introduced there as well.
I dare say that I haven’t noticed that much of a change in things and that could either be because LLMs are just that good at Reddit content, or that because Reddit was already so botted and manipulated it didn’t really change much.
I sure hope the money that Reddit made makes up for the readers who are fleeing.
So the pessimist in me can see the Internet being affected by the free-vs-premium formula: "basic" Internet with ads, tracking, AI fillers, limited access to +18 content, in the worst form comes with these pre-defined sites and "premium" that's free of these limitations but it also in time tries to squeeze more money from users - like "premium but with ads"
I know there are no girls on the internet, but this AI crap is on another level. Even if find a trustworthy creator, I might be seeing a fake video of them. Say I like MKBHD reviews, I will need to pay attention if I am really watching his video on his official channel.
My guard will have to be up so much, all the time, I actually don't think it will even be healthy to "consume content" anymore. Why live a life where almost everything I see can be a lie? Makes me not want to use any of this anymore.
While I generally agree with your whole comment, I feel like this part has been true for years on social media well before AI generated content hit the scene.
Maybe I am overall in a bad mood regarding all this, but this recent article https://time.com/7026050/chatgpt-quit-teaching-ai-essay/ also struck a chord. Do I really wanna spend my time reading/watching machines talking to each other? How long until browsing Reddit or HN will be worthless?
Do I wanna get old with lower cognitive abilities and become this? https://slate.com/advice/2024/10/grandparents-misinformation...
It effectively kills the non-walled internet as an information repository.
https://commons.wikimedia.org/wiki/Category:Pavo_cristatus_(...
Wikimedia is a fantastic resource.
It also includes Google Images and Flickr.
https://search.creativecommons.org/
I found the peachicks on Commons by searching "peacock" and then following categories up the tree. If people use the wrong search engine with naïve search terms, I don't know what to tell ya.
This is a parallel example of why reference librarians are still worth consulting, because they will guide you to the library's resources and databases, and demonstrate how to use search queries.
Yandex images search is flawless though.
A really quick fix is to search with “-ai” and that Google doesn’t do this implicitly for images is really strange.
So do humans.
If Google prioritizes AI slop, Google will be deprioritized.
I massively rate Kagi, but this is way less than ideal.
If we could subscribe and suggest content along our interest graphs, we would control the algorithm and could prune slop with ease.
It'd be incredibly awesome if news, forums, and social media worked like BitTorrent.
DDG lets you turn ads off completely.
My search lists are curated very well through my settings and even just using the recommended block list keeps a lot of junk out of my search results. If I find a bad site, I can block it from all future results pretty quickly. I also can use regex on the URL's in the search result to redirect things like Reddit to old.reddit automatically. It's very nice.
The masses are fully here now. They're too passive to know or care what's going on. They stick with the path of least resistance: Google, Amazon, Reddit, Twitter, etc. No matter how hostile or shitty those options become.
We have to put aside the way we've thought about the internet before now because it doesn't apply anymore. There will be no more MySpace -> Facebook. The internet is no longer made up of a high enough percentage of conscientious and deliberate users to make a difference.
The problem is going to get worse as hallucinations are used as training data because even the AI companies can't tell the difference between AI content and human content.
There was a period in the past when human spam was a problem that was not trivial to solve.
As always, modern problems require modern solutions.
AI? This shall pass too. Internet will find its way.
Our best bet is to have scraped all that data, and give you a temporal parameter to search, like:
+"Sponge bob" year:2012
I'm starting to think that all this AI stuff has finally pushed the ads-based Internet past its tipping point.
I feel I could be motivated to work on a walled garden with moderation paid for by subscription fees. What would it be worth to you to have an entirely new online experience free of all the enshittification of the past 15 years?
Personally, I pay for Kagi just to have a small taste of what that could be like. But what if not just the search engine, but also all the sites be funded entirely by a subscription fee paid to the service profider? What if privacy could be a foremost feature of that world? What if advertising and astroturfing were strictly forbidden, and human authors would have to be vetted by other humans to be allowed a place in this world? "This content is Certified ads- and AI-Free(tm)."
I really don't know how well something like that would turn out in 2024, but I feel I wouldn't be alone in wanting to give it a try.
It has to be done in a decentralised way to ensure no enterprise controls who is trusted and who isn't.
I know there are exceptions. There are answers I've wanted that can be found within the first few minutes of the first video on Youtube, which I've gone days without discovering because I'm video-averse. But I suspect that the habit is, on average, more benefit than detriment.
On the sad side the TikTok and YouTube ones that likely led to all of this aren't labeled and are present, not to mention the complete lack of "I want the AI things automatically filtered, I'm not interested in trends I'm searching for actual things right now" button. Without something like that it will become harder to use Google to find new content.
I mean people obviously like the content, it's cute enough to get shared around so much to make itself popular in these images and to trigger the post on X about it. Nothing wrong with that... but if it's not easily filterable for what the user is actually trying to find then Google has somewhat failed at its goal.
Have people tried searching for other animals? Maybe this isn't a case of Google being inundated with AI-generated photos, but just something to do with the results for this particular phrase.
Search and Internet is dead. It will be. There is no going back with AI. We must to learn how to deal with. You too should rethink how to approach the Internet, how to surf it.
If search is dead, are there any solutions to it? I use more RSS source now, because this is human created content. I navigate more to "word of mouth".
Mostly I think about how something like that is going to be signed into law by some state and it'll require everything you do to be linked to your government issued ID card so they can "prove" you're not spreading AI misinformation and all the horrendous unintended side effects that will spread from there.
...
"Asynchronous, symmetrically anonymized, moderated open-cry repute auction. Don't even bother trying to parse that. The acronym is pre-Reconstitution. There hasn't been a true asamocra for 3600 years. Instead we do other things that serve the same purpose and we call them by the old name. In most cases, it takes a few days for a provably irreversible phase transition to occur in the reputon glass - never mind - and another day after that to make sure you aren't just being spoofed by ephemeral stochastic nucleation."
Fantastic book. I read it twice so far, highly recommended. So many little off-handed conceptual gems everywhere.
Searching for 'baby dog' would probably get you garbage images too. (it does)
In any case the web and Google's index of it is crowdsourced. If the web associates this image and that phrase, what are they supposed to do about it?
https://trends.google.com/trends/explore?date=today%203-m&ge...
Tech: "Gosh we better tune our algos so these images are even MORE indistinguishable from the real thing"
Evidently the road to hell is paved with novelty image generators.
Same, as far as I could tell all AI garbage with weird saturation and colors and uncanny valley .... they look weird / didn't work for me.
I predict the word meat fucked with the AI
And then I remembered that I was on duckduckgo.
Or is this one of those fundamental attribution error things:
- MY product is a powerful tool for creators who wish to save time
- THEIR product is just a poorly-though-out slop generator
Does it occur to people to instead be part of something real and visceral, and not just blame social media's ad-driven impression model, not pretend they are only part of a trend for which they can't be totally blamed?
You could say what you say about anyone at any time. Where do you draw the line? I guarantee you'll be guilty of the exact same thing. I don't want to generalize, but IMO this sentiment of yours, I hear most loudly from software engineers far removed from ordinary non-technical end users: is making beautiful new LISPs and CNIs and Python package auditing tools the only valid work with seemingly no tradeoffs?
I don't sincerely believe that people who are working on Kubernetes features or observability tools are bad people. Do high drama personalities who engage in a mode of discourse of "wow" and "shockingly" say valid things too? Yeah. But it's as simple as, log in your own eye before you worry about the thorns in others. Exceptionally ironic because the poster is vamping about "Attribution errors." Another POV is, shysters project.
No, everything is not the same as everything else.
I am absolutely not far removed from non-technical end users. They are my client base, ultimately. As a freelancer I focus on building real things that make things better for people whose faces and voices I get to know. GenAI will be useless to them, because it is antithetical to what they do.
And that focus is only getting keener; I want nothing to do with the AI-generated web.
So what I'm hearing is, "I agree very strongly with the people who pay me." Or to put it in your words:
"MY product is a powerful tool for creators who wish to save time."
"THEIR product is just a poorly-thought-out slop generator"
Sure. And THEIR products are just thoughtless slop generators.
- I am a thoughtful technologist, building real things for real people, concerned about others and the social impact of my work;
- they are greedy and ignorant, destroying society for short-term personal gain, no matter what the consequences.
It's human nature to put badness on an abstract them, but we don't get anywhere that way. It's good for getting agreement (e.g. upvotes), because we all put ourselves in that sweet I bucket and participate in the down-with-them feeling. But it only leads to more of what everyone decries.
I did not make any claims about myself at all, until I was separately accused of being something or other by someone projecting onto me whatever it was they needed to feel better about themselves.
Second, you have rate-limited me with the "posting too fast" thing so I couldn't reply to your comment or other ad hominem, even though I was posting at a rate no faster than the discussions about OpenSCAD and FreeCAD I had been involved with earlier (considerably less, I would say).
It's IMO really classless to use your administrative privileges to silence people after you accuse them of something but before they can respond, but I am not surprised to see that.
I will repeat again: I think it is really clear to me, and really to everyone I have me outside this bubble, that there is no fine distinction to be drawn between content generating AI projects that are "good" and those that are contributing to "slop". It's all slop-generation; e.g. NotebookLM is no better or cleverer than Midjourney.
Every tool HNers are excited about is going to be used to make the world's culture, and the web, worse.
I'd encourage you and those reading to consider this.
Sure, you can't make much of a change by yourself. But you don't have to be part of what amounts to inflicting automated cultural vandalism on an unprecedented scale.
Goodbye.
That said, I don't really think this is a tide any individual market actor can reasonably stem. It's going to require some pretty fundamental changes in the way we use the internet.
You talk about being a part of something "real and visceral" but you're complaining about the demise of being able to sit at your desk and see pictures of wildlife. Maybe it's okay that google image search dies and makes people go out and find the wildlife they want to see.
The internet, even in its best format (e.g. ad-free, free access information for all; and communication with all of humanity) has a ton of real downsides. It's not clear to me that AI should be strangled in its infancy to save the internet (which does _not_ exist in that "best" format).
I don't think that is what will happen if google images dies.
Sick and tired of giving parasites benefit of the doubt they've long sucked dry.
In early 2010s when Instagram, Twitter, Facebook started getting big, all the websites and apps had this process of discovery that you had to go through to make it fun for yourself. It obviously turned some people off of it, and made the onboarding a bit harder, but you needed to follow some people, send some friend requests, and in the end you would mostly see things you've actively wanted to see. Even when the algorithms started sorting the timelines, it would still be (mostly) within the things you've chosen to see. Even Youtube's recommendation algorithm was pretty simple, and it would suggest extremely similar videos.
I think it changed around 2016, when the algorithms started trying to determine what you like, based on your interaction with other things, rather than your explicit action of saying "i want stuff from this person/channel/etc.". I'm sure a significant chunk of us have worked on similar algorithms, so you get the gist of it. But this change resulted in users getting attention from the global audience (because in order for algo to detect what you like, it has to throw in suggestions from everywhere).
I get that forums have existed for decades, and people were getting Reddit karma since 2000s, but it was still more deliberate action when you wanted to see something. TikTok, YouTube and Instagram changed the entire playing field in the last 6 years or so, where your real life "social score" didn't have to be depend on whom you know in real world for anyone. It translates into - you can generate posts, content, whatever you wanna call it, for everyone rather than actively getting someone's attention. Like, going viral on YouTube was a big thing at some point. There's some ongoing meme-like comments saying "you would be invited to Ellen's show in 2010", which is kinda true because breaking out of the "only seen by people whom you know" box was extremely rare.
Well, now, everyone, technically has a chance, which incentivizes people to constantly push out content. It doesn't matter, if you're doing it for just social media clout, or financial motives, and etc. It's just possible for something to go "big", albeit for minuscule benefits from it. So there's constant churn of... content. And now AI is just making it even simpler to create such content. But again, resulting in even further decrease of social importance of such pictures/videos/texts.
I understand there's always a group of people that "write/create/paint for themselves", which I understand. I'm on a similar boat. But the if majority of creators have different incentives, the platforms will cater for them. And in this case, platform is the whole Internet, and incentives are "financial, and seeking global attention". Right now, it takes about a minute to create a video and post it on any of the websites, which was basically impossible back in the day. That barrier of entry, combined with one's deliberate discovery what, I think, was making the internet look more fun.
I'm not touching the subject of ad-infestation in every corner, and it definitely accelerated the downward spiral of average quality of content. But in the end, I blame ourselves for choosing this path, because we could've put pressure on global-algorithms of YouTube, TikTok, and etc. We chose to not to do so, because, well, it still gives us dopamine hits.
Like what solutions are we gonna come up with to solve it? Is the human side of the internet (however we create it) going to become more pure? Perhaps in discovering ways to avoid low quality AI content, we'll also find ways to escape from destructive recommender systems and monetized advertisements as well. Strange as it sounds, solving this problem could lead us to a much brighter future!
(also noteworthy for the 'Publication' section near the bottom)
We have been creating our own reality even before AI.
This is already impossible because it's impossible to enforce. You can't stop something running on a random laptop, and you can't stop models running on server farms in, say, North Korea.
If it cant verify the source it could label it suspect :-). Just thinking here ... you got any other ideas or we are just going to let the Internet die by the hands of AI as Neil DeGrasse Tyson predicts https://www.youtube.com/watch?v=SAuDmBYwLq4 or you just gonna downvote someone who tries to come up with solutions.