Oh, I should mention though. There was no advertising at all. They didn't make any money off me. It was 100% Gemini which I recognize as not long-term feasible.
Oh, I should mention though. There was no advertising at all. They didn't make any money off me. It was 100% Gemini which I recognize as not long-term feasible.
If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content?
I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me.
Are you actually saying you'd be OK with that?
The last one is the worry.
The middle is... where we all start.
The first is not a bad place to be, all things considered!
Empirically, however, LLMs don't strip out the author: the big models know a lot about what I've written even with search disabled. Ex: https://claude.ai/share/8cbcdf88-a360-421a-8c06-ae7b7992e866
(I was pointing out that "they'll train on a version that strips out you as the author" seems to be is incorrect about how training works)
My experience being on the searching end is that these things are terrible about attribution of where they find anything. Which has bad consequences not just for authorship, but for correctness (which is the usual reason I'm poking at them -- they're being wrong again). This makes a lot of sense when you consider the massive, massive compression that's got to occur during training, but it's still frustrating.
>> Without searching, what can you tell me about exmadscientist? Fine if you get this hallucinate or get things wrong, I just want to see what level of knowledge you have trained-in.
>Working purely from trained-in knowledge, with the caveat that this is exactly the kind of long-tail detail I might garble:
>"exmadscientist" rings a bell as a username on Hacker News. My impression is of a fairly prolific commenter who writes knowledgeable, detailed comments about electrical engineering and hardware topics — things like PCB design, electronics manufacturing, component selection, test equipment, and general "how things actually get built in industry" commentary. The comments I associate with that handle tend to be practical, opinionated, and grounded in real professional experience rather than hobbyist speculation.
>Beyond that, I don't have confident specifics: I couldn't tell you a real name, employer, location, or any particular famous comment or post they've made. It's also possible the same handle exists on Reddit or other forums and I'm blending impressions together.
>So the honest summary of my knowledge level: I recognize the username and have a general "HN hardware/EE person" association, but nothing more granular than that — and even the association could be a mild confabulation, since usernames like this sit right at the edge of what a model reliably retains.
(I don't know if you have a blog or otherwise write on the internet; I just asked it about your HN handle)
Yeah, authors don't want to be recognized as authors, they don't want any reward for their work, they don't want to amass pool of loyal readers, interact with them, etc.
All they want is for halucinating AI to take excerpts of their work and compile it with random sh!t.
GENIUS
LLMs just use everything, generate similar code with no attribution and keep users from visiting, so no bragging rights or attention.
Worse, there are some PRs that seem fully generated ...
So i mostly stopped sharing and started pulling my old repos offline.
At this pace, i don't want to compete with a clone of myself in the future that will do my work for much cheaper.
Sure. But you can see that for some people (myself included), writing for peers is part of the joy? And that if instead a megacorp places an opaque computer program between the author and the readers, that joy might be ruined?
I might never blog/publish code again amidst all this. I never had ads on my sites. I am not alone in this.
Sites like Wikipedia or developer documentation pages which exist to distribute knowledge for its own sake don't have any reason to care whether that knowledge is consumed by a human or a computer being used by a human.
I'm sure the billionaire class would love a return to patronage based libraries, NDAs on authors of books, and the elitism they would feel with a return to private libraries locking away all kinds of knowledge that would happen if patronage become the only way authors could make money (such as with AI just regurgitating their works, or if the stupid 'do away with copyright' people got their way).
Cheap access came from the invention of cheap printing . The laws were passed to restrict it.
Making the avenue of creating for the average person also the avenue for the most income was huge in creating our modern literature landscape.
Evidence for this statement?
> If the only money is in private works for private libraries, that is where the quality stuff is going to go.
Evidence that this ever happened?
> Making the avenue of creating for the average person also the avenue for the most income was huge in creating our modern literature landscape.
That only leads to higher quality (as you claim) if your definition of higher quality is "what the average person buys".
When people wont see others blogs, they wont start writing own. When there will bw no ome to actually read it, they will go to do something else.
But maybe that is the future.
Every country starts erecting their own towers of babel that we talk at, and it constantly compresses our conversations down to the most effective distribution of weights.
At some point talking at the machine becomes a high status job, and we give respect to the people who whisper to it the most.
Theres many people I know who write notes that I know would be great to read. However they never publish them.
So… are we saying that the only public writing in the future is meant to be consumed by the machine?
Some of the proposals to address this include charging bots for access to web resources, but they will also have repercussions for regular users. I don't see how you solve this cleanly.
Not sure of the effectiveness but it's there.
I think the main thing Cloudflare is trying to do is block direct traffic from frontier labs and then start charging them for access. They might end up shooting themselves in the foot, as this simply empowers sketchy residential-proxy outfits to undercut Cloudflare and sell the data to labs for less.
The problematic bots are all disguising themselves as Chrome and sending requests from millions of residential proxy IPs, and the only real solution to those is some sort of captcha or PoW page on first visit.
There could be open source tooling to create custom private "closednets", with
- trust ring mechanism to allow invitations, flagging, banning, and banning those that invite people who were banned
- the rules of the closednet
- search engine with opt-in scraping
- portal (remember the 80s?) with all the registered nodes, perhaps by service category such as public git repo hosts, web sites etc.
etc.
The first closednet could be Hacker News.
If it had any real value, anyway.
Small, truly private communities could be an interesting thing though.
Yeah, Network by Humans for Humans. Thats why Im not interested in all those IoT/Auto networks when you just connect and stuff automagically configure. It looks nice at first glance, but you loose control. F2F works way better in that matter, like RetroShare, but I never investigated it much.
VPNs are great for torrenting but any serious website like an online bank or web email provider will turn you away. They claim it's for bots but really it because they only want customers they can track.
Yep. IMO, this is so far the biggest AI-inflicted damage to the web. A bit of anecdata - wikipedia (and all other wikimedia sites) are blocking my Firefox since about a week, with a "please respect our bot policy" message. Outright block, not even a captcha.
It took me a while to figure out they don't like me disabling some SSL ciphers, so now "JA4 browser fingerprint" is not matching user-agent. Funnily enough curl (what I would imagine a bot would use) pulls exact same URLs from exact same client IP, just fine.
But we already have the latter case that exists - ad blockers. Ad blockers literally serve up the word-for-word original content minus the ads.
Sure, humans would benefit.
It took them searching, reading themselves, maybe even understanding something in the process, to complete a 360° revolution of their squirrel cages in time T.
Now they can omit searching, skip reading to the regurgitated answer, throw away understanding, and complete a full revolution in T/N, where N is a heuristic value directly proportional to the amount of skin in the AI hype.
But the catch is that the squirrel cage must run non-stop still.
obviously new content still has value because it remains the source layer for LLM agents. it just wont be ads giving you revenues thats all.
That expectation is a problem, has always been a problem, and Tim Berners Lee never mentioned anything about a reward structure when coming up with the WWW.
Should be "You're thinking".
This should be "This should be 'You're thinking'." don't you think? Why bother correcting someone's grammar with a sentence fragment? You're just trading one mistake for another. I'm hoping someone finds a grammar error in my post, because continuing this would be hilarious.
Reflexively, I think it should be more like ...
javascript: `This should be "You're thinking".` ;
// to preserve the original character use and to avoid '...'...' parse foos
// however `"...".` also possibly deserves a [sic] to critique the original
// i.e. ~grammar police say the period belongs within the quote marks, no?
... but then that's just me, in [my] quirks mode.In that case the creator should welcome AIs with open arms; a human reader will forget eventually, but the AI will preserve the knowledge forever.
no, only some mangled form of it
A lot of us hate that with a white-hot passion akin to the eye-melting intensity of an arc welder.
If you don't understand that, then yes, I expect you're quite happy with LLMs, and indeed find inexplicable the reactions of those who are emphatically not happy with genAI.
It's not even about votes, I don't even need votes. Whenever I write something that I am happy with, I read and reread it imagining I am reading it as a third person. Sometimes it forces me to rework my arguments. I wouldn't write to convey my ideas through a chatbot. And that's what this post is about-- killing the internet and with it decimating any audience you might have accrued if you had something to say and you published it on a website.
It’s like bragging about a new highly addictive psychedelic drug that a dealer gave you a taste of for free. The effects are awesome today, you feel so fun and free! Never mind that it’s destroying your body and that the dealer will eventually charge you or demand you pay in other ways, that’s a problem for another day. Weeee!
However, when it comes to non-creative things (such as hooking up some hardware ABC to some software XYZ using RST), LLMs might be better at digesting the factory manuals (hopefully THOSE were written by humans) and explaining it in a way the user understands for their specific case.
I had the similar experience to yours yesterday and it lead nowhere. Funnily enough I was also trying to configure a vpn on a router, google didn't return anything useful (besides a blog post clearly written by AI and with absolutely no information in it). Claude managed to give some interesting pointers, but its suggestions were not working and I also noticed that it started to hallucinate badly about ipv6 and gave me some suggestions that were just plain untrue. Claude Opus is smart, usually when it gets so convinced about something is after researching the internet and not just based on its training data. I wonder where it got so convinced about it. Maybe reading some other hallucinated blog post like the one I stumbled upon?
As time goes on, more and more people will recognize this problem and we'll develop new ways of measuring information quality and trustworthiness. Nothing about this problem is fundamental, it's just that we're in the middle of a very chaotic transition.
To be more precise, I hate the SEO shithole the internet has become, that Google serves up, that Google facilitated, indirectly created.
(I really don't have any tears to shed if there is a death of the Corporate Internet™.)
We're likely at the "golden age" of LLM-assisted web searching and summarization.
Hopefully open models keep it cracked open, but expecting enshittification is always the safe bet these days.
That is why we are seeing paid streaming services with ads.
Very happy with Kagi personally
> One you pay for yourself !
SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.
If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
They do this because they benefit from their site being visited or the information they are providing being noticed.
> If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
I see no reason it would have that effect. It does, however, create different incentives for the search provider to improve the signals indicating page relevance since the user is the priority instead of advertisers.
>>>> Is there some search engine that you think could have become popular and not ended up with SEO optimization?
>>> One you pay for yourself !
>> SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.
> They do this because they benefit from their site being visited or the information they are providing being noticed.
>> If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
> I see no reason it would have that effect.
Agreed.
It's nice to be the customer instead of being the product, for once.
We used to know better, Standard Oil vertical integration was dismantled.
That is to say, it can only get worse from here.
They're a good jumping off point, but I need to delve into the original sources just like I did when I used Google.
It seems to get shell scripts right most of the time.
He said, proud of his own ignorance.
I would be the first to admit my ignorance on the absolute majority of topics. There is a limited number of things I can learn in life, and kubernetes won't be one of them - I'm just not interested in it (and all the other infra stuff, to be honest), as long as it works.
> Oh, I should mention though. There was no advertising at all. They didn't make any money off me
Are you sure about that? Even if you didn't see ads ( remember people pay even if you don't click - just like a billboard ) - they are still profiling you to better sell you ads in the future, and using your interaction as free training data.
YMMV.
But yeah, text remix machines are not a long-term solution to that problem.
It's the same reason why we don't give students the answers to things, we teach them to find the answers.
And somewhere between ten and seventy percent of what it teaches you will be anywhere from subtly wrong in minor ways to utterly flaming bullshit, and unless you go through the boring, slow, old-fashioned learning methods, you'll have no way to know.
I end up using the gemini api for with search enabled for the cases that I don't have access to good grounding data even in agentic tasks.