AI paid for by Ads – the GPT-4o mini inflection point
batchmon.com
batchmon.com
But this also means that because we've exhausted the human generated content by now as means of training LLMs, new models will start getting trained with mostly the output of other LLMs, again because the web (as well as books and everything else) will be more and more LLM-generated. This will end up with very interesting results --not good, just interesting-- akin to how the message changes when kids the telephone game.
So the snapshot of the web as it was in 2023 will be the last time we had original content, as soon we will have stop producing new content and just recycling existing content.
So long, web, we hardly knew ya!
If the AI generations are correct, is it really that bad? If they're bad, I feel like they're destined to fall to the bottom like the accidental Facebook uploads and misinformed "experts" of yesteryear.
In a certain sense, it doesn't really need it. I like to think of the Library of Babel as a grounding thought experiment; technically, every truth and lie could have already been written. Auguring the truth from randomness is possible, even if only briefly and randomly. The existence of LLMs and tokenized text do a really good job of turning statistics-soup into readable text.
That's not to say AI will always be correct, or even that it's capable of consistent performance. But if an AI-generated explanation of a particular topic is exemplary beyond all human attempts, I don't think it's fair to down-rank as long as the text is correct.
If somebody had a site that we were not indexing and wanted to be, they could pay a human to review it every few months.
You can record as many albums as you like as well, but the DJ needs to like your music before they play it on the radio.
The thing with the AI content boom is that if there’s 1000x more of it than there is genuine indie stations, it gets harder to find the real content. Piping things through a top25 filter doesn’t fix that, or actively makes it worse due to the incentives to monopolize / plan the system.
I mean us as in a network of trusted individuals.
For example, i've been appending "site:reddit.com" to some of my Google queries for a while now —especially when searching for things like reviews— because, otherwise, Google search results are unusable: ads disguised as fake "reviews" rank higher than actual reviews made by people, which is what i'm interested in.
I wouldn't be surprised if we evolve some similar adaptations to deal the flood of AI-generated shit. Like favoring closer-knit communities of people we trust, and penalizing AI sludge when it sips in.
It's still sad though. In the meantime, we might lose a lot of minds to this. Entire generations perhaps. Watching older people fall for AI-generated trash on Facebook is painful. I hope we acted sooner.
For all I know your reply is also a botted response to promote reddit reviews as trustworthy and bot-free :P
To put it another way: who defines the trust network?
Or another way: every trust network will be invaded.
Or another way: trust is already actively exploited and has been for decades (or longer, if you want to go there....)
I guess for me, so far at least, some sites feel much more legit and human that the obviously bot-ridden mess that are the likes of Twitter/Instagram/FB. Like for example here or on Lobsters (and more on the latter) i have the feeling that it's mostly people talking with people. On the couple of relatively-small subreddits i visit, i feel the same too.
But i could be wrong of course. Maybe the tone of an HN poster is super easy for an LLM to copy; there's a reason why "shit HN says" exists after all. The only reason i have to believe otherwise is that, in comparison, Instagram or Twitter bots are so obvious and bland, and those companies have way more resources to throw at AI than HN or reddit :P
Doesn't mean anyone buys it.
We’re doomering in this here thread.
/s
The AI is already tainted with human output.... If you think its spitting out garbage it's because that's what we fed it.
There is the old Carlin bit about "for there to be an average intelligence, half of the people need to be below it".
Maybe we should not call it AI rather AM, Artificial Mediocrity, it would be reflection of its source material.
This is true for the median, not necessarily for the average.
How many people are below the average IQ?
Likewise with record labels if platforms like Spotify which allow self-publishing get overwhelmed with Suno slop, which is already on the rise (there's some conspiracy theories that Spotify themselves are making it, but there's more than enough opportunistic grifters in the world who could be trying to get rich quick by spamming it).
https://old.reddit.com/r/Jazz/comments/1dxj409/is_spotify_us...
There is also a rapidly growing industry of people whose job it is to write content to train LMs against. I totally expect this to be a growing source of training data at the frontier instead of more generic crap from the internet.
Smaller models will probably stay trained on bigger models, however.
Do you have an example of this?
How do they differentiate content written by a person v/s written by LLM, I'd expect there is going to be people trying to "cheat" by using LLMs to generate content.
Honestly, not sure how to test it, but this is B2B contracts, so hopefully there's some quality control. It's part of the broad "training data labeling" business, so presumably the industry has some terms in contracts.
ScaleAI, Appen are big providers that have worked with OpenAI, Google, etc.
https://openai.com/index/openai-partners-with-scale-to-provi...
This is just an absurd idea. We're going to just stop producing new content?
Observable: ChatGPT quite often used to just outright says "As a large language model trained by OpenAI…", which is a dead giveaway.
The actual training process makes the model output be the likeliest output, and the introduction phrase you quoted would not come out of this process if there was no RLHF. See GPT3 (text-davinci-003 via API) which didn't have RLHF and would not say this, vs. ChatGPT which is fine-tuned for human preferences and thus will output such giveaways.
I see no reason to believe it wouldn’t be a pendulum situation.
That’s how GANs work, after all.
Human generated content will be outpaced by AI generated content by a large margin, so even though there'll still be human content, it'll be meaningless on aggregate.
It's a very "not ready for primetime" feature
Usually doing `g $query` right after gives me at least some useful results (even when using double quotes, which aren't guaranteed to work always).
Happens about 200 times a day (0.04% of queries), very painful for the user we know, still trying to find root cause (we have limited debugging capabilities as not storing much information). it is on top of our minds.
Maybe give an option to those users who are reporting bugs to pass more debug info if the user agrees.
That's a bit of fantasy given the amount of poorly written SEO junk that was churned out of content farms by humans typing words with a keyboard.
The internet is an SEO landfill (2019) https://news.ycombinator.com/item?id=20256764 ( 598 points by itom on June 23, 2019 | 426 comments )
The top comment is:
> Google any recipe, and there are at least 5 paragraphs (usually a lot more) of copy that no one will ever read, and isn't even meant for human consumption. Google "How to learn x", and you'll usually get copy written by people who know nothing about the subject, and maybe browsed Amazon for 30 minutes as research. Real, useful results that used to be the norm for Google are becoming more and more rare as time goes by.
> We're bombarding ourselves with walls of human-unreadable English that we're supposed to ignore. It's like something from a stupid old sci-fi story.
That, to me, is the biggest difference. Previously I was mostly sure that something I read couldn’t have been generated by a computer. Now I’m fairly certain that I would be fooled quite frequently.
Shorter version: intentional bullshtting never ends, it's in human, and AI, nature. Like it or not. Having several sources used to help, but now with flood of generated content it may be not the case anymore. If used right this has real affect on business. That's how small sellers live and die on Amazon.
- there isn't a world government to enact such laws
- people would break those unenforceable laws
But perhaps I'm wrong. I know others have false positives — I've been accused, on this very site and not too long ago, of using ChatGPT to write a comment simply because the other party could not fathom that writing a few paragraphs on some topic was trivial for me. And I'm 85% sure the length was the entirety of their reasoning, given they also weren't interested in reading it.
In orchards bathed in morning light, Where verdant leaves and branches sway, The tangerine, a gem so bright, Awaits the dawn of a golden day.
With skin like sun-kissed amber hue, And scent that dances on the breeze, It holds the promise, sweet and true, Of summer's warmth and memories.
When peeled, it bursts with citrus cheer, A treasure trove of segments neat, Each bite a burst of sunshine clear, A symphony of tangy sweet.
Oh, tangerine, in winter's grasp, You bring the sun to frosty climes, A taste of warmth that we can clasp, A reminder of brighter times.
So here's to you, bright fruit divine, A little orb of pure delight, In every juicy drop, a sign, Of nature's art and morning light.
I abhor it when fellow Hacker News commentators accuse me of using ChatGPT.
Do twisted dreams linger, of what it might mean to be a taste on the memory of a forgotten alien tongue?
Is its sacred role seen -- illuminated amongst the greens and unique chaotic chrominance bouncing ancient wisdom between the neighboring leaves?
The tangerine -- victim, pawn, and, ultimately, master ; its search for self in an infinitely growing pile of mixed up words truly complete. There is much to learn.
Fuck tangerines, man. Those little orange bastards are a pain in the ass to peel. You spend 10 minutes trying to get that leathery skin off, your fingernails getting all sticky with that goddamn citrus juice. And then when you finally wrestle one of those fuckers open, you got all those little white strings hanging off everywhere. It's like dental floss from hell.
And don't even get me started on how those tangerine slices always shoot juice in your eye when you try to eat them. It's like getting maced by a tiny orange asshole. You ever get tangerine juice in your eye? Shit burns like the fires of hell itself. Makes you want to claw your own goddamn eyeballs out.
Nah, fuck tangerines and their whole stupid family tree. Oranges, clementines, satsumas - they can all go straight to fruit hell as far as I'm concerned. Give me a nice safe banana or an apple any day over those sadistic little citrus grenades. Tangerines are the work of the devil, plain and simple. Fuck writing poems about them little torture devices.
this rant didn't remind me of George Carlin but I still laughed anyway
How long will it be, before humans reading mostly LLM output, adopt that same writing style? Certainly, for people growing up today, they will be affected.
The only one that wants 1 answer per view is from a propaganda perspective. Where truth is politicized and no longer facts, but opinions.
One of the many surprising things to me about ChatGPT when it was first released was how well, in its default style, it imitated the bland but well-organized writing style of high school composition textbooks: a clearly stated thesis at the beginning, a topic sentence for each paragraph, a concluding paragraph that often begins "In conclusion."
I mentioned that last point—the concluding "In conclusion"—as an indicator of AI writing to a university class I taught last semester, and a student from Sweden said that he had been taught in school to use that phrase when writing in English.
If I see HN comments that have final paragraphs beginning with "In conclusion" I will still suspect that an LLM has been used. Occasionally I might be wrong, though.
An entertaining informative style of speech can detract from clearly communicating substance. (Of course, the audience rarely wants substance.)
Just imagine 180M users of chatGPT having an estimated 1B sessions per month. The model is putting 1-2Trillion tokens into people's brains. People don't assimilate just the writing style and ideas, but also take actions into the real world influenced by the model. Sometimes they create useful discoveries or inventions that end up on the internet and in the next scrape. Full cycle.
From what I’ve seen (tutoring high school kids), the picture is much bleaker. They use ChatGPT to write for them but they have no writing style of their own. They can barely put a sentence together just to write the prompt!
Pro Tip: Use a model like llama3 to ‘humanize’ text.
Llama is trained with Metas data sets so you get more of a natural sounding, conversational tone.
I think a lot of the material was from standardized testing.
This very structured writing style. Many paragraphs, each discussing one aspect, finished by a conclusion. This is the classic style taught for (American at least) standardized testing, be it SAT, GRE, TOEFL, et al.
The yellow stem toward my hand?
Come, let me clutch thee:
I have signal not, and yet I taste thee still.
In fact the really bad spammers were already re-using prompts/templates, think of how many of those recipe novellas shared the same beats. "It was my favorite childhood comfort food", "Cooked with my grandma", blah blah blah
why does that distinction matter?
Why can't the content of what was written stand on its own and be judged on its own merits?
Actually original commentary in a discussion is bloody hard to come by.
Human output signal might be wildly different from person to person if judged on originality. But LLM output is then pure noise. The internet wad already a noisy place but humans are “rate limited” to a degree an LLM is not.
> A recipe is a statement of the ingredients and procedure required for making a dish of food. A mere listing of ingredients or contents, or a simple set of directions, is uncopyrightable. As a result, the Office cannot register recipes consisting of a set of ingredients and a process for preparing a dish. In contrast, a recipe that creatively explains or depicts how or why to perform a particular activity may be copyrightable. A registration for a recipe may cover the written description or explanation of a process that appears in the work, as well as any photographs or illustrations that are owned by the applicant. However, the registration will not cover the list of ingredients that appear in each recipe, the underlying process for making the dish, or the resulting dish itself. The registration will also not cover the activities described in the work that are procedures, processes, or methods of operation, which are not subject to copyright protection.
Recipes were an easy way to avoid some copyright claims. Copy the list of ingredients, and write a paragraph about how your grandmother made it from a secret recipe that turned out to be on the back of the box.
----
I can still think of content farms and the 2010s and the sheer bulk of junk they produced.
And in trying to find some other examples, I found https://web.archive.org/web/20170330040710/http://mediashift...
> The former “content creator” — that’s what Demand CEO Richard Rosenblatt calls his freelance contributors — asked to be identified only as a working journalist for fear of “embarrassing” her current employer with her content farm-hand past. She began working for Demand in 2008, a year after graduating with honors from a prestigious journalism program. It was simply a way for her to make some easy money. In addition to working as a barista and freelance journalist, she wrote two or three posts a week for Demand on “anything that I could remotely punch out quickly.”
> The articles she wrote — all of which were selected from an algorithmically generated list — included How to Wear a Sweater Vest” and How to Massage a Dog That Is Emotionally Stressed,” even though she would never willingly don a sweater vest and has never owned a dog.
> “Never trust anything you read on eHow.com,” she said, referring to one of Demand Media’s high-traffic websites, on which most of her clips appeared.
What It's Like To Write For Demand Media: Low Pay But Lots of Freedom (2009) https://news.ycombinator.com/item?id=1008150
The extra fluff relates to copyright by making wholesale copying of articles illegal. It's not about making the recipe copying legal.
The SEO stuff is true too.
The difference is that back then there was an effort from companies like Google to fight the spam and low quality content. Everyone was waiting Matt Cutts( back then head of web spam and search quality at Google) to drop a new update so they can figure out how to step up their game. So at one point you could't afford to just spam your domain with low quality content because you would be penalised, and dropped from the search engines.
There is nothing like that today everybody is on the bandwagon of AI, somehow chatting with pdf documents is now considered by the tech bro hype circle as a sign of enlightenment a beginning of a spark of intelligence...
Isn't an LLM just a form of compressing and retrieving vast amounts of information? Is there anything more to it than that?
Don't think LLM itself will ever be able to out compete competent human + LLM. What you will see is that most humans are bad at writing books so they will use LLM and you will get mediocre books. Then there will expert humans that use LLM and are experts to create really good books. Pretty much what we see now. Difference is future you will a lot more mediocre everything. Even worse than it is now. I.e, if you look at Netflix there movies all mediocre. Good movies are the 1% that get released. With AI we'll just have 10 Netflix.
Both will happen, with dire effects to the internet as a whole.
Perhaps, perhaps not. The best performing chess AI, are not improved by having a human team up with them. The best performing Go AI, not yet.
LLMs are the new hotness in a fast-moving field, and LLMs may well get replaced next year by something that can't reasonably be described with those initials. But if they don't, then how far can the current Transformer style stuff go? They're already on-par with university students in many subjects just by themselves, which is something I have to keep repeating because I've still not properly internalised it. I don't know their upper limits, and I don't think anyone really does.
Not a specific LLM's limits, the limits of LLMs as an architecture.
Your brain is also based on statistics. We also get stuck in a rut because the "right" answer is no longer statistically significant.
And yet this is not what limits our cognition.
Current LLMs are slow to update with new info, which is why they have cut-off dates so far in the past. Can that be improved to learn as fast (from as little data) as we do? Where's the optimal point on inferring from decreasing data before they show the same cognitive biases we do?
(Should they be improved, or would doing that simply bring in the same race dynamics as SEO?)
The problem with LLMs is that there is one and it is always the same. Sure, you can get different ones and train your own, to a degree.
"LLM" isn't an architecture. The transformer architecture used by all the leading LLMs is Turing complete.
That’s because it’s not really true. There are glimpses of this but it trips up too often.
The web will be different, and I don't count SEO out yet, but... maybe we'll like AI as a middleman better than what's on the web now.
I’ve seen this take before and I genuinely don’t understand it. Plenty of people create content online for the simple reason they enjoy doing it.
They don’t do it for the traffic. They don’t do it for the money. Why should they stop now? Is not like AI is taking away anything from them.
I personally give this experiment a 7, maybe 8 out of 10.
It’s gonna be a mess I can tell you already but it’s not going to be impossible.
There’s plenty of people who love writing and won’t stop.
I’ve given up on the internet as a place to share my passions and hobbies for the most part, and while LLM’s weren’t the only reason, this current trend is a significant factor. I focus most of my attention on talking directly with people. And yes that does mean the information I share is guaranteed to be lost to time, but I’d rather it be shared in a meaningful manner in the moment than live on in an interpreted zombie form in perpetuity.
I think giving up on the web because of AI is the wrong move. You should still create and focus more on connecting with others directly, when online. Get in touch, write emails, sign guestbooks.
I’m personally having great exchanges daily with people from all over via email and that won’t stop because of stupid ChatGPT or whatever.
And don’t get me wrong, it’s awesome to spend more time offline so if you want to do down that path it’s great.
I just don’t think it’s the only solution.
So why harm your audience and your own baseline preferences just to spite a system that will never notice the attack?
Also there are plenty of people who create content because they love it, and also need to be able to make a living at it, because doing so at the level of quality they want is time consuming and expensive.
But mostly because even people who produce content because they love it want to share that content with the world and that will be nigh impossible when the only content anyone sees, and that any platform or algorithm surfaces, is AI generated. Why put in the effort and heart and work to create something only for an AI to immediately clone it for ad revenue? Why even bother?
And in doing that you also prevent real humans from accessing that same content. Look, I have no simpathy for AI companies. I wrote about it before on my site, will probably write again. The current situation sucks. But giving up is not the right answer imo.
> Also there are plenty of people who create content because they love it, and also need to be able to make a living at it, because doing so at the level of quality they want is time consuming and expensive.
Fair but those are the minority. I'd argue the vast majority of people create content because they enjoy the process and earn a living in other ways. I run a newsletter where I interview people with blogs and so far, after a year running it, not a single person has told me they blog for a living. Every single one is doing it for passion. And I suspect that's true for the vast majority of people out there. The bulk of internet content (when it comes to creative content that is) is created by people who do it as a hobby.
> But mostly because even people who produce content because they love it want to share that content with the world and that will be nigh impossible when the only content anyone sees, and that any platform or algorithm surfaces, is AI generated. Why put in the effort and heart and work to create something only for an AI to immediately clone it for ad revenue? Why even bother?
Why even bother? Because there are people out there who care. And the assumption that "the only content anyone sees, and that any platform or algorithm surfaces, is AI generated" is a wrong one imo. I can assure you that there are PLENTY of people out there who still value original content, still value connecting with real human beings doing things because they love the craft. Assuming everything is doomed is not helpful.
Is it going to be harder? Yes. Are there solution? Yes.
Most likely we will see an arms race where some companies try to filter out AI content while others try to imitate humans as best they could.
Putting aside the question of whether dragnet web scraping for human generated content is necessary to train next gen models, OpenAI has a massive source of human writing through their ChatGPT apps.
We are now much closer to that world than we ever were before.
And there’s obviously other people who do this too https://github.com/kagisearch/smallweb/blob/main/smallweb.tx...
I don’t get much traffic but I don’t mind. The thing that really made it for me is sites like this http://www.math.sci.hiroshima-u.ac.jp/m-mat/AKECHI/index.htm...
They just give you such an insight into another human being in this raw fashion you don’t get through a persona built website.
My own blog is very similar. Haphazard and unprofessional and perhaps one day slurped into an LLM or successor (I have no problem with this).
Perhaps one day some other guy will read my blog like I read Makoto Matsumoto’s. If they feel that connection across time then that will suffice! And if they don’t, then the pleasure of writing will do.
And if that works for me, it’ll work for other people too. Previously finding them was hard because there was no one on the Internet. Now it’s hard because everyone’s on it. But it’s still a search problem.
Do you think a collection of them can spot one another's hallucinations?
Do you think that, on occasion, some hallucinations will at least directionally be under explored good ideas?
This is true only if you assume that combining existing thought patterns is not new thinking. If they can't learn a certain pattern from training data, indeed they would be stuck. However, their training data keeps growing and updating, allowing each updated version to learn more patterns.
That is a naive, flawed way to do it. You need to filter and verify synthetic examples. How? First you empower the LLM, then you judge it. Human in the loop (LLM chat rooms), more tokens (CoT), tool usage (code, search, RAG), other models acting as judges and filters.
This problem is similar to scientific publication. Many papers get published, but they need to pass peer review, and lots of them get rejected. Just because someone wrote it into a paper doesn't automatically make it right. Sometimes we have to wait a year to see if adoption supports the initial claims. For medical applications testing is even harder. For startups it's a blood bath in the first few years.
There are many ways to select the good from the bad. In the case of AI text, validation can be done against the real world, but it's a slow process. It's so much easier to scrape decades worth of already written content than to iterate slowly to validate everything. AlphaZero played millions of self games to find a strategy better than human.
In the end the whole ideation-validation process is a search for trustworthy ideas. In search you interact with the search space and make your way towards the goal. Search validates ideas eventually. AI can search too, as evidenced by many Alpha model (AlphaTensor, AlphaFold, AlphaGeometry...). There was a recent paper about prover-verifier systems trained adversarially like GANs, that might be one possible approach. https://arxiv.org/abs/2407.13692v1
The pre-AI internet will be like scientists looking for pre-nuclear steel.
What makes it impossible for AI to succeed?
Of course, it means a flood of crap content.
Itch.io has almost no crap filters so all you find is crap. Steam lets anyone publish but you rarely come across any crap. Many PC game devs know that the income overwhelmingly comes from Steam vs every other site put together.
Unfortunately, this just gives more power to the walled gardens.
I run a content site that is fully AI generated, https://tryspellbound.com
It writes content that's worth reading, but it's extremely expensive to run. It requires chain of thought, a RAG pipeline, self-revision and more.
I spent most of yesterday testing it and pushed it to beta, but the writing feels stilted and clearly LLM generated. The inflection point will come for content people actually want to read, but it's not going to be GPT-4o mini.
How does spellbound work?
It's an entirely different game once you can generate useful content worth reading with AI. People will even pay you good money for it.
It's all garbage so no one notices when some of it suddenly disappears off the face of the earth and gets replaced with other garbage: but for the ones making it, their revenue essentially goes to 0 overnight.
Making good revenue but mostly I'm just very bullish based on current revenue and willing to bootstrap
so, status quo? this sort of content only has value because google links to it when people search, and because google runs an ad network that allows monetizing it. google is also working furiously to provide these same AI-generated answers in their SERP, so they can eliminate this and monetize the answers directly instead of paying out to random third parties.
i'm pretty skeptical that this ai-generated content will ever be monetizable in the way the article suggests, simply because google is better at it. if you're a human making your living by writing articles that are indistinguishable from ai-generated content, then you might be harmed by this but for most people this inflection point is not going to be a noticeable change.
> I'm going to take the median across all categories, which is an estimated annual revenue of $1,550 for 50,000 monthly page views.
> This is approximately ~$0.00022 earned per page view.
The problem is... this doesn't take into account a million AI generated sites suddenly all competing for the same amount of eyes as before, driving revenue to zero very quickly. It'll be worth something for a bit and then everyone will catch up.
It seems to assume a world where SEO entrepreneurs where ready to churn out million-page sites, but the cost per query were blocking them. There is no marginal cost, no SEO cost to adding another page, as long as a couple people visit it and "pay it off".
In the real world, it doesn't work like that. Whatever monstrosity was created like this would not do well in the search engines. So no meaningful threshold has been passed, in terms of the cost for AI generation.
People are creating lots of AI content, but not like this - not bottom tier generic SEO pages which will barely rank and aren't that compelling in an already saturated Internet.
Incidentally the real money seems to be in generating AI images and, eventually, video: much better return for your money.
HN is not a reflection of the world.
I'm optimistic about synthetic data giving us another big unlock, anyway. The text on the internet is not that reasoning dense. And they have a snapshot of pre-2023 that is fixed and guaranteed not to decay. I don't think one extra year of good quality internet is what will make or break AGI efforts.
The harder bottleneck will be energy. It's relatively doable to go from 1GW to 10GW but the next jump to 100GW becomes insanely difficult.
> an estimated annual revenue of $1,550 for 50,000 monthly page views.
> This is approximately ~$0.00022 earned per page view.
No, this is $0.002583 earned per page view, a ~12x difference. Looks like the author divided by 12 twice.
I still enjoy commenting on HN and writing some thoughts on my blog. I'm pretty sure that there are many other people too.
At some point everything that is not cryptographically singed by someone I know and trust needs to be considered AI generated.
Maybe AI-generated content might have better quality than generated by humans. But then it's likely that I'm under the influence of some bigger corporation that just needs some eyeballs.
So, on the margin, this will drive human created content out since it is now less profitable to do it by hand than it was before.
That means this will devalue AI content and ad impressions, but not necessarily human content, if people in fact value it.
Lot cost models just lowered the bar of entry.
For example, a person could write a shareware game over a few weeks or months, sell it for $10, buy advertising at a $0.25 customer acquisition cost (CAC) and scale to make a healthy income in 1994. A person could drop ship commodities like music CDs and scale through advertising with a CAC of perhaps $2.50 and still make enough to survive in 2004. A person could sell airtime and make speaking appearances as an influencer with a CAC of $25 and have a good chance of affording an apartment in 2014. A person can network and be part of inside deals and make a million dollars yearly by being already wealthy in a major metropolitan city with a CAC of $250 in 2024.
The trend is that work gets harder and harder for the same pay, while scalable returns go mainly to people who already have money. AI will just hasten the endgame of late stage capitalism.
Note that not all economic systems work this way. Isn't it odd how tech that should be simplifying our lives and decreasing the cost of living is just devaluing our labor to make things like rent more expensive?
As a simple example, imagine a Prisoner's Dilemma, except neither side knows defecting is an option (so in effect both players are playing a single-move game where "cooperate" is the only option). Landing on cooperate-cooperate in this case is easy (indeed, it's the only possible outcome). But as soon as you reveal the ability to defect to both players, the defect-defect equilibrium becomes available.
We will need humane solutions to this, because the non humane ones are starting to become visible (armed drone swarms driven by AI).
I 'd rather be exploited by google
Arc is an excellent example because AFAIK it's still free, and I haven't heard a single complaint about throttling, availability, etc., and they've since gone on to treat it as a marketing tentpole instead of experiment.
https://platform.openai.com/docs/guides/rate-limits?context=...
This has really dystopian vibes, since it centralizes opinion and “factuality” in an authoritative but potentially extremely biased or even manipulatively deceptive manner.
OTOH it will provide opportunities for competitive solutions to query answering.
News sites are already often shit and parasitic. I mean parasitic because if you go to a free news site (say Yahoo news, etc) you often see rewritten articles that originated from paid sites (e.g. NYT). The pure ad-supported sites are typical enshitification that degrades journalism and increases sensationalism because they don't need to write unique articles, but you should sensationalize them to drive up views. You also don't have to hire journalists to get story details. So news most people read degrades and you get very limited views.
The problem here is that this paradigm barely works because you have to pay real people to write those rephrased articles. So while it costs more to run the NYT where you need to hire investigative journalists and send people to physical places, there is a bound on that difference. But if you paste in a NYT article into GPT4 and ask it to summarize it, you'll get very similar quality to yahoo news (or even CNN, MSNBC, or Fox. Which all also do this leeching, but less of an issue). I'm sure people realize how easy it is to scrape NYT and then post the GPT output. This is in spirit no different than if you just used archie.is, but large scale.
The same is true for many tutorial sites or cooking sites, etc. I'm sure many of you also get annoyed at the google search results that are just stackover flow posts embedded on a different site or the Medium articles (especially paid ones) that are also just SO posts and can show up higher in the listing.
The issue becomes: how do we generate and disseminate new information in this paradigm? Okay, free blog posts aren't "hurt" because they have no income, but people build reputation through them and it gets many people jobs. But what about others that do make a living through this? Is this not similar Jack Conte's (Patreon co-founder/CEO and 1/2 of the band Pomplamoose) argument about creating content "for the algorithm" vs for "yourself/your fans/fun/etc". That it is taking some of the human elements out of the art/entertainment/content. (Can totally disagree with his argument btw). Personally I'm on the side of Jack. Our goal shouldn't (now) be to just serve people search results or just generate content for content's sake, but to now focus on serving people high quality content and high quality results. Google indexed the entire internet. People gamed the system (SEO) and now google results are shit, youtube results are shit, and everything is shit. We don't need more content (who uses page 2 on Google?), but we need to have better content. [1]
I think we need to ask: is this what we want? If not, then what are we going to do about it?
If we are okay, then I think someone should create a super-website where you just have information about just about everything. There definitely is utility in it. But the question is at what cost.
[0] https://youtu.be/hwn6-8XpIuE
[1] I think most people want this. But the problem is you're not going to find market forces showing this because there is no product doing this. Or if there are, they aren't well known and could be confusing to use and/or a wide variety of problems (UI/UX do matter). But it requires reading between the lines and market research a la talking to people and finding out what they want, not a la data. You need both.
We're trying to do that with PulsePost (https://pulsepost.io) and the biggest challenge is unique content. Given a keyword or a niche topic, AI models tend to generate similar content within similar subjects. Changing the temperature helps to a degree but the biggest difference comes from adding internet access. Even with same prompt, if the model can access the internet, it can find unique ideas within the same topic and with human review it becomes a high value article.