78 karma · joined August 3, 2025
There might be an inevitability to this - there is always a form PC policing, or euphemism, or re-brands, or rhetoric. But there is always the danger that what is considered a well-meaning tweak to language masks or evolves into repression.
From TFA "Text watermarking is meant to help identify whether content was generated by AI, but inside an agent it also becomes part of the generation process that produces decisions."
The struggle in repression is real, even if it's unconscious. The psychoanalyst asks to speak with FREE ASSOCIATION and the inability to do this in SPEECH is what reveals unconscious problems. Repression, though, is a STRATEGY that we most definitely sign-on to. It FEELS unjust to be asked to take responsibility for something we didn't INTEND to do, but that, unfortunately, is our task in life.
Political Correctness has a similar idea that by adjusting the terminology we use, we can purge biases and historical implications and speak in a purer way.
Psychoanalysis has its own idea of repression - where a person struggles to BLOCK our associations between ideas, memories, and words in order to try to stop one thought from being contaminated by another, intolerable thought.
All of these attempts to control language are fundamentally misguided at best, often have severe unintended consequences, and are genuinely immoral at worse.
I remember in the early days of PGP this was put forward as a vision. I remember the various key-holding systems (that were hard to use). I did a key-party with my brother and his friend in the '90s.
Why did this fail?
This is not always the psychological situation, but I think it might be a lot of the time.
I also think that until we recognize this, we won't be able to help people to not do it. Calling them lazy or being bewildered by them or feeling rage or contempt towards them or threatening them is not going to be practically helpful.
I'm not totally sure I understand the "distribution/registries" thing though? This also related to GitHub and HF, right? And Netflix. (You mentioned Spotify.) How they host stuff and then leverage that position to build add-ons and lock-in, lobby for laws. Companies PAY for this "bandwidth."
I always thought that BitTorrent would eventually take off and make these kinds of sites irrelevant, that it would democratize bandwidth, the last leg of the battle for universal accessibility.
I like your idea. I wish you luck.
Or: the CCTV is well hidden, but the people who installed it left behind some mud that got on someone's shoes and that made them late...
THIS is where the regulation needs to start.
But if the point is to be "censorship-free" then why respect licenses at all? They are among main choke points today. If authoritarians use licenses to censor political, artistic, scientific, etc., speech that they want to block, does that make the censorship more respectable?
When Anthropic sues a Chinese lab for IP infringement and get a court to put a bar on that software, does it THEN get pulled from Pirate Face?
I know that an awful lot of international negotiations have become focused more and more on questions of "IP" - licensing battles are already intensely politicized and it's hard to imagine a future where it doesn't get much much worse. Imagine N Korea coming after you for violating a license that they worked hard to control and leverage.
I haven't wrapped my mind around this
Culture has always had a kind of violence to it, but not FROM the THIEF: you HAVE to speak the language that was already made; you have to understand and participate in their song-structures, their plot devices, their ways of attesting.
The "IP owners" should be paying ME!!!
We are trying to make images in the AI era be as reliable as they were pre-AI? But they were not reliable pre-AI, it was just more difficult to intentionally doctor them.
A lot of progress is in harnesses and distillation and weird research that requires creativity. The "Frontier" always tries to frame things in terms of "the frontier" in order to limit how we look at things.
We don't know the true costs of pre- and post- training, running data-centers, hiring people, making business deals, etc., so even the (maybe) quantifiable question of cost vs capabilities is not at all clear. The efficiency of businesses ultimately gets measured by their ability to maximize a commercial return; but this is NOT a good metric for other important things, including maybe "how good is this"? (Where "good" is open to debate but maybe not quantifiable.)
It had been true that the "best" operating systems required a lot of business money. And encyclopedias. And ... (list some other good historical examples here!) ... Are Linux and Wikipedia the "efficient frontier"? In a sense they are both infinite performance at zero cost. Hard cash outlays sometimes stop being the determinant for high-effort endeavors.
I remember in the 1990s Encyclopedia Britannica's CDs and MSN (Microsoft Network(tm)) among many others argued that the web would never be able to match the curated, bespoke, expensive products they offered. It was IMPOSSIBLE for that crap to be matched!
It is a POLITICAL question.
I remember so well when "Open Source" branding started with Bruce Perens - all the business arguments. When you need a database, someone ELSE'S business decisions (how THEY are going to make money) never end up helping YOU. If you use proprietary extensions and become dependent on them, you inevitably will get BURNED when their business needs diverge from your needs. - They close shop - They refuse to interop with something you need - They demand that you obey their arcane rules - They rug-pull - They get hacked as only they can - They lie to you - They stab you in the back
THANK YOU, Nandakishor Mukkunnoth, for putting in the work to help to clarify this stuff!
You are like a firefighter compared to their fire-insurance racket.
Can I add NOTES about pages? This might be a good spot to do that...? Maybe the interface can be in a web page instead of terminal?
Before Google took off there was a vibrant ecosystem of FOSS dev around search, all different little aspects of it. Then after Google people stopped fiddling with search, search became "solved" or maybe "must be coded by the big boys". Shame.
Thank you for this, looooong time coming
2) I do not believe the "billions" number is OpenAI/Anthropic's training costs. I suspect it includes business expenses (including the big $$ to the guy who came up with "frontier model") and infrastructure, etc. That includes the data-center costs for running the cloud. And the "R&D" expenses which includes who-knows-what. And the settlement payment for data access. Etc. Why are the cost breakdowns not available to the public? Not because of thoughtfulness, altruism, care for the human race, but because of business plans.
3) "Piggybacking off OpenAI/Anthropic" - Scraping the web is "piggybacking" too, and so is buying existing data or even paying for new data. The "L" in LLM stands for "language" which is our common heritage.
But what does this have to do with anything anyway? People argue that truly Free (FOSS) LLMs couldn't be be developed because of costs, but I do not think that that is obvious. This used to be the argument against Linux and Wikipedia.
The AI gods have tried to make it seem like they are not just running their programs in the cloud, but that the AI is so "powerful" that it requires a super computer to run.
But my 5 year old Mac is running pretty good AI, shocking fast, surprisingly powerful! The "natural" course of dev (Moore's law, etc) makes it pretty clear where this is going...
Remember how DeepSeek v.whatever cost ~$5m
Stable Diffusion 1.5 was reportedly $70k in compute.
https://hugovergnes.github.io/little-lm-3-8b/
The training and cloud costs are often conflated. R&D costs too (which are hard to compare to open systems).
A lot of compute is clearly "wasted" where they are not focusing on optimizations, etc.
We can't trust any of their denials. FB denied everything year after year.
AI regulation needs to start here. Forced interop, forced source code licensing, harsh penalties for privacy violations or conspiracy to access private data. Block lobbying. Etc. These are the kinds of old-fashioned solutions we need for this kind of old-fashioned evil!
Model: FB. FB scraped other websites on a massive scale, then spent big on legal lobbying to block others from scraping. FB slurped our address books and spied on our friends. FB bought other companies and mixed the databases. FB made an art & science out of generating "sticky engagement" (they literally acted like trying to addict kids was a worthy "academic" goal, suitable for "serious" investigation thet they consider legitimate "science"). They mastered the cookie and have researched web fingerprinting techniques running 24/7/365.25. Recall that FB recently backdoor-installed a webserver onto every iPhone they could in order to circumvent tracker-blocking.
We aren't just disclosing by chatting. The AI companies now run binaries on all of our computers. They are 1000% non-transparent about everything. They make up new econ-jargon (like "run-rate") to make it seem like they are disclosing. They are constantly doing complex international lobbying and mucking in international relations. They have powerful propaganda/spin centers generating stories, ,manipulative warnings, and misleading info.
This is NOT a comment on AI tech. I like AI, and I support the right of people (programmers) to scrape the open web.
But in short: these are good, old-fashioned tech companies that we have seen over and over ... and over. They are positioned to be the next M$, the next FB (IBM, AOL, lol). Did you follow the latest Steve Balmer news? Do you read Pro Publica?
I get on my knees and PRAY...