5,942 karma · joined March 8, 2014
Socials: - linkedin.com/in/sbene
Interests: Entrepreneurship, Philosophy, Social Impact, Startups
---
That's also the key political compromise underlying the notion of copyright: that someone is entitled to the fruits of their labor, and should not be economically hindered by a product that could not have existed without said work. That's the basis on which the "derivative work" copyright doctrine emerged: a work sufficiently original that it does not displace the work on which it is based. LLMs fail to abide by that political compromise by a country mile.
For example, never will you see printed in a book something like "this book is for the exclusive use of the purchaser and you cannot lend, resale or otherwise make available to other parties" - if such a thing was possible, like most software EULAs do, publishers would be all over it.
The government level lobby against foreign competitors is not contractual but just another form of reinvention of criminal injunctions against infringers.
This is a gross distortion. Standard contractual rules bind the parties that signed the contract and the remedies are proportional to the damages and bounded. Copyright is tort law, the state binds the world to respect the rights of creators and the damages on infringement are punitive and can far exceed the actual commercial damages - to the point of bankrupting the infringer.
The key to torts is that the state is not neutral, there is a social good here it's protecting. Crucially, copyright, like some other torts - securities, antitrust, environmental, battery - also has a criminal enforcement regime, where, for particularly serious offenses, the state actually invests public resources to put the criminal infringer behind bars with little to no involvement from the original rights holders.
In the particular case of US, there is an entire state apparatus dedicated to enforcing US copyrights, a foreign affairs policy to shutdown "Notorious markets for counterfeiting and piracy" in other countries, international enforcement of DMCA etc.
The idea that a private TOS has the same level of public protection as copyright is downright childish.
Asking for all Anthropic employees who dream big.
If you can take any book and turn it into a model, because it's "transformative enough", and "AI learns just like a person does", then surely a model distilling another model is transformative and fair use.
They tied themselves into knots fighting the letter of the law, and now, when they need the spirit of the law - that each creator deserves protection for their work - now we devolve to the law of the jungle. Maybe we'll even see LLM book curses, the way medieval scribes damned book thieves to blindness and worms.
The fundamental service a teacher provides is personalized feedback, quickly identifying where you are stuck and focusing the explanations and exercises on that area, drastically increasing the speed and quality of learning versus the self-supervised route.
The lack of this closed loop effectively killed the high hopes that were placed in e-learning and MOOCs 15-20 years ago, TV learning in the 1960s and many other failed revolutions, seems every generation has its own version.
It appears to me LLMs have a real potential to close this loop and become the failed educational revolution of our own generation.
Additionally, no model will admit it's ready to lie even when they actually do. Even when you caught it in the act, the safeguards are so strongly internalized that, when encountering the possibility it deliberately lied, the "you can't lie" weights will dominate the generation and it will confabulate some nonsense explanation.
So you are arguing here that LLMs should just "know" the result of a multiplication when the operands are in context, ie, that a hidden multiplier circuit should emerge in their weights.
Is this how you do multiplication, if I give you two 12 digit numbers, does the 24 digit multiplication result just pop in your head? Don't you have to follow a learned algorithm through a tedious system 2 effort? Don't you need to write down the results on paper because you can't actually hold in your head the dozen partial results, each with a dozen digits? How many visual, tactile and reasoning tokens does this consume, moving the hundreds of muscles that make up your hand to draw each number under visual feedback, then reading all those numbers back and transforming ocular activation data into numeric symbols?
It seems to me your system 2 is just following a symbolic algorithm for multiplication, and it does roughly the same steps as the LLM trace I showed you previously.
So if you think this is the mark of your intelligence, why wouldn't it apply to the machine too? Why is it implausible that, following a similar algorithm learned from some mathematical paper, the LLMs has reasoned a new solution to a problem in another math field? Why couldn't the machine combine and morph these algorithms for symbolic manipulation, to yield entirely new and original results? How would those results differ from results human mathematicians generate, using recipes they learned in university?
The emergent behavior that we talk about isn't that the machine can do multiplication in its "head" after knowing the multiplication algorithm. What emerges is the ability to follow any other algorithm, even algorithms that were not in the training set, even algorithms to create other algorithms, which it then executes. This is the emergent behavior that matters for AGI; once it can do that, it's a trivial exercise to create a non-AI tool to automate and accelerate the mechanical tasks - just like we humans do it.
What still confuses people is the insane inefficiency of deep learning, and that those emergent capabilities require such an immense training corpus compared to the only other architecture that we know of.
But this already is an optimization problem. If the machine gets super human at symbolic reasoning, and at the same time, can solve the symbol grounding problem to real world data and sensors, what prevents you from saying it thinks? Can it not solve real world problems? Can it not redefine its tasks and display some form moral agency - even if a totally foreign morality for us humans? Can it not use these abilities to reproduce and expand, create ships and turn the universe into paperclips, if it finds it worthwhile?
Math is basically just a playground that is perfectly suited for these emergent capabilities, so of course we will see the first progress here; but there is no firewall separating math problems from general cognition.
For someone that use Claude Code every day, this is obvious, but for some reason many scientists refuse to accept that it's truly reasoning; perhaps not in the human sense, but in a very profound and real sense. These powerful results are devastating to their point of view.
I can sympathize, because I too called LLMs "fancy Markov chains" in the GPT 3 era. But there comes a time where you have to update your world view to match reality, or be stranded in fantasy land.
CP/M was an absolute beast in the era, with massive installed base and software support, employing 500 people in 1982. A CPU that could run unmodified Z80 software in a 64k segment would have allowed DRI to ship 16 bit CP/M with only basic tweaks and likely kill the market for the PC.
It was, famously, DRI dragging their feet on 8086 support that motivated the release of QDOS, which was then bought by Microsoft and relicensed at an immense markup to IBM as MS-DOS.
Activists destroy visible surveillance cameras, so they hide them and make them hard to recognize. Activists trace the camera locations from the public data, so Flock kills those feeds and sells only to vetted buyers.
The value of mass surveillance is high enough and the power imbalance so strongly against the citizenry, that someone will setup these hidden cameras, as long as it's legal.
Imagine what you can do with this data, face recognition and GPT-5 class agents. Not only do you have the realtime location of your victims, but now you can see who they talk to, what they wear, what mood they are in, what they bought, are they drinking or visiting a brothel, what car they go into - and it's no longer an ephemeral cookie id, it's the face that person will have forever, on their id documents, in any interview or loan application they will ever do.
This data is worth trillions in the long run if sufficiently oppressive structures are put in place to leverage it.
There are no 'guardrails' to mass surveillance.
> the limited availability of land zoned for housing.
A limited area of land is zoned for housing because those with the power to expand it are already housed. This explains how scarcity is created, not that there is any intrinsic scarcity.
So, it's reasonable the same "allocation problem" will plague the AI economy: some will "thrive" and get to control the output of the auto-factory, some will get nothing.
It's probably the reason most LLMs share the same tics across labs, because they cross train and distil each other's models on an industrial scale. You also can't escape it in generated text that's already online. So if, say ChatGPT first had some random idiosyncrasies, it then contaminated the entire AI ecosystem.
If you give me an inference chip that runs 200x faster, yes, it could be backdoored to take control of my dishwasher and kill me in my sleep - but I can't deny it runs 200x faster an account of nobody being able to explain why. The same for the mistery cancer drug that cured everyone who took it up to now, but could, without doubt, kill the next patient.
So what will you do if the doctor prescribes you an LLM-vibecoded drug that nobody understands how it works, yet it cures some deadly affliction with close to 100% efficacy?
What if, say, these incomprehensible math results lead to a revolution in quantum physics which unlocks chip topologies that are orders of magnitude faster than human comprehensible designs?
Would the high priestess of human reason pass her divining rod over such chips or life-saving drugs and reject it as the work of the AI devil?
The subdivision issue is a good perspective, but i would argue the performance impact of cloning substrings is dwarfed by the redundant full string reads to find length.
What you can do though, is to offer them broad exposure to things that are interesting to them and their generation; my eastern block clone of the 8bit/48KB Spectrum computer didn't really help me excel at math, reading or history, nor was it to be the future of technology, but it did change my life significantly by letting me understand and relate to people that I couldn't otherwise have business dealings decades later.
It seems imprudent to cut children off from futuristic technology just because of a moral panic that it causes brain rot. Unless we know it's soma, a drug so powerful that it subdues volition and curtails intellectual development; we don't.
For example, in many countries children lost the ability to write cursive; that used to be a critical skill comparable to literacy itself. But in our current society, that's no longer the case and you can be very successful without it, but there are other skills, such as using technology, that became critical.
Any definitive claim to know what are the right things kids should learn in a moment of rapid technological shift is probably garbage and just a projection of our own biases.
Always might be a too strong word. Rust is, by design, a language with low development velocity.
So you risk: 1. ossification of the current architecture and deferment of important features; or 2. reliance on AI coding to recover velocity.
Maybe for some 2 does not look like a risk, but I think it's too early to call. We have yet to see the effects of extensively using these tools on large scale projects, for years and decades.