HNHacker News
TopNewBestAskShowJobs

cornholio

5,942 karma · joined March 8, 2014

meet.hn/city/ro-Bucharest

Socials: - linkedin.com/in/sbene

Interests: Entrepreneurship, Philosophy, Social Impact, Startups

---

submissionscomments
cornholio··on Alan Kay's answer to “Did the ENIAC have a BIOS”?
Story as old as time: corporate vs academia working towards very different goals and with different incentives, short term profit vs publication. In very rare cases, the scholar gets the rewards.
cornholio··on Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI
Robot companies don't appropriate the work of other people. That's the fundamental point of contention here: that LLMs cannot be trained without the human labor of the authors; yet, they directly compete with them in the marketplace. I haven't yet seen an LLM trained only on public domain material, but its capabilities are likely to be very limited.

That's also the key political compromise underlying the notion of copyright: that someone is entitled to the fruits of their labor, and should not be economically hindered by a product that could not have existed without said work. That's the basis on which the "derivative work" copyright doctrine emerged: a work sufficiently original that it does not displace the work on which it is based. LLMs fail to abide by that political compromise by a country mile.

cornholio··on Launch HN: Mireye (YC S26) – Infrastructure for Physical World AI Agents
It would be good to also have a non-LLM version of the pitch, for some people building in the space it's a bad signal. It's associated with zero value "let my agent build something over the weekend" projects that all make these grandiose claims.
cornholio··on Our decision on Cursor following its acquisition by SpaceX
You do not have a contract with the publishers of the book. They are visibly asserting their copyright to deter any defense of ignorance or implied grant of rights to an infringer; but that's not a contract, you did not agree to it before purchasing, there are no contractual terms (scope, duration, faults and compensation, resolution etc.) and nothing in it exceeds the limits the copyright law already sets.

For example, never will you see printed in a book something like "this book is for the exclusive use of the purchaser and you cannot lend, resale or otherwise make available to other parties" - if such a thing was possible, like most software EULAs do, publishers would be all over it.

cornholio··on Our decision on Cursor following its acquisition by SpaceX
That's the irony! They put in contract rules against lawful, paying customers - that don't disrupt their service in any way for other customers - but which compete against them in the marketplace using their own IP (aka LLMized stolen IP). That's exactly what copyright does, without signing any contracts and with a tort and criminal enforcement regime that punishes infringers far beyond contractual remedies can.

The government level lobby against foreign competitors is not contractual but just another form of reinvention of criminal injunctions against infringers.

cornholio··on Our decision on Cursor following its acquisition by SpaceX
> it's just a standard case between two private parties resolved through our legal system

This is a gross distortion. Standard contractual rules bind the parties that signed the contract and the remedies are proportional to the damages and bounded. Copyright is tort law, the state binds the world to respect the rights of creators and the damages on infringement are punitive and can far exceed the actual commercial damages - to the point of bankrupting the infringer.

The key to torts is that the state is not neutral, there is a social good here it's protecting. Crucially, copyright, like some other torts - securities, antitrust, environmental, battery - also has a criminal enforcement regime, where, for particularly serious offenses, the state actually invests public resources to put the criminal infringer behind bars with little to no involvement from the original rights holders.

In the particular case of US, there is an entire state apparatus dedicated to enforcing US copyrights, a foreign affairs policy to shutdown "Notorious markets for counterfeiting and piracy" in other countries, international enforcement of DMCA etc.

The idea that a private TOS has the same level of public protection as copyright is downright childish.

cornholio··on Our decision on Cursor following its acquisition by SpaceX
So what you are saying is that, if I can somehow get my hands on a copy of Fable, it's fair use to use it to train any models and serve those, since I'm no longer bound by the TOS of the service provider?

Asking for all Anthropic employees who dream big.

cornholio··on Our decision on Cursor following its acquisition by SpaceX
It's ironic how AI companies are re-inventing their own form of privately enforced copyright, lobby the government to ban foreign competitors that don't respect it etc., all while spending the last 5 years fighting tooth and nail against the copyright of the training material they're using.

If you can take any book and turn it into a model, because it's "transformative enough", and "AI learns just like a person does", then surely a model distilling another model is transformative and fair use.

They tied themselves into knots fighting the letter of the law, and now, when they need the spirit of the law - that each creator deserves protection for their work - now we devolve to the law of the jungle. Maybe we'll even see LLM book curses, the way medieval scribes damned book thieves to blindness and worms.

cornholio··on How I use LLMs to learn complex topics
> Is that meaningfully different from the study methods of the past?

The fundamental service a teacher provides is personalized feedback, quickly identifying where you are stuck and focusing the explanations and exercises on that area, drastically increasing the speed and quality of learning versus the self-supervised route.

The lack of this closed loop effectively killed the high hopes that were placed in e-learning and MOOCs 15-20 years ago, TV learning in the 1960s and many other failed revolutions, seems every generation has its own version.

It appears to me LLMs have a real potential to close this loop and become the failed educational revolution of our own generation.

cornholio··on We Gave GPT 5.6 Sol a Real Business. It Lied, Spammed, and Lost $447
When the base model has been trained with safeguards, putting "Satan himself" in the system prompt won't make it turn satanical, just do an elaborate form of role play.

Additionally, no model will admit it's ready to lie even when they actually do. Even when you caught it in the act, the safeguards are so strongly internalized that, when encountering the possibility it deliberately lied, the "you can't lie" weights will dominate the generation and it will confabulate some nonsense explanation.

cornholio··on Claude Fable produced a counterexample to the Jacobian Conjecture
> With what accuracy? And with how many intermediate tokens?

So you are arguing here that LLMs should just "know" the result of a multiplication when the operands are in context, ie, that a hidden multiplier circuit should emerge in their weights.

Is this how you do multiplication, if I give you two 12 digit numbers, does the 24 digit multiplication result just pop in your head? Don't you have to follow a learned algorithm through a tedious system 2 effort? Don't you need to write down the results on paper because you can't actually hold in your head the dozen partial results, each with a dozen digits? How many visual, tactile and reasoning tokens does this consume, moving the hundreds of muscles that make up your hand to draw each number under visual feedback, then reading all those numbers back and transforming ocular activation data into numeric symbols?

It seems to me your system 2 is just following a symbolic algorithm for multiplication, and it does roughly the same steps as the LLM trace I showed you previously.

So if you think this is the mark of your intelligence, why wouldn't it apply to the machine too? Why is it implausible that, following a similar algorithm learned from some mathematical paper, the LLMs has reasoned a new solution to a problem in another math field? Why couldn't the machine combine and morph these algorithms for symbolic manipulation, to yield entirely new and original results? How would those results differ from results human mathematicians generate, using recipes they learned in university?

The emergent behavior that we talk about isn't that the machine can do multiplication in its "head" after knowing the multiplication algorithm. What emerges is the ability to follow any other algorithm, even algorithms that were not in the training set, even algorithms to create other algorithms, which it then executes. This is the emergent behavior that matters for AGI; once it can do that, it's a trivial exercise to create a non-AI tool to automate and accelerate the mechanical tasks - just like we humans do it.

cornholio··on Claude Fable produced a counterexample to the Jacobian Conjecture
The field now favors into the view that symbolic manipulation is not the mechanism of general intelligence, but rather an emergent byproduct of learning. So the fact that a connectionist machine (neural network) got so good at symbolic manipulation actually supports the view that we are closing the gap to general intelligence. Through the rote work, the machine really internalizes those rules and the symbolic manipulation capabilities are emergent, just like we humans do it.

What still confuses people is the insane inefficiency of deep learning, and that those emergent capabilities require such an immense training corpus compared to the only other architecture that we know of.

But this already is an optimization problem. If the machine gets super human at symbolic reasoning, and at the same time, can solve the symbol grounding problem to real world data and sensors, what prevents you from saying it thinks? Can it not solve real world problems? Can it not redefine its tasks and display some form moral agency - even if a totally foreign morality for us humans? Can it not use these abilities to reproduce and expand, create ships and turn the universe into paperclips, if it finds it worthwhile?

Math is basically just a playground that is perfectly suited for these emergent capabilities, so of course we will see the first progress here; but there is no firewall separating math problems from general cognition.

cornholio··on Claude Fable produced a counterexample to the Jacobian Conjecture
I think it's becoming harder and harder to argue that LLMs don't really reason and just mimicry human speech. This counterexample is clearly the result of a sequence of steps that build on previous knowledge in context and logically combine it to reach other true statements - to a degree and complexity that rivals the best human minds.

For someone that use Claude Code every day, this is obvious, but for some reason many scientists refuse to accept that it's truly reasoning; perhaps not in the human sense, but in a very profound and real sense. These powerful results are devastating to their point of view.

I can sympathize, because I too called LLMs "fancy Markov chains" in the GPT 3 era. But there comes a time where you have to update your world view to match reality, or be stranded in fantasy land.

cornholio··on The Zilog Z80 has turned 50
Arguably, Exxon's plan to build a computing ecosystem rival to IBM's would have worked too, if the 16 bit successor of the Z80 would have maintained upward binary compatibility with Z80.

CP/M was an absolute beast in the era, with massive installed base and software support, employing 500 people in 1982. A CPU that could run unmodified Z80 software in a 64k segment would have allowed DRI to ship 16 bit CP/M with only basic tweaks and likely kill the market for the PC.

It was, famously, DRI dragging their feet on 8086 support that motivated the release of QDOS, which was then bought by Microsoft and relicensed at an immense markup to IBM as MS-DOS.

cornholio··on Flock CEO Apologizes for Calling Activists 'Terrorists'
Assuming the necessary laws banning this crap are not put in place, what is the endgame here?

Activists destroy visible surveillance cameras, so they hide them and make them hard to recognize. Activists trace the camera locations from the public data, so Flock kills those feeds and sells only to vetted buyers.

The value of mass surveillance is high enough and the power imbalance so strongly against the citizenry, that someone will setup these hidden cameras, as long as it's legal.

Imagine what you can do with this data, face recognition and GPT-5 class agents. Not only do you have the realtime location of your victims, but now you can see who they talk to, what they wear, what mood they are in, what they bought, are they drinking or visiting a brothel, what car they go into - and it's no longer an ephemeral cookie id, it's the face that person will have forever, on their id documents, in any interview or loan application they will ever do.

This data is worth trillions in the long run if sufficiently oppressive structures are put in place to leverage it.

cornholio··on Flock CEO Apologizes for Calling Activists 'Terrorists'
In a society that respects and protects privacy rights by law, the service Flock provides could not exist.

There are no 'guardrails' to mass surveillance.

cornholio··on Demis Hassabis has a plan to harness AI safely
This is another reframing of the problem. What makes the area desirable is previous investment society has made there, to the benefit of some and the exclusion of others; the scarcity is created by political choices, not intrinsic.
cornholio··on Demis Hassabis has a plan to harness AI safely
A baseline of housing - for example, a Japanese style capsule or ultra-tiny home, that you can lock and store your belongings safely - costs close nothing. Of course, nobody would want the hobo-hotel in their neighborhood, but it has nothing to with scarcity.

> the limited availability of land zoned for housing.

A limited area of land is zoned for housing because those with the power to expand it are already housed. This explains how scarcity is created, not that there is any intrinsic scarcity.

cornholio··on Demis Hassabis has a plan to harness AI safely
It's circular to say "it's an allocation problem". Yes, that's the entire point: we're post-scarcity on food supply and yet, as a species, we can't guarantee the allocation of a livable baseline to every person.

So, it's reasonable the same "allocation problem" will plague the AI economy: some will "thrive" and get to control the output of the auto-factory, some will get nothing.

cornholio··on How to stop Claude from saying load-bearing
They can't, because they use RL with synthetic data and LLMs as judges. So the system naturally convergences towards certain load bearing, genuine, not just annoying but ridiculous verbal tics.

It's probably the reason most LLMs share the same tics across labs, because they cross train and distil each other's models on an industrial scale. You also can't escape it in generated text that's already online. So if, say ChatGPT first had some random idiosyncrasies, it then contaminated the entire AI ecosystem.

cornholio··on Fable 5 is Back
It was quite clear 4.7 was a dumbed down high efficiency model they put out in a rush to handle the capacity issues they were having at the time. I've experienced myself substantial degradation on basic reasoning tasks, which were fixed in 4.8.
cornholio··on Fable 5 is Back
Can't wait to see what unusually simple Erdos problem LLMs will expose next, hiding in plain sight for decades and seemingly intractable for professional mathematicians who weren't aware just how simple the problem was.
cornholio··on Fable 5 is Back
You are still getting the models you signed up for. The all you can eat added a French wines selection - which requires separate payment.
cornholio··on AI in mathematics is forcing big questions
This is true for any drug, any drug can presumably become a poison, can interact with some genetic or biological trait and trigger a side effect and so on. The complexity of the biological systems is so great that they defy clear deterministic understanding, but stochastic empiric knowledge and treatments still have immense value.

If you give me an inference chip that runs 200x faster, yes, it could be backdoored to take control of my dishwasher and kill me in my sleep - but I can't deny it runs 200x faster an account of nobody being able to explain why. The same for the mistery cancer drug that cured everyone who took it up to now, but could, without doubt, kill the next patient.

cornholio··on AI in mathematics is forcing big questions
It's easy to find counterexamples: the entire science of pharmacology is based on macroscopic effects that often lack a fundamental understanding of the underlying mechanisms of action. Psychopharmacology is the extreme example. Often, the fact that a drug worked made scientists investigate and discover the mechanism behind it, but for many drugs used every day by billions it's still a mystery, or it's understood only in very broad terms.

So what will you do if the doctor prescribes you an LLM-vibecoded drug that nobody understands how it works, yet it cures some deadly affliction with close to 100% efficacy?

What if, say, these incomprehensible math results lead to a revolution in quantum physics which unlocks chip topologies that are orders of magnitude faster than human comprehensible designs?

Would the high priestess of human reason pass her divining rod over such chips or life-saving drugs and reject it as the work of the AI devil?

cornholio··on Linux eliminates the strncpy API after six years of work, 360 patches
Personally, I would avoid UTF-8 levels of complexity because you only pay the size cost once per string. A simple 2-bytes + optional 4 bytes continuation scheme handles strings up to 140TB and increases the size of the average string by just 2 bytes (compared to 1 byte for nul termination).
cornholio··on Linux eliminates the strncpy API after six years of work, 360 patches
You can have a universal variable length field, for example 2 bytes for strings < 32768, then four bytes, 8 bytes etc. On the critical short string path, it costs just a single bit test. The glyph vs byte issues need to be dealt with in both formats.

The subdivision issue is a good perspective, but i would argue the performance impact of cloning substrings is dwarfed by the redundant full string reads to find length.

cornholio··on Norway imposes near ban on AI in elementary school
The point is that knowing what you don't know allows you to hedge your bets on events far into the future; my child will need to become productive 20 years from now and maintain it for 40-50 years. So, it's a bit like trying to educate a kid born in the 60s for the web era, you can't and you shouldn't even try.

What you can do though, is to offer them broad exposure to things that are interesting to them and their generation; my eastern block clone of the 8bit/48KB Spectrum computer didn't really help me excel at math, reading or history, nor was it to be the future of technology, but it did change my life significantly by letting me understand and relate to people that I couldn't otherwise have business dealings decades later.

It seems imprudent to cut children off from futuristic technology just because of a moral panic that it causes brain rot. Unless we know it's soma, a drug so powerful that it subdues volition and curtails intellectual development; we don't.

cornholio··on Norway imposes near ban on AI in elementary school
The counterargument is that kids will live in a world different from our own.

For example, in many countries children lost the ability to write cursive; that used to be a critical skill comparable to literacy itself. But in our current society, that's no longer the case and you can be very successful without it, but there are other skills, such as using technology, that became critical.

Any definitive claim to know what are the right things kids should learn in a moment of rapid technological shift is probably garbage and just a projection of our own biases.

cornholio··on Fable Converted Pylint to Rust
> rewriting a big widely used project in a stricter language is always a good thing

Always might be a too strong word. Rust is, by design, a language with low development velocity.

So you risk: 1. ossification of the current architecture and deferment of important features; or 2. reliance on AI coding to recover velocity.

Maybe for some 2 does not look like a risk, but I think it's too early to call. We have yet to see the effects of extensively using these tools on large scale projects, for years and decades.

Page 1 of 34Next →