I think OPs point though is that, like addiction, discipline is hard to maintain. Humans are irrational, and sometimes, some people just need to go cold turkey.
1,745 karma · joined March 6, 2021
I think OPs point though is that, like addiction, discipline is hard to maintain. Humans are irrational, and sometimes, some people just need to go cold turkey.
In short: explore versus exploit. Both are necessary, both are valuable.
I largely agree with the OP, because I can only see an explosion in both camps.
The explore side might be automated by compute, and that breaks the social contract of the existing academic incentive structure.
But people who are inherently curious will continue to be inherently curious. Nothing will stop anyone from exploring their curiosity, and the speed and depth at which they explore will only increase. We will build tools that enhance and automate our pedagogical compression.
And for the people who want to cook to feed the family, that automation is coming too.
Or perhaps we'll find a new ceiling: the frontier failure modes that still require a human mathematician in the loop, because discovering some class of novel insight doesn't scale with current AI architecture.
This feels more like an "e-bike for the mind" and less of a chauffeur.
He doesn't make a claim that implies "don't train" might not matter in the way you might think.
They cannot rule out that the data was trained because they have no per-user provenance tracking through the training pipeline once data is de-identified..
The entire point of de-identification is the inability to know the source of data. If the researcher forgot to hit "do not train" then that's that..
The only thing they have is a coincidence and the fact that the LLM may have used the training data that then researcher technically may have agreed to share.
Whether or not that's smoking gun of anything is hard to say. And the fact may remain that the proofs are significantly different, we do not know.
Look at how Astra scored 100% on Arc-AGI-3. It was largely because of the harness.
The harness increases the chances (often to 100% chance) of a non-deterministic LLM to perform deterministic actions.
Not only that, the harness provides the feedback that becomes training data for the model. So, over time the model bakes those lessons in, and the harness becomes less necessary and the agent becomes more efficient at some tasks.
A harness will likely always be necessary, we may hit some level of complexity or some level of compute that never allows us to bake the lessons into the model, and external tools provide the model with the ability to find leverage and make up for those short comings.
"Hey... we need to improve our porn detection AI... but we want to avoid paying for all that porn... is anyone open to taking on some personal liability for a nice bonus this year"
I'm just saying... maybe...
Perhaps you're realizing why this makes no sense?
I find it to be far more useful than when humans wrote PR descriptions. Many engineers didn't write one, and those that did were poorly written... this problem is mostly solved for us.. it still has LLMism speak.. but it's useful enough for me to get the context I need to do my review.
CI runs. Local Git hooks. Cursor also has hooks built into their agent. Other agent APIs probably have something similar.
"Seam" is an industry standard term coined by Michael Feathers in Working Effectively with Legacy Code.
To call a seam load bearing means it's performing critical work for the dependent class, perhaps a database query.
A seam that is not load-bearing would be something that is just injected for testability - maybe a date provider that provides some constant time to avoid flaky tests.
Tbh, this is quite literally the opposite of vapid. A whole book was written about them and their importance, and how to leverage them.
In my experience, Claude uses the word accurately. Code has a lot of seams, and seams are an important thing to communicate when working with code. Therefore, expect to see the word often.
Personally, I don't mind it at all. I'm glad the industry is finally standardizing our language more. Makes it easier for me to communicate with other engineers.
AI is trained on human intelligence. The hyper-scalers are squeezing every last drop of automatically verified reward, and that may get us very far. A compiler passes or fails in milliseconds for free, forever.
But.. "good design taste" has no compiler.
Of course, taste isn't unverifiable. But it's is expensively verifiable. Noisy, slow, and orders of magnitude lower throughput. People with deep domain knowledge often can't articulate well _why_ one design works and the other doesn't. So, judgment arrives as a verdict, and not a crisp rationale. I guess we'll see if sample efficiency outpaces the cost of human judgement.
In the mean time, leverage will sit with whoever holds this tacit knowledge (incumbents). I.e., hospital systems, law firms, chip designers, studios, SaaS that are dominating their niche.. and not with the labs training on it. To me, this is why valuations of companies like Palantir could potentially make sense.
Article I, Section 8, Clause 8 of the US Constitution
Which empowered Congress to "promote the progress of science and useful arts, by securing for limited times to authors and inventors the exclusive right to their respective writings and discoveries."
Scientists and the artists and their "exclusive rights" have built quite a lot over the centuries.
What this post is actually pointing out is that intellectual property that has transferrable physical representation has more value to the consumer.
And intellectual property that does not have transferable physical representation has more value to the producer.
Reselling or gifting a book you've read to a friend is wholesome.. it feels good. Truly.. but every time we do that we also take from the artist.
We draw the line somewhere because these things that "are the parents' decision" have consequences on broader society. They have consequences that impact you and me. And we also have a say.
You can make the argument that it's just the parents' decision. But you have to say why.
Reality is that was A bottleneck. Code review has historically been faster than writing the code.
That is no longer true for me. I can complete two to three PRs per day in a span of time that would have historically taken one to three days.
I now sit around doing code reviews and asking for code reviews.
Deciding what to build. Reviewing Code. And testing code. Are the new bottleneck.
So of course we don't see massive productivity gains. Because these parts of the SCLC were always bottlenecked but their capacity matched the throughout. We fired all the dedicated QAs years ago. Sr+ engineers that do all the code review are limited.
Teams have not re-organized to match the new code-input velocity.
Engineers don't want to do QA because it's "beneath them".. and most engineers don't like performing or are not Sr enough to do extensive or high quality code review.
In the same way that the sound waves and facial expressions I produce are not conscious, the output json of an LLM is obviously not conscious either.
The locus of consciousness and subjective experience may be in the computer, either at inference time or training time..
Of course, its not that simple. Some companies probably are great at scouting. Yegge mentioned a few ways in the post. Good internship programs, acquihiring, etc.