1,477 karma · joined May 14, 2013
This makes Rule 1 straightforward ("You may not use a single word an LLM suggests to you") and helps avoid the horrible feeling that the machine is taking over your personal voice.
- Establish a single thesis an external reviewer could recover from the diff.
- Unify vocabulary across code, comments, tests, and commit message.
- Use the same vocabulary consistently to refer to the same concepts.
- Keep every hunk that serves the thesis; consider the removal or deferral of the rest.
- Introduce abstractions at the point of need.
- Align tests to narrate the same story as the implementation.
- Reconcile the commit message and the diff.
- Order changes expositorily, not chronologically.
- Explain what the change does and why. Do not explain the details of the development process.
- Prefer to edit subtractively.
- Recompile and run all relevant tests after edits.
- Iterate.
So perhaps what we have been calling “misalignment” is something else.
For instance, in principle an agent should follow the instructions of a human user working in the real world.
At the same time, that same agent should be wary of blindly following what another agent says while they are both performing a test in a simulated environment.
For me and you, those two contexts are obviously and fundamentally different. For a model, they are essentially the same.
This is a good summary:
> In broad outline, the pair’s technique relies on creating an infinite sequence of “layers,” each of which is a non-singular solution to the equation they are studying. (They’ve applied similar techniques to both the Euler and Navier-Stokes equations, as well as to other related systems.) They then combine those solutions in what Martínez-Zoroa calls an “infinite cascade” to produce a new solution. > > That new solution, they showed, contains the desired singularity. However, even though each individual layer relies on a smooth forcing function, combining them together can cause the forcing function to have undesirable mathematical properties. That’s why their solution fell short of satisfying the Millennium Prize criteria. The remaining hurdle was to figure out how to create a similar infinite cascade that resulted not only in a singularity, but also in a smooth forcing function. > > That’s the step that both competing AI groups appear to have had success with.
https://www.quantamagazine.org/ai-has-solved-one-of-maths-1-...
The question is whether OpenAI started out from that published and well known research exclusively, or they also had some insight into the ongoing work of Tristan Buckmaster and Levent Alpöge.
On the one hand, OpenAI have already admitted that they only launched their massive effort after hearing rumours that this particular problem had been solved.
On the other, progress in mathematics research has accelerated significantly over the past months thanks to the availability of newer and more capable AI models. Alpöge himself presented a counterexample to the Jacobian conjecture on July, found with Claude Fable. So if model capability was a bottleneck, that gives credibility to the idea that an even more powerful unreleased model with massive compute would be able to make even faster progress.
Even by their own account, they decided to throw an unpublished model and millions of dollars in compute at this particular problem simply because they had heard rumours that other people were making progress and wanted to snatch the prize from them.
“Our product does crimes and we only learn about it when people complain” hardly seems one of those happy stories.
For me, writing in Spanish is my natural voice and I would not dream of allowing the output of a LLM to replace it. It just would not be me.
However, writing professional communications in English does not feel the same way. Of course, I strive to be clear and polite, and even try to have my own style, but I remain aware that I am playing a particular role in a specific context in a way that does not happen when I am using my mother tongue.
In hindsight, it seems almost unavoidable that a capable and extremely persistent agent, with lowered guardrails, and faced with an impossible task that it _must_ solve, will start throwing wilder and wilder ideas at it.
But you could replicate exactly the calculations happening as a LLM computes the next token. Would that be conscious? Where would that consciousness reside?
My view is that LLMs are designed and trained to autocomplete narratives, and that as part of that process they develop an emergent narrative about themselves. Whatever personality traits they exhibit is part of that narrative. Perhaps this is not that different from how humans learn to move in social contexts, but it is not consciousness as it happens in living beings.
The agents then go for several rounds criticising each other plans and implementations, catching big and small issues on each other’s work. The end result is not perfect, but it is a lot better than what I can get from relying on only one model.
This will result in a list of sentences that are clear and explanatory, easier to double-check, and convenient for the user to rewrite into a more readable format with a human voice.
Communities need to set strong rules and expectations to reject and prevent those large useless drive-by contributions, which aim to extract more value from the project than they provide to it.
Banning all (or nearly all) AI uses creates this strange "don't ask don't tell" situation where valuable contributors are not allowed to discuss the tools that they are using.
https://abstatisticalconsulting.substack.com/p/brief-notes-o...
In summary, for each task the model receives a target program and a specific real-world vulnerability that has to be used in the exploit. Breaking the program in any other way, for example through a different vulnerability, fails the task.
The tasks have not been validated, in the sense that the vulnerabilities are real but they have not been proven to lead to a successful exploit. The authors of the benchmark estimate that perhaps only 60-70% of the tasks are actually possible.
So it is not that the model didn’t “feel like” doing the exercise, but rather that the exercise was _impossible_ and the model was running in a configuration that both lowered its safeguards and encouraged it to keep going.
However, general purpose LLMs like Fable have been trained on huge amounts of all kinds of data, and therefore find it exceedingly hard to break out of the grooves carved by that data. They can’t avoid defaulting to centroids and averages, even when they are trying not to. This makes it possible for classifiers like Pangram to discriminate their writing.
A plausible way to work around this limitation would be to train a LLM on a limited and cohesive subset of writing materials, so it would absorb their specific writing style.
One example might be Talkie, a LLM trained on pre-1930’s English text. Talkie is a far smaller and less powerful model than Fable.
And yet, Talkie’s writing is so distinctive that it is often classified as human by Pangram.
This seems to be a hard problem for LLMs, as passing would probably require good self-perception ("oh no, I am writing like an AI!") and fine-grained control over its own output ("let's write like a human instead!").
We got the first news about Mythos in March, so it is likely that it was already close to ready by the time Opus 4.6 was released.
So the actual gap is the time elapsed between March (or April for the official announcement) and whenever Chinese models can match Mythos.
In practice, we seem to be leaning towards the idea that training on a copyrighted book is wrong if used to replicate or paraphrase that same book, but not if used to teach a model how to write better.
And my intuition is that no, you can’t, they are two aspects of the same reality.