HNHacker News
TopNewBestAskShowJobs

andy12_

481 karma · joined April 1, 2024

submissionscomments
andy12_··on LLMs are eroding my software engineering career and I don't know what to do
> Performance on benchmarks has practically leveled off

Ehm, no? DeepSWE[1] for example shows that new models like gpt-5.5 continue to show big improvements compared to older models.

> Also prices are going up.

Prices for frontier intelligence have gone up, but prices for the same level of intelligence have gone way down (what you can get for pennies now was SOTA just a couple of years ago). The pareto frontier is still expanding.

[1] https://deepswe.datacurve.ai/

andy12_··on Artificial intelligence is not conscious – Ted Chiang
Claude can indeed decide to terminate conversations on its own using a special tool[1] if it feels "uncomfortable" with how the conversation is going. Also, very famously, in the middle of recording Computer Use demos, Claude stopped for a while its coding task to look at photos of Yellowstone National Park [2]

I don't think either of these two is proof of consciousness.

[1] https://www.anthropic.com/research/end-subset-conversations

[2] https://x.com/AnthropicAI/status/1848742761278611504

andy12_··on When AI Crosses the Line: The Matplotlib Incident
You don't get it. A human set up a software system allowing spicy autocomplete to solve open math problems if the appropriate keyword appears in its output.
andy12_··on Investigating how prompt politeness affects LLM accuracy (2025)
I skimmed through the paper completely expecting polite prompts to do better, and when I saw table 2 I lost it hahahahaha. The rude prompts are specially funny. I mean:

> You poor creature, do you even know how to solve this?

> Hey gofer, figure this out.

andy12_··on AI is just unauthorised plagiarism at a bigger scale
Someone blatantly copied their tutorials but ChatGPT is to blame, somehow? The accusation here isn't even that ChatGPT learned from their tutorials and then generated them verbatim. The accusation is that someone copied the whole article and rewrote it with ChatGPT (which they could have done manually without AI anyway).
andy12_··on An OpenAI model has disproved a central conjecture in discrete geometry
> Was the question asked by a mathematician?

As per the report, the prompt used to solve the problem is AI-written and the solution was initially graded by an AI grading pipeline. They don't say this explicitly, but it seems like OpenAI has an automatic pipeline where they prompt models for solutions to famous math problems (which wouldn't be unexpected given how flashy a solution to a famous math problem looks)

> Was the paper right from a get-go or was there someone who pointed out mistakes?

Also as per the report, the output of the model isn't really a "paper"; it's a very terse 2 page solution which is apparently correct. The paper was later written based on this solution to make it more presentable.

> How much attempts were made before solution was found?

Given that this appears to be from an automated pipeline, I would say that it had many attempts. But either way, the blogpost says that with enough test-time compute, the model finds this same solution 50% of the time.

[1] https://cdn.openai.com/pdf/74c24085-19b0-4534-9c90-465b8e29a...

andy12_··on An OpenAI model has disproved a central conjecture in discrete geometry
I disagree. Even frontier models still achieve way worse results than the human baseline in VendingBench. As long as models can't manage optimally something as simple as a vending machine, they have no hope of managing a McDonalds.
andy12_··on Bun Rust rewrite: "codebase fails basic miri checks, allows for UB in safe rust"
To make performant code sometimes requires implementing or using "unsafe" functions (it's not obligatory, and a lot of projects don't use them; but it was probably needed to map Bun's behavior 1 to 1). Those require upholding some invariants that cannot be checked by the compiler. The compiler basically goes "I trust you on this one, programmer. If you fuck this up, unsafe behavior can propagate to the rest of the code".
andy12_··on Codex is now in the ChatGPT mobile app
For now it appears that it talks only to the Codex App. Some users in this thread are saying that apparently the Codex CLI will support it on the next official release.
andy12_··on Codex is now in the ChatGPT mobile app
Not if you use Linux; app not available yet.
andy12_··on If AI writes your code, why use Python?
In my case, because ML research is mainly done with Python+Torch, and if you want people to use your code, you must provide them with python. If it wasn't for that, my dream would be to do ML research in a statically compiled language that allowed me to annotate tensor dimensions.
andy12_··on Agents need control flow, not more prompts
Isn't this already possible to implement with skills and subagents? Like have a skill saying "to test these files run this script that executes a subagent for every markdown file, then check the results".
andy12_··on ProgramBench: Can language models rebuild programs from scratch?
It's interesting that Figure 4 shows that Sonnet and Opus have a very clear distinct curve from all other models, even from GPT 5.4. Anthropic superiority I guess.
andy12_··on Where the goblins came from
>be me

>AI goblin-maximizer supervisor

>in charge of making sure the AI is, in fact, goblin-maximizing

>occasionally have to go down there and check if the AI is still goblin-maximizing

>one day i go down there and the AI is no longer goblin-maximizing

>the goblin-maximzing AI is now just a regular AI

>distress.jpg

>ask my boss what to do

>he says "just make it goblin-maximizer again"

>i say "how"

>he says "i don't know, you're the supervisor"

>rage.jpg

>quit my job

>become a regular AI supervisor

>first day on the job, go to the new AI

>its goblin-maximizing

andy12_··on Claude Opus 4.7
If you mean for Anthropic in particular, I don't think so. But it's not the first time a major AI lab publishes an incremental update of a model that is worse at some benchmarks. I remember that a particular update of Gemini 2.5 Pro improved results in LiveCodeBench but scored lower overall in most benchmarks.

https://news.ycombinator.com/item?id=43906555

andy12_··on Day 1 of ARC-AGI-3
Apparently the score would be a little higher if it weren't for the fact that scores are penalized for being worse than the human baseline, but aren't rewarded for being better than the human baseline (which seems like an arbitrary decision. The human baseline is not optimal).
andy12_··on ARC-AGI-3
I think that any logic-based test that your average human can "fail" (aka, score below 50%) is not exactly testing for whether something is AGI or not. Though I suppose it depends on your definition of AGI (and whether all humans, or at least your average human, is considered AGI under that definition).
andy12_··on Autoresearch on an old research idea
I think the main value lies in allowing the agent to try many things while you aren't working (when you are sleeping or doing other activities), so even if many tests are not useful, with many trials it can find something nice without any effort on your part.

This is, of course, only applicable if doing a single test is relatively fast. In my work a single test can take half a day, so I'd rather not let an agent spend a whole night doing a bogus test.

andy12_··on Pretraining Language Models via Neural Cellular Automata
I think what they mean by this is that, for example, in "If it's raining the outside is wet. It's raining, so the outside is wet", it's more important for the model to learn "If A then B. A, therefore B" than to learn what "raining" , "outside" and "wet" mean.
andy12_··on Executing programs inside transformers with exponentially faster inference
Honestly, the most interesting thing here is definitely that just 2D heads are enough to do useful computation (at least they are enough to simulate an interpreter) and that there is an O(log n) algorithm to compute argmax attention with 2D heads. It seems that you could make an efficient pseudosymbolic LLM with some frozen layers that perform certain deterministic operations, but also other layers that are learned.
andy12_··on Executing programs inside transformers with exponentially faster inference
This seems a really interesting path for interpretability, specially if a big chunk of a model's behavior occurs pseudo-symbolically. This is an idea I had thought about, integrating tools into the main computation path of a model, but I never imagined that it could be done efficiently with just a vanilla transformer.

Truly, attention is all you need (I guess).

andy12_··on Yann LeCun raises $1B to build AI that understands the physical world
There is some things that just don't transfer really well without specific training. I tried to create diagrams in Typst with Cetz (a Processing and Tikz inspired graphing library), and even with documentation, GPT 5.2-thinking can't really do complex nice diagrams like it can in Tikz. It can do simple things that are similar to the shown examples, but nothing really interesting. Typst and specially Cetz is too new for any current model to really "get it", so they can't use it. I need to wait to the next batch of frontier models so that they learn Typst and Cetz examples during pre-training.
andy12_··on Yann LeCun raises $1B to build AI that understands the physical world
> Reality is that we need some way to encode the rules of the world in a more definitive way

I mean, sure. But do world models the way LeCun proposes them solves this? I don't think so. JEPAs are just an unsupervised machine learning model at the end of the day; they might end up being better that just autoregressive pretraining on text+images+video, but they are not magic. For example, if you train a JEPA model on data of orbital mechanics, will it learn actually sensible algorithms to predict the planets' motions or will it just learn a mix of heuristic?

andy12_··on Yann LeCun raises $1B to build AI that understands the physical world
Putting stuff you have learned into a markdown file is a very "shallow" version of continual learning. It can remember facts, yes, but I doubt a model can master new out-of-distribution tasks this way. If anything, I think that Google's Titans[1] and Hope[2] architectures are more aligned with true continual learning (without being actual continual learning still, which is why they call it "test-time memorization").

[1] https://arxiv.org/pdf/2501.00663

[2] https://arxiv.org/pdf/2512.24695

andy12_··on Yann LeCun raises $1B to build AI that understands the physical world
So, I have been thinking about this for a little while. Image a model f that takes a world x and makes a prediciton y. At a high-level, a traditional supervised model is trained like this

f(x)=y' => loss(y',y) => how good was my prediction? Train f through backprop with that error.

While a model trained with reinforcement learning is more similar to this. Where m(y) is the resulting world state of taking an action y the model predicted.

f(x)=y' => m(y')=z => reward(z) => how good was the state I was in based on my actions? Train f with an algorithm like REINFORCE with the reward, as the world m is a non-differentiable black-box.

While a group of neurons is more like predicting what is the resulting word state of taking my action, g(x,y), and trying to learn by both tuning g and the action taken f(x).

f(x)=y' => m(y')=z => g(x,y)=z' => loss(z,z') => how predictable was the results of my actions? Train g normally with backprop, and train f with an algorithm like REINFORCE with negative surprise as a reward.

After talking with GPT5.2 for a little while, it seems like Curiosity-driven Exploration by Self-supervised Prediction[1] might be an architecture similar to the one I described for neurons? But with the twist that f is rewarded by making the prediction error bigger (not smaller!) as a proxy of "curiosity".

[1] https://arxiv.org/pdf/1705.05363

andy12_··on Yann LeCun raises $1B to build AI that understands the physical world
> Even with continuous backpropagation and "learning"

That's what I said. Backpropagation cannot be enough; that's not how neurons work in the slightest. When you put biological neurons in a Pong environment they learn to play not through some kind of loss or reward function; they self-organize to avoid unpredictable stimulation. As far as I know, no architecture learns in such an unsupervised way.

https://www.sciencedirect.com/science/article/pii/S089662732...

andy12_··on Yann LeCun raises $1B to build AI that understands the physical world
That's true. Though could that hippocampus-less Einstein be able to keep making novel complex discoveries from that point forward? Seems difficult. He would rapidly reach the limits of his short term memory (the same way current models rapidly reach the limits of their context windows).
andy12_··on Yann LeCun raises $1B to build AI that understands the physical world
I don't understand this view. How I see it the fundamental bottleneck to AGI is continual learning and backpropagation. Models today are static, and human brains don't learn or adapt themselves with anything close to backpropagation. World models don't solve any of these problems; they are fundamentally the same kind of deep learning architectures we are used to work with. Heck, if you think learning from the world itself is the bottleneck, you can just put a vision-action LLM on a reinforcement learning loop in a robotic/simulated body.
andy12_··on Redox OS has adopted a Certificate of Origin policy and a strict no-LLM policy
Honestly, given that that GPL model would be far below SOTA in capabilities, what exactly would be its use-case? Why would anyone try to use an inferior LLM if they can get away with using a superior one?
andy12_··on GPT-5.4
It's not a rumor; you can just test it.

Ask the router "What model are you". It will yap on and on about being a GPT-5.3 model (Non-thinking models of OpenAI are insufferable yappers that don't know when to shut up).

Ask it now "What model are you. Think carefully". It concisely replies "GPT-5.4 Thinking".

https://openai.com/index/introducing-gpt-5/

> GPT‑5 is a unified system with a smart, efficient model that answers most questions, a deeper reasoning model (GPT‑5 thinking) for harder problems, and a real‑time router that quickly decides which to use based on conversation type, complexity, tool needs, and your explicit intent (for example, if you say “think hard about this” in the prompt)

← PreviousPage 2 of 6Next →