HNHacker News
TopNewBestAskShowJobs

roboboffin

153 karma · joined September 15, 2024

submissionscomments
roboboffin··on Genie 3: A new frontier for world models
If human need drives the creative process, then there will always be a human in the loop. Instead, each human becomes the “random seed” that initialises the process based on their own unique make-up. This is only different from how things work now, in that humans are also creating the artefact.

Similar to how synths meant we no longer need to play an instruments by plucking strings, it hasn’t affected the higher level creativity of creating music, only expanded it.

roboboffin··on Genie 3: A new frontier for world models
Depends what you mean by creativity. In some ways, AI is not creative at all, everything is generated by mapping text to visuals using diffusion modelling via a shared latent space. It has no agency or creative thought of its own.

Humans have demonstrated time and again, even things beyond our experience can be explored by us; quantum mechanics for example. Humans find a way to map very complex subjects to our own experience using analogy. Maybe AI can help us go further by allowing us to do this on even more complex ideas.

roboboffin··on Genie 3: A new frontier for world models
In theory, creativity is an infinite space. As technology advances it allows humans to explore more and more complex things; take the advancement of music as an example, synths, loops etc.

If humans are not stretched to their limits, and are still able to be creative, then the tools will help us find our way through this infinite space.

AI will never be able to generate everything for us, because that means it will need infinite computation.

roboboffin··on 'It is a better programmer than me': The reality of being laid off due to AI
I just had a thought. It used to be that complex C++ systems used to take so long to compile that developers used to go and have a coffee etc. This was before distributed compiling.

Maybe it will return to that, the job will have a lot of waiting around, and “free” time.

roboboffin··on 'It is a better programmer than me': The reality of being laid off due to AI
I’m not saying people will be laid off, although this is what the article is about. So, I think people will still be prompting, but if you can prompt an agent and it can happily code away, what are you supposed to be doing ? Watching it do its work ? The only option is that you will have to generate ideas of new work constantly to drive value. This is something that generally happens over time now, but as implementation becomes quicker; idea generation will have to accelerate.
roboboffin··on 'It is a better programmer than me': The reality of being laid off due to AI
The difference for me, is that things are changing too fast to keep up. For example, if a large part of your job is taken away seemingly overnight, by a new model, your whole job could change in a heartbeat.

What preparation are you supposed to do for this ? Previously, change was relatively slow and it was reasonable to keep up in your own time. I believe that is no longer possible.

roboboffin··on 'It is a better programmer than me': The reality of being laid off due to AI
I think workplaces will have to allow people time to adapt. So that if your particular skill set is replaced by AI, you have the ability to retrain to a part that isn’t.

Ultimately, large part of many jobs are repetitive, and can be replaced by pattern matching. The other side, creating new patterns, is hard and takes time. So, employers will have to take this into account. They may be long periods of “unproductive” time, or more risky evaluation to try new ideas.

roboboffin··on Magistral — the first reasoning model by Mistral AI
No worries, I wasn’t saying to you directly.

I agree 15 disks is very difficult for a human, probably on a sheer stamina level; but I managed to do 8 in about 15 minutes by playing around (I.e. no practice). They do state that there is a massive drop in performance at this point.

roboboffin··on Magistral — the first reasoning model by Mistral AI
Not sure why I am being downvoted. I am simply saying that we know there is a defined algorithm for solving Tower of Hanoi, and the source code for it is widely available. So, o3 producing the code as an answer, demonstrates even less intelligence, as it means it is either memorized or copied from the internet. I don't see how this point counters the paper at all.

I believe what they are trying to show in that paper, is that as the chain of operations approaches a large amount (their proxy for complexity), an LLM will inevitable fail. Humans don't have infinite context either, but they can still solve the Tower Of Hanoi without need to resort to either pen or paper, or coding.

roboboffin··on Magistral — the first reasoning model by Mistral AI
I think that their point was that the problem is easily solvable by humans without code, and shows the ability to chain steps together to achieve a goal.
roboboffin··on How Does Claude 4 Think? – Sholto Douglas and Trenton Bricken
Here is a link to the press release about the drug discovery:

https://www.futurehouse.org/research-announcements/demonstra...

roboboffin··on Anthropic warns fully AI employees are a year away
I think phrases like “going rogue” apportion too much agency to the AI. It is more likely that it just spirals out of control, and causes damage that way.

I can’t imagine that the AIs will just be let loose without strict monitoring, as this article alludes to.

roboboffin··on DeepSeek-R1 speeds up llama.cpp code by x2
I wonder if code becomes so complex that humans cannot understand, there may be a way to fool the LLMs into creating backdoors in the software.

For example, a open source LLM is produced, used everywhere, and it subtly inserts some malicious coding. Not saying this is happening now, but could happen.

Then when AGI comes along, this would shift to understanding the motivations of the AI and how they align with human ethics.

roboboffin··on 'Mainlined into UK's veins': Labour announces public rollout of AI
This research seems to show that two X-rays can be identified as being from the same person, rather than identifying the person directly.

However, I have no doubt that AI could easily de-anonymize data fully given enough data points.

roboboffin··on rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
Yeah, that's what I thought.
roboboffin··on 30% drop in O1-preview accuracy when Putnam problems are slightly variated
The actual quote is that they are easier than IMO; but maybe not directly comparable.

https://x.com/littmath/status/1870848783065788644?s=46&t=foR...

I think it’s more probable that it would have solve the easier problems first, rather than some hard and only some easier; although that is supposition.

Reading this thread and the blog post gives more idea about what the problems might involve.

It’s difficult to judge without more information on the actual results, but that means we cannot draw any strong conclusions either way on what this means.

roboboffin··on 30% drop in O1-preview accuracy when Putnam problems are slightly variated
I think the problems it solved were understood to be well known undergraduate problems.

https://xenaproject.wordpress.com/2024/12/22/can-ai-do-maths...

roboboffin··on Does current AI represent a dead end?
I guess the point I am trying to make, is that paradoxically the more an AI company's products are integrated into the economy, the less value they can extract from the economy. As a large amount of the world's economic output is just dealing with the human factor.
roboboffin··on Does current AI represent a dead end?
One thing I thought recently, is that a large amount of work is currently monitoring and correcting human activity. Corporate law, accounting, HR and services etc. If we have AGI that is forced to be compliant, then all these businesses disappear. Large companies are suddenly made redundant, regardless of whether they replace their staff with AI or not.
roboboffin··on OpenAI O3 breakthrough high score on ARC-AGI-PUB
However, the way it is progressing is that the SOTA is saturating the current benchmarks; then a new one is conceived as people understand the nature of what it means to be intelligent. It seems only natural to concentrate on one benchmark at a time.

Francois Chollet mentioned that the test tries to avoid curve fitting (which he states is the main ability of LLMs). However, they specifically restricted the number of examples to do this. It is not beyond the realms of possibility that many examples could have been generated by hand though, and that the curve fitting has been achieved, rather than discrete programming.

Anyway, it’s all supposition. It’s difficult to know how genuine the results is, without knowledge of how it was actually achieved.

roboboffin··on OpenAI O3 breakthrough high score on ARC-AGI-PUB
Interesting that in the video, there is an admission that they have been targeting this benchmark. A comment that was quickly shut down by Sam.

A bit puzzling to me. Why does it matter ?

roboboffin··on The Matrix: Infinite-Horizon World Generation with Real-Time Interaction
Not an expert, but I think procedurally generated terrain is generally fractal in nature, and is reproducible in that sense from a seed that is used in the generation. It is therefore recursive, as fractals are recursive.

A traditional neural network is a universal function approximator, however it is not recursive in nature, unless it is some sort of RNN. The transformer architecture, which this seems fairly similar to this one, is also not recursive in nature; although I believe, limited recursion can come about through CoT.

Therefore, I don't believe this could match the reproducibility, in an infinite sense, of a traditional procedural generator.

roboboffin··on Were RNNs all we needed?
I understand your point. I apologise, if I am coming across pendantic.

My point is computers already follow algorithms, and algorithms contain reasoning; but the computers are not reasoning themselves. At least, not yet!

roboboffin··on Were RNNs all we needed?
It’s not reasoning, it retrieval of a pattern, and that pattern may contain reasoning.

The prompt engineering is the real reasoning, provided by the human.

roboboffin··on Were RNNs all we needed?
For example, papers like this call into question whether or not a LLM can plan:

https://arxiv.org/html/2409.13373v1

This is a basic form of reasoning, to plan out the steps needed to execute something.

roboboffin··on Were RNNs all we needed?
I'm not sure that's true at all. There are several well known researchers that say LLMs are in fact not doing reasoning.
roboboffin··on OpenAI DevDay 2024 live blog
Hopefully, it won’t cause a plethora of nuisance phone calls. As the cost tends to zero, then it will be much easier to spam people; even more so than now.
roboboffin··on OpenAI DevDay 2024 live blog
Is it that most models are based on the transformer architecture ? And so performance improvements can then we used throughout their different products ?
roboboffin··on Training Language Models to Self-Correct via Reinforcement Learning
I think this is similar to this point: https://news.ycombinator.com/item?id=41601738

That the chain-of-thought diverges from accepted truth as an incorrect token pushes it into a line of thinking that is not true. The use of RL is there to train the LLM to implement strategies to bring it back from this. In effect, two LLMs would be the same and would slow diverge into nonsense. Maybe it is something that is not so much of a problem anymore.

Yann LeCun talks about how the correct way to fix this is to use an internal consistent model of the truth; then the chain-of-thought exists as a loop within that consistent model meaning it cannot diverge. The language is a decoded output of this internal model resolution. He speaks about this here: https://www.youtube.com/watch?v=N09C6oUQX5M

Anyway, that's my understanding. I'm no expert.

roboboffin··on Training Language Models to Self-Correct via Reinforcement Learning
Is this similar to the effect that I have seen when you have two different LLMs talking to each other, they tend to descend into nonsense ? A single error in one of the LLM's output and that then pushes the other LLM out of distribution.

I kind of oscillatory effect when the train of tokens move further and further out of the distribution of correct tokens.

Page 1 of 2Next →