Gemini 2.5 Pro reasons about task feasibility
everything.intellectronica.net
everything.intellectronica.net
I added a section about that to my review last night describing two of the larger examples: https://simonwillison.net/2025/Mar/25/gemini/#update-it-s-ve...
(It's always risky saying anything like this on a forum like Hacker News because it's inevitable someone will find a way to argue that the examples are trivial/unrealistic/show I don't know what I'm doing/clearly just regurgitated from StackOverflow/etc, but I'll take the risk anyway.)
Here it one-shotted a fully functional LISP interpreter for me: https://everything.intellectronica.net/p/the-little-lisper
Turns out OpenAI's llms are pretty decent at coding x86_64 bios bootloaders in assembly, but as soon as you go off script from the two main examples online, it falls apart really quickly, as it's crystal clear it has no idea what is actually going on or the limitations of how bootloaders (and 2, and 3 stage bootloaders) work.
Quite a few happy people at Google today I bet.
Which leads me to wonder, it's not like the Gemini 2 models were terrible either - they consistently were in top 5 if not top 3, now they've smashed past everything with a +40 elo.
Are we starting to see Google apply their compute/resources/data/money to assert dominance? What next from the recently-pretty-quiet Open AI? Are we getting to the stage where well-funded startups like Anthropic et al simply cannot compete with "google-scale" for general purpose models and end up as coding-only niche models? Sure you can throw GPUs at the problem and burn more investor cash, but are Google starting to run away with it with their data and infrastructure advantages? Who even comes close when you factor in data? Meta are the only people I can think of, but their data must be quite narrow (basically social graph and short-form videos and ad click data?)
Exciting times.
https://github.com/simonw/files-to-prompt https://llm.datasette.io/
What is the risk? Do you lose anything? Its like saying its risky to say these AI are not useful since people might say you are just using it wrong etc.
Reminds me of the time when I discovered GTD. Don’t worry, we will find a way to become overwhelmed again.
“Create a music synthesizer that I can run in my browser. Give it original functionality, with buttons and other controls that create pleasing but unexpected sounds, melodies, rhythmic patterns, etc. as well as various sliders and dials that can be used to modify the sounds and patterns.”
Here is the result:
https://www.gally.net/temp/202540326geminisynthesizer/index....
Here’s an excerpt from the reasoning trace:
> Brainstorm "Original" and "Unexpected" Functionality:
> Standard Synths: Oscillators (sine, square, saw, triangle), filters (lowpass, highpass), envelopes (ADSR), LFOs. These are common. How to make them unexpected?
> Generative Aspects: Instead of just playing notes the user clicks, what if the synth generates patterns?
> Melody Generation: Use algorithms like random walks, cellular automata, or simple probability rules.
> Rhythm Generation: Euclidean rhythms, probabilistic triggers, evolving patterns. ...
> Interface Concept: How can the controls themselves be part of the "unexpected" experience? Buttons that trigger complex events, sliders with non-linear responses, dials that control abstract concepts.
[0] https://aider.chat/ [1] https://github.com/dnakov/anon-kode [2] https://openrouter.ai/ [3] https://t3.chat/
https://github.com/danielmiessler/fabric?tab=readme-ov-file#...
This is more rudimentary and works on the CLI, but I've had good results with it using both Gemini Pro and local models.
https://www.gally.net/temp/202540327geminisynthesizer-v2/ind...
This was the prompt I gave to it (through a spoken interface, thus the length and repetition):
“Attached is a website I had you create for me yesterday based on the prompt that appears in another attached file. In that latter file I've also included your thinking process in response to my prompt as well as your explanation to me of how this synthesizer is supposed to work. I am basically happy with the synthesizer you created for me. It works very well, and the output is fascinating to listen to. But I would like the music produced by it to be more melodical and contrapuntal, that is, with more distinct notes that can be perceived forming melodies while still having the random and unexpected and creative generation of those melodies. I would also like to have a broader frequency range of tones that are being produced. For example it would be nice to have something like a bass line. Continue to make the music unexpected and creative and generative. That was one aspect of the music that was very positive for the first result: the fact that I could keep listening to the produced music for a long period of time and not get bored by it. So try to make the tone soundscape richer, more complex and with more sense of melody and counterpoint. Also add any more controls you can think of to make the, to give the user even more ways in which to affect the output, such as more fine tuning on the degree of tonality vs. atonality, conventional harmonic structures vs. unconventional harmonic structures, clear rhythmic patterns vs. unconventional rhythmic patterns, etc.”
The first result had a lot of digital clipping in the output on my M1 Mac mini. After some back and forth with Gemini about possible causes and solutions, it added a limiter and some more controls. The problem persists on the Mac mini. On my M4 iPad with Safari, the sound is clean. I kind of like it.
My feeling is that we have the pieces to build AGI. Like humans, we don't need a 400IQ person to solve all problems ('AGI'). What we have is coordination problems and in LLM land it's 'the glue' that's missing. Hopeful it's a matter of patterns/best-practices emerging.
Yes! I share the feeling that once LLMs get good enough at some abstraction level, you can always put another "level" on top that should abstract what already works into bite sized pieces. Hassabis also mentions this in a recent podcast, different levels of abstraction. We'll probably see some tooling in this space shortly, to coordinate between the different levels. And then RL it and watch it demolish planning tasks benchmarks.
We might very well already be at the point where every level is achievable, we just have to glue them together.
It almost certainly can. Try asking Gemini 2.5 Pro to do that and see what happens.
I mean, it's nice when the models can integrate the step-by-step internally... but I feel people have been missing out on the complex interactions by expecting it all in one adhoc prompt.
One app that comes to mind is Google's Conversational agents. The routing is just done by referencing another agent in the instructions, no need to explicitly link beyond the prompt.
Memory for one, not only do models need to be able to have long term, short term memory but they also need to be able to selectively forget. Hallucinations are still a big problem, you can easily (unintentionally) put the models in situations where they make up facts. Context limits - comprehension limits are still effectively 8-10k even though the token limits have been raised to infinity.
You either fine-tune which is a very lossy process that degrades generality or you do in-context learning/RAG. Forgetting in its current form would be eliminating obsolete context, not forgetting would be using 1 million input tokens to answer "what is 2+2?".
In any case, any external mechanic to selectively manage context would be far too limiting for AGI.
But for AGI, we're indeed missing an short term memory system with the ability to record the passage of time and filter out information not relevant to the task at hand, but I don't think they should be neural networks like we humans have. Neural networks for storing information is the only thing biology had to work with, but that doesn't mean it's the best solution for AGI, and I don't think the path to AGI is an complicated end-to-end neural network model.
AGI, no matter the level of consciousness* you aim for, will probably end up being more like an OS where processes are agents that work together. You'd have long running agents, short running agents, agents that analyze data, agents that apply algorithms, agents that come up algorithms, agents that criticize and fact check, agents that classify memories of other agents, agents that produce data for other agents to use in generating new models and supervising agents and interface agents that runs continuously to interact with the world and / or users.
*= which i define as the ability to understand that you are an entity existing in an environment that can be affected by an action, and also the ability to understand that an observed change in the environment might have been due to a previous action that you remember doing. This understanding can come on different levels and is mainly due to how detailed and fleeting your short-term memories are.
> I have never seen an LLM do this
Interestingly, many of the program we use provide a finite set of functionality that we can discover over time. But LLM's are different: you can't explore them because the input space is too big. Therefore, they can surprise us for a long time. That's cool!
Maybe being an LLM psychologist is a job with future.
I've seen this happen with GPT-4 with zero shot prompts. Similar to the author "negotiating" allowed it to continue with an iterative approach.
The model is unlikely to know its own limits. Hopefully these refusals are amenable to prompt engineering: “even if the task seems infeasible, try anyway.”
And hopefully next-gen models are trained to have more faith in themselves :)
(this is followed by long spec of the RB-338 which I also generated and is too long to include here, and a screenshot).