Computing Inside an AI
willwhitney.com
willwhitney.com
Oh boy I can't wait for GPT Electron, so I can wait 60 seconds for the reply to come back and then another 60 seconds for it to render a sad face because I hit some guard rail.
Back in the pre-LLM days these types of thought pieces made sense as a call to action because the economics of creating sophisticated proof's of concept was beyond the abilities of any one person. Now you can create implementations and iterate at nearly the speed of thought. Instead of telling people about your idea, show people your idea.
The issue is training a multi modal model that can make use of said gui.
I don't believe that there is a better general interface than text however so I won't bother.
But they mention things like Oasis in the article that use a specialized model to generate games frame-by-frame.
What?
Are you trying to say it’s too expensive for a single worker to make a POC, or that one person can’t make a POC?
Either way that’s not true at all…
There have been one person software shops for a long long time.
Isn’t that limiting our perspective of AIs models to being computers and what computers can’t do, so the model can’t do.
Nah, there's a better option. Instead of a computer, we could... go for treating it as a person.
Yes, that's inverting the whole point of the article/discussion here, but think about it: the main limitation of a computer is that we have to tell it step-by-step what to do, because it can't figure out what we mean. Well, LLMs can.
Textual chat interface is annoying, particularly the way it works now, but I'd say the models are fundamentally right where they need to be - it's just that a human person doesn't use a single thin pipe of a text chat to communicate with the world; they may converse with others explicitly, but that's augmented by orders of magnitude more of contextual inputs - sights, sounds, smells, feelings, memory, all combining into higher-level memories and observations.
This is what could be the better alternative to "LLM as computer": double down on tools and automatic context management, so the user inputs are merely the small fraction of data that's provided explicitly; everything else, the model should watch on its own. Then it might just be able to reliably Do What I Mean.
AI as an animal we are trying to tame. Why does it have to be a machine metaphor?
Perhaps AI is a ecosystem upon which we all interact at the same time. The author pointed out that the one-on-one interaction is too slow for the AI - perhaps a many-to-one metaphor would be more appropriate.
I agree with the author that we are using the wrong metaphors when interacting with AI but personally I think we should go beyond repeating the mistakes of the past by just extending our current state, I.e. going from a physical desktop to a virtual „desktop“.
Selecting a metaphor implies that ones imagination is - at least partially - constrained by the metaphor. AI as a powerpoint would make using AI for anything other than presentations seem unusual since that what powerpoint is used for.
Also when the original author "models as computers" what does "computer" represent? A mainframe computer the size of small apartment, a smartphone, a laptop, turing machine or some collection of server racks. Even the term "computer" is broad enough to include many forms of interaction. I interact with my smartphone visually while with my server rack textually, yet both are computers.
At least initially, AI seems to be something completely different, almost god-like in its ability to provide us with insightful answers and creative suggestions. God-like meaning that judged from the outside, AI has the ability to provide comforting support in times of need, which is one characteristic of a god-like entity.
Powerpoint wasn't built to be a god-like provider of answers to the most important questions. It would indeed be a surprising if a PP presentation made the same impact as religious scriptures - to thousands/millions of people, not referring to individual experiences.
The authors also introduce projected pre-norm and layer-norm hash to facilitate their proofs, another sense in which it is an upper-bound on the current approach to AI, since these concepts are not standard. Nonetheless, the paper shows how allowing a number of intermediate decoding steps polynomial in input size is already enough to run most programs of interest (which are in P).
There are additional issues. This work relies on the concept of saturated attention, however as context length grows in real world transformers, self-attention deviates from this model as it becomes noisier, with unimportant indices getting undue focus (IIUC, due to precision issues and how softmax assigns non-zero probability to every token). Finally, it's worth noting that the more under-specified your problem is, and the more complex the problem representation is, then the quickly more intractable the induced probabilistic inference problem. Unless you're explicitly (and wastefully) programming a simulated turing machine through the LLM, this will be far from real-time interactive. Users should expect a prolog like experience of spending most of their time working out how to help search.
Trivia: Softmax also introduces another problem: the way softmax is applied forces attention to always assign importance to some tokens, often leading to dumping of focus on typically semantically unimportant tokens like whitespace. This can lead to an overemphasis on unimportant tokens, possibly inducing spurious correlations on whitespace, this propagating through the network with possibly unexpected negative downstream effects around whitespace.
This isn't a thing we have, or will have.
It's like saying that a computer with infinite memory, CPU and power can certainly break SHA-256 and bring the world's economy down with it.
The fact that you need infinite time for some of the stuff doesn't mean you can't do any of the stuff.
It just needs to be able to compute enough programs to be useful.
Even our current infrastructure of precisely defined programs and compilers isn't able to compute all programs.
It seems reasonable in the future be able to give an LLM the python language specification, a python program, and it iteratively returns the answer.
nowhere in this post does the author say that it's ready with the current state of models, or he'd use a foundation model for this. why the hate?
True but..
> instead of building the website, the model would generate an interface for you to build it, where every user input to that interface queries the large model under the hood
This to me seems wildly _more_ lossy though, because it is by its nature immediately constraining. Whereas conversation at least has the possibility of expansiveness and lateral step-taking. I feel like mediating via an interface might become too narrow too quickly maybe?
For me, conversation, although linear and lossy, melds well with how our brain works. I just wish the conversational UXs we had access to were less rubbish, less linear. E.g. I'd love Claude or any of the major AI chat interfaces to have a 'forking' capability so I can go back to a certain point in time in the chat and fork off a new rabbit hole of context.
> nobody would want an email app that occasionally sends emails to your ex and lies about your inbox. But gradually the models will get better.
I think this is a huge impasse tho. And we can never make models 'better' in this regard. What needs to get 'better' - somehow - is how to mediate between models and their levers into the computer (what they have permission to do). It's a bad idea to even have a highly 'aligned' LLM send emails on our behalf without having us in the loop. The surface area for problems is just too great.
ChatGPT has this feature: forking occurs by editing an old message. It will retain the entire history, which can still be navigated and interacted with. The UX isn’t perfect, but it gets the job done.
But note the free tiers for groq and cerebras are very generous.
It's a database. The WYSIWYG example would require different object types to have different UI components. So if you change what a container represents in the UI, all its children should be recomputed.
Need direct association between labels in the model space and labels in the UI space.
This is an article of faith. Posts like this almost always boil down to one or two sentences like this, on which the entire rest of the post rests. It's postulated as a concrete fact, but it's really not.
We don't know how much better models will get, we don't know that they will get good enough to accomplish the tasks the author is talking about, hell, we don't even know if we'll ever see another appreciable increase in model quality ever again, or if we've already hit a local maximum. They MIGHT, by they also might not, we don't know.
This post would be a lot more honest if it started with "Hey, wouldn't it be neat if..."
After every response, ask the model for some alternatives, things that could be adjusted, etc. have it return a json response. Build the UI from that. Clicking on an element just generates a prompt and sends it back into the model
It's like smart replies or the word suggestions that are popping up on my virtual keyboard now but with a richer UI. It's not perfect, but I think it would be an improvement for many things.
Here's a thread https://x.com/xundecidability/status/1867044846839431614
And example function for Gemini written in shell, where the system prompt is the function definition that interacts with the model. https://github.com/irthomasthomas/shelllm.sh/blob/main/shelp...
Tangentially, I have considered the possible impact of thermodynamic computing in its application to machine learning models.
If (big if) we can get thermodynamic compute wells to work at room temperature or cheap microcryogenics, it’s foreseeable that we could have flash-scale AI accelerators (thermodynamic wells could be very simple in principle, like a flash cell)
That could give us the capability to run Tera-parameter models on drive-size devices using 5-50 watts of power. In such a case, it is foreseeable that it might become more efficient and economical to simulate deterministic computing devices when they are required for standard computing tasks.
My knee jerk reaction is “probably not” but still , it’s a foreseeable possibility.
Hard to say what the ramifications of that might be.
Instead of showing a “discoverable” palette of buttons and widgets which is limited by screen space just ASK the model what it can do and make sure it can answer. People obviously don’t know to do that yet so a simple on screen prompt to the user will be necessary.
Yes we should have access to “sliders” and other controls for fine tuning the output or maintaining a desired setting throughout generations but those are secondary to the ability of the models to make sweeping and cohesive changes and provide many alternatives for the user to CHOOSE from before they get to the stage of making fine grained adjustments.
And as to the sentence itself, I'm unclear on what exactly it's saying; people have been using other people at tools from before recorded history. Leaving aside slavery, what is it that you would say that HR departments and capitalism in general do?