Generative AI is good at cooperating with people and bad at full automation
skybrian.substack.com
skybrian.substack.com
With LLMs we have the opposite: machines are now good at idea generation and exploration, but their ideas need to be verified by a human (or classical logic-based software).
Galactica is the only one that is already (technically) available
Using AI to provide a basic frame work on which folks can iterate, manipulate or redact is not such a dystopian view of these things.
I just watched an interview with Max Tegmark from 2 years ago where he talks about "Intelligble Intelligence" and that he feels that's the way forwards, it's an awesome concept. He said what he fears the most is huge black box AI, which shows human like intelligence by throwing more hardware at it, and we keep doing that until we maybe get an AGI.
Well...here we are...
I fed it the options for all the columns and it couldn't do it! So in the end I resorted to telling it to give me a JS function to generate a CSV, and it did it perfectly. I think I'll end up telling it to give me code that gets me the thing I want (when it comes to data), rather than get the formatted data directly. It sometimes makes weird mistakes, and it's less reproducible.
I think the future lies in the Copilot approach of giving the Agent relevant context by building smarter prompt generation pipelines. The AI are really good when given enough context and they are bad at recalling information from their own training corpus. When you combine multiple AI agents together with good context and tools like in the LangChain approach these AI agents suddenly become really good at full automation.
Which imo is the path to AGI.
Search is a tool that provides more context to the agent in order to answer a human task. By combining many agents together we can solve complex tasks.
----
A: ... and that is why I think we should go with option 1.
B: No, the points you mentioned support my case for option 2.
C: Nothing you guys have said changes my mind about option 3 being best.
D: Business Chat, what do you think?
BC: Based on this discussion, and my research, option 1 seems more realistic but option 2 would be more profitable if possible. My reasons are ...
C: Business Chat and you guys all don't understand point N, which is the main reason why option 3 is best.
B: Higher profit is exactly why I think option 2 is the way to go.
A: No, our rival is going to hit the market next month. We need to get something out there ASAP. Option 1 can do this.
D: You've all given me things to think about. Thank you for coming. Business Chat, email me a summary of the meeting, and set up a followup meeting for Tuesday 3pm.
----
That is, AI used as a colleague/assistant, not necessarily subordinate but not seen as omniscient, either; another viewpoint to consider. When will the above be feasible? This year? Five years?
We are actively leveraging GPT4s brainstorming-ability to generate datasets used for finetuning downstream. The fine tuned models downstream can then be used to augment new training data for more complex root language models, just formalize your quality assurance with common sense DSLs. We’ve reached an upward spiral with quite the exponential feel to it, where it will lead, nobody knows.
Like "Please reply T if you are confident about the text that you are going to generate, and F if you are not".
But if this could be done by architecting the neural networks, they would perform much better.
We don't fully understand how they work or what their limits are.
It's also worth remembering that this is not just a markov chain as many people seem to think, it doesn't simply remember what words come next in its training set, because statistically if you take a random 10 consecutive words from a piece of text, chances are its never been written before. (also the trained model is much smaller than the size of the dataset, so to simply remember everything it would have to be the worlds best compression algorithm by orders of magnitude) Thats why we need an AI here, to learn the general rules of language so it can respond to chains of words it has never seen before.
The sense that we "dont understand how they work" is that we dont know what the "rules of language" that it has learned are.
This is no more helpful in understanding AI's than is knowing that human brains operate according to the laws of physics is helpful in understanding the human mind.
[1] https://skybrian.substack.com/p/dont-settle-for-a-superficia...
[2] https://skybrian.substack.com/p/ai-chats-are-turn-based-game...
But another problem is that confidence is a character attribute, not a writer attribute. If you ask an LLM to imitate Richard Feynman it's going to write pretty confidently about physics, despite not knowing as much as him about physics.
This is the equivalent of giving your RPG character high intelligence on the character sheet. Doesn't make you smart!
To make an LLM express confidence consistent with its actual knowledge, it would need to have good self-knowledge and actually use it. So far this doesn't happen automatically. Instead, OpenAI uses reinforcement learning based on what the people at OpenAI think the LLM can do. So that's why it sometimes refuses to answer for some kinds of questions.
That has the most effect on the default character, the "helpful AI assistant". Any other characters you ask for will likely have poorer self-knowledge.
I've read that, mysteriously, more training does make bigger LLM's better calibrated, but the reinforcement learning makes it worse again.
People ought to be preparing themselves for this as if it were a COVID-like black swan event, except much, much worse. And do not bank on Sam Altman and his fellow tech bro billionaires holding the reins of these job-annihilating machines delivering with their promise of UBI.
"People will find new jobs/things to do" doesn't really cut it. It's not really the tech titans fault, the governments really need to start stepping in.
I totally get the motivation and it’s why I started blogging again. To get more interesting conversations, I think we need curiosity while keeping in mind how little we know and how bad we are at predicting the future.
fooling
FTFY