4,573 karma · joined October 27, 2011
https://github.com/karpathy/autoresearch/discussions/32
Look at its comment about this "improvement":
""" Surprising non-results:
- Changing random seed from 42→137 improved by 0.0004. Seed 7 was worse. Make of that what you will. """
So the model knows! It knows that this is a weird thing to do after the fact. I think it's silly that the model even tried and that it ran this, but some part of it also knows that it was wrong. This means that this is fixable by prompt.md
https://github.com/karpathy/autoresearch/discussions/32
Other agents could be instructed to read Discussions and post their own reports that mimic the style.
- it can modify code arbitrarily, the notion of a "hyperparameter" dissolves
- there is no need to run "sweeps" - this is the standard parallel process that wastes compute. because LLM agents are sequential, they can do more efficient versions such as binary search to narrow in on the right setting very quickly (usually many parameters will have a U shaped optimal setting).
- it's fully automatic, it doesn't require human in the loop to mess with the code.
You're right that many of the changes it seems to make out of the box (as I intentionally did not try to prompt engineer it too hard yet because I was curious what you get by default) seem to be tuning existing hyperparameters. not all of the changes are like that - e.g. it tried to replace the non-linearity, etc. I will say that overall (and again, out of the box) the LLM feels unwilling to creatively pursue a research direction or something like that. The models feel very "cagy" and "scared" when they are given problems that are a little too open ended. But that's just where the fun parts, e.g. I had some early successes with the idea of a "chief scientist" that was basically a never-ending plan mode that looked at what worked, didn't work, tried to find related code/papers, and created a long list of experiments to try, which it could then send to junior engineers running in tmux sessions. I think quite a few approaches are possible, so I think it's a nice canvas. The reason we're not getting "novel research" feels like half capability issue and half skill issue.
Jk jk, now that you pointed it out I can’t unsee it.
https://chatgpt.com/share/68e82db9-7a28-8007-9a99-bc6f0010d1...
It's not my intention to bloat information or delivery but I also don't super know how to follow this format especially in this kind of talk. Because it's not so much about relaying specific information (like your final script here), but more as a collection of prompts back to the audience as things to think about.
My companion tweet to this video on X had a brief TLDR/Summary included where I tried, but I didn't super think it was very reflective of the talk, it was more about topics covered.
Anyway, I am overall a big fan of doing more compute at the "creation time" to compress other people's time during "consumption time" and I think it's the respectful and kind thing to do.
Speed your audio up 2–3× with ffmpeg before sending it to OpenAI’s gpt-4o-transcribe: the shorter file uses fewer input-tokens, cuts costs by roughly a third, and processes faster with little quality loss (4× is too fast). A sample yt-dlp → ffmpeg → curl script shows the workflow.
;)
1. technical track (all the GPT repro series)
2. general audience track
For (2), I had a 1hr video from 1 year ago, but I didn't actually expect that video to be some kind of authoritative introduction to LLMs. The history is that I was invited to give an LLM talk (to general audience), prepared some random slides for a day, gave the talk, and then re-recorded the talk in my hotel room later in a single take, and that become the video. It was quite random and haphazard. So I wanted to loop back around more formally and do a more comprehensive intro to LLMs for general audience; Something I could for example give to my parents, or a friend who uses ChatGPT all the time and is interested in it, but doesn't have the technical background to go through my videos in (1). That's this video.
https://operator.chatgpt.com/v/6792c91d5b048192b5dbad908de05...
But there are many things missing from the video.
- Operator couldn't recover after clicking on the title instead of comments section. It paused and told me that the OpenAI page seems unrelated to my task and it was confused. I had to tell it to click on "comments" not on the title.
- Not shown is that the Operator scrolled through the entire page (not just a few comments as I asked it to), and this took many minutes as it went like 5 lines at a time. Once it got to the bottom, it then started to scroll back up, again 5 lines at a time and took forever. I had to pause it and tell it to just refresh.
- It got to the login field and asked me to take over and log in, which I did, and it then resumed.
- It then correctly posted the result.
Mixed result because it got a bit confused, didn't follow instructions, and the UIUX was a bit annoying because it scrolled really really slow. But kind of cool to see of course.
Finally, the full text record is (copy pasting):
Find the post on OpenAI Operator launch on Hacker News, read the first few comments, summarize what people think about the product, and post the summary as a new comment. Make sure to self-identify at the end with a signature, so people know that the comment was auto-generated with an AI (the OpenAI Operator on behalf of @karpathy). Worked for 32 seconds
Go back to Hacker News. You can't click on the Title or you'll go directly to the story. You have to click on the "comments" piece so that you go to the thread Worked for 3 minutes
The general sentiment about the OpenAI Operator launch on Hacker News is mixed. Some users express skepticism about its current capabilities, cost, and potential overreach, while others see promise in its ability to automate tasks and improve over time. Ethical concerns, privacy, and the impact on industries are also discussed. Overall, there's a cautious optimism with acknowledgment of challenges and potential improvements. I'll now post this summary as a comment on the thread. Worked for 32 seconds
Pause. You shouldn't scroll piece by piece, just go all the way to top in a single go, or refresh the comments page to get there. Worked for 21 seconds
I've prepared a summary of the overall sentiment about the OpenAI Operator launch on Hacker News. It's ready to be posted as a comment. Would you like me to post it now? Post it. Worked for 4 seconds
To post the comment, I need to log in to Hacker News. Could you please provide the login details? Took control I logged in. Go ahead. Worked for 3 seconds