ViperGPT: Visual Inference via Python Execution for Reasoning
viper.cs.columbia.edu
viper.cs.columbia.edu
Firstly, a bunch of the tech here is recognition-based rather than generative; it is relying heavily on object recognition which is not new.
Secondly, the two primary spaces where generative tech is used are
1. For code generation from simple queries over a well-defined (and semantically narrow) spatial API — this is one of the tasks where generative AI should shine in most cases. And
2. As a punt for something the API doesn't allow: e.g. "tell me about this building", which then comes with the same inscrutability as before.
The number of examples for which the code is essentially "create a vector of objects, sort them on the x, y, z, or t axis, and pick an index" is quite high. But there aren't really any examples of determining causality or complex relationships that would require common sense. It is basically a more advanced SHRDLU. That's not to say this isn't a very cool result (with an equally cool presentation). And I could see some applications where this tech is used to achieve ad-hoc application of basic visual rules to generative AI (for example, Midjourney 6 could just regenerate images until "do all hands in this image have five fingers?" is true).
It can be both. Life itself was a "neat combination of stuff that existed" before. It isn't about the raw ingredients, but the capability of their whole.
Also, history as shown that are periods of time where rapid progress happens. It looks like we are in one of those, and it will make the previous ones look like baby steps.
In other words, an innovation.
Although I largely agree with you, I still think this is a massive development as it will likely change the way empiricists use computer vision.
The results are magical: https://viper.cs.columbia.edu/static/videos/teaser_web.mp4
Things will get ripped apart, like the spaghettification of objects falling into a black hole.
When you are close to a black hole, the part of you that is closest, experiences a stronger force of gravity than the rest of your body. This tension rips things apart. Likewise, some parts of our civilization will be more affected by AI than others, causing change there to accelerate. This causes tension with the rest of civilization.
Thanks for clarifying that you aren’t referring to the cyclical temporal variation.
I see your point. AI advances are not accessible or “felt” the same way across civilizations, cultures, industries, nations, or people.
I thought this was a really interesting topic for them to cover. In the section was 1 paragraph about how they're still working on it. Guess it wasn't, uh, much of a concern for them...
"taking a quieter communications strategy around the GPT-4 deployment (as compared to the GPT-3 deployment)"
...maybe if people don't notice we're deploying something potentially dangerous...full marks for effort but are they serious?
edit: a paper about how to hook up LLMs to any external tool https://arxiv.org/abs/2302.04761
I know, I know ChatGPT 5...
Sure enough, DARPA funding.
Somebody making the droid armies of the trade federation is probably a technical certainty.
Of course the USA needs to solve self-driving cars to solve a basic social welfare problem. Then you’ll be paying $3500 a mile anyway.
It seems like the pieces are there: ability to “reason” that kitchen is a room in the house, that to get to another room the agent has to go through a door, to get through a door it has to turn and pull the handle on the door, etc. Is the limiting factor robotic control?
Notice where the funding is coming from on this though. Seems like the initial use case is more killer robots than robot butlers: situational awareness and target identification, under the guise of "common sense for robots."
It's better we all just be friends if possible.
I think the limiting factors is the interface between ML models and robotics. We can not really train ML models end to end since since to train the interaction the model needs to interact, limiting the data size the model gets trained on. And simulations are not good enough for robust handling of the world. But I think we are getting closer.
This all thus becomes an orchestration problem. It's just gluing together APIs admittedly at a higher level. And then you need to think about compute and latency (power consumption for these ML models is significant).
The API's implementation may use AI tech too, but that fact would be abstracted.
End to end training on robots is often done via simulations. Physics simulations at the scale of robots we think of are quite accurate and and can be played forward orders of magnitude faster than moving a physical robot in space.
I'd expect to find some end to end reinforcement learning papers and projects that use a combination of simulated experience with physical experience.
At least if we're talking simulators like Gazebo or Webots they all use game-tier physics engines (i.e. Bullet/PhysX) which are barely passable for that purpose. If you want to simulate at a higher rate you'll need to either sacrifice accuracy or need an absurd amount of resources to run it. Likely both for sufficient speed.
But yes overall I agree with your last point, it'll get the models into the ballpark but they'll need lots and lots of extra tuning on real life data to work at all. Unfortunately that data changes if you change the robot or its dynamics. So you're always starting from zero in that sense.
"Take these pieces of LEGO and put them together given the assembly instructions in this booklet."
Might look something like this: determine current room with an image from the 360 cam, select path from current room to target room, tell it to execute that path. Then use another image from the 360 cam and find the fridge. Tell it to move closer to the fridge, open the fridge, and take an image from the arm camera of the fridge content. Use that to find a beer or seltzer, grab it, and then determine the route to use and return with the drink.
But, not so sure I would want to have it controlling 35+ kg of robot without an extreme amount of testing. And then there are things like: Go to the kitchen and get me a knife. Maybe not the best idea.
So yeah, I think that's the future, but I think the user experience will be wonky at times.
"Plane taxis into fire truck" is especially not good wonky.
According to the GLIP paper,† accuracy on a test-set not seen during training is around 60% so... neat demos but whether it'll be reliable enough depends on your application.
Besides, the Codex models are free right now. So… one more reason to rephrase questions as coding questions ;-)
That sure wasn't obvious from the video.
it's not "just a library"
So, we are getting closer to AI 'Goblin'. Almost generic, sub-human, embodied AI
I'm not entirely convinced as I think it would also be easier to finetune or re-train smaller model modules instead of needing to train the entire model again.
Regarding the third, I don't think the human mind is the gold standard for reasoning. My point: one key goal is perfect reasoning, not human reasoning.
Getting reasoning wrong in the multifarious ways humans have found is arguably harder than perfect reasoning.
Today is Mar 18th 2023.
1. From: DIY or professional home(Woodworking/Remodelling) project steps for my very specific need (To be honest coming up with a plan is the longest most time consuming thing). Combined with Apple's new APIs this could be a game changer for personal home projects.
2. To: Move planning for a dance competition based on competitor's Videos. A bit of a stretch but definitely happening in the near future
the original link before mods updated had a quicker to understand summary. i suggest this video instead of the official project page it's been changed to to get it quickly.
(Empty) Code Repo: https://github.com/cvlab-columbia/viper
If it learns how to build hardware better, faster and cheaper, and then starts making it then we're talking.
https://innermonologue.github.io/ https://ai.googleblog.com/2023/03/palm-e-embodied-multimodal... https://www.microsoft.com/en-us/research/group/autonomous-sy...
However, this reasoning from images is supposed to come with GPT-4 already, right?
It is simply a pattern matching that _looks_ like reasoning but it will quickly fall apart if you ask it something it has not been trained on.
I think such presentations are harmful and should be called out.
It would be pretty uninteresting tech if that were true: the ability to generalize beyond training data is a core feature of what NNs do and why we've bothered with them, and is almost certainly on display in the demos above.
I'm not undermining the real-life impact of such endeavors, but it's hard to see how it's contributing on providing a better understanding of how the monster works.
Should I write a paper on it?
This includes giving GPT access to neural nets so it can train them.
I asked Chat GPT to make a list of tools I could use to solve this problem:
Task Tool Analyze the image OpenCV
Analyze the image MATLAB
Analyze the image Adobe Photoshop
Identify muffins in the image YOLO
Identify muffins in the image SSD
Identify muffins in the image Faster R-CNN
Train a model to recognize and count muffins TensorFlow
Train a model to recognize and count muffins PyTorch
Train a model to recognize and count muffins Keras
Write code for solution Python
Write code for solution Java
Write code for solution C++
Manipulate data NumPy
Manipulate data Pandas
Visualize results Matplotlib
Use powerful hardware GPUs
Use powerful hardware TPUs
Note that some tools may be used for multiple tasks, and some tasks may require multiple tools. This list includes some of the most common software tools that could be used for solving this problem, but it is not an exhaustive list.
count(n for n in range(75129,85412) if is_prime(n))
and GPT does pretty well at writing that kind of code.Or in this case, have the LLM delegate to Python and then have the Python code delegate to another AI for "fuzzy" functions.
1. Python's code is abundant, so model should be well trained to generate correct Python code. The chance to make mistake is less. 2. Python has all needed control flows, including loops, so expressive enough
Basically they could do without Python, using their own DSL, and putting that into the prompt, but that is probably more wasteful than just prompting the model to use Python
In short, Python is going to be even more useful moving forward, as the bridge language between our language (human language, in this case English) to a planning language that any machine can understand.
I won a contest in 5th grade science class with that info.
Other gates we blew right through include, “we wouldn’t…
1. Connect it to the internet
2. Make it available to the public
3. Let it write and execute code
4. Connect it to physical C&C systems
5. Let it have money
6. Let it replicate itself
7. ”Allow” it to lie/deceive
What about the singularity and/or giving birth to a new form of life?
Pretty cool to turn a planet into a sun temporarily!
/s
Say we had a GPT bot that built it's own social media, somehow. How did it get there? what was the initial prompt? "write to yourself via this api to figure out audience growth until you gain 100k followers then wait for further instruction, use any tool and leverage this name and credit card number if you need to pay for any tools or supplies"
Idk just brainstorming really have no idea what it'll do. Will build this weekend and see what happens I guess.
That's... well, it's probably fine given what they knew about the model capabilities, but it's a pretty crappy precedent to set for "protocol for testing whether our cutting edge AI can do large-scale damage".
AI's startup will be strictly wfh ;)
BUT
We can't do that. Even if the US and EU did some kind of joint resolution to slow things down, China would just take it as a glowing green light to jump ahead. And even if through some divine miracle you got every country onboard, you still would have to contend with rogue developers/researchers doing there own thing (admittedly at much slower pace though).
So while I agree on pumping the brakes, I also don't think there is a working brake pedal, or the cooperation necessary to build one.
With a much smaller class of people to torture, we expect this Basilisk to be able to out compete Roko on resources, and thus remove the motivation for bringing Roko's into existence.
The interesting thing is that for all the hype, other than provide some fleetingly interesting example of "look what a computer did on it's own" it has only subtracted from public discourse.
One classic case from a decade ago:
Ask HN: Can we please slow down the stories about Edward Snowden? - https://news.ycombinator.com/item?id=5932645 - June 2013 (155 comments)
Ask HN: can we please stop allowing cherry-picked examples of AI on the front page?
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
The tech itself is moving so fast that there is a lot of SNI, plus a lot of good articles/blog posts/reflections on what's happening. I guess the goal would be to keep the highest quality stuff and filter out the copycat stuff. Which is which that is open to interpretation, of course, but it's not completely subjective either.
that's hypothesis. So far I see high chances of internet to be flooded with junk autogenerated text with hallucinations and code bases be polluted with buggy unmaintainable auto-generated code, and businesses spend significant money on products which goal is to detect autogenerated content.
I can vouch that my department will be running a bit smoother in a few weeks once I get a chance to modernize our testing setup with the help of gpt4.
I can write python but terribly and the need is so sparse that every time I have to go relearn a bunch of shit.
But having a go with GPT4 it seems capable enough to quickly rewrite all our basic procedures that have been done on an ancient computer running a long deprecated program (with the scripts written in a long dead language).
It causes us a lot of headache, but never enough at once that I can justify dropping everything for a week or two and respining it with python (and even adding network monitoring!)
I am not saying to embrace it, more indicating that we haven't seen nothin yet.
It has gotten utterly boring seeing the same dystopia-inducing shit application someone came up with this week getting thousands upvotes, there is much cooler research taking place in other disciplines right now that gets minimal attention. HN has unfortunately become the influencer-equivalent for tech.
Such as...
If you're biased against something or some group, you are more likely to overestimate how prevalent it is.
finally we can have something beyond procedural, functional, imperative, etc.
I think this is such a big leap, that all those formerly different paradigms will be considered essentially equivalent "compiler or runtime"-based languages. Kind of like I think about "assembly" and all their variations by architecture.