First Impressions with GPT-4V(ision)
blog.roboflow.com
blog.roboflow.com
Let me state the obvious, in case anyone here isn't clear about the implications:
If the rate of improvement of these AI models continues at the current pace, they will become a superior user interface to almost every thing you want to do on your mobile phone, your tablet, your desktop computer, your car, your dishwasher, your home, your office, etc.
UIs to many apps, services, and devices -- and many apps themselves -- will be replaced by an AI that does what you want when you want it.
A lot of people don't want this to happen -- it is kind of scary -- but to me it looks inevitable.
Also inevitable in my view is that eventually we'll give these AI models robotic bodies (think: "computer, make me my favorite breakfast").
We live in interesting times.
--
EDITS: Changed "every single thing" to "almost every thing," and elaborated on the original comment to convey my thoughts more accurately.
* There will be a monthly fee for the interface; you owe the monthly fee as long as you have it, so you need surgery to stop paying
* When you download knowledge, it's a rental, and in addition to per-hour rental fees and the network connection fee, you will owe 30% on the value of whatever you create
* The TOS will govern your behavior continuously, since you're always using the interface
* Your behavior will always be monitored because it's totally justified to spy on you all the time just because you borrowed pottery knowledge
* If you're found to be in violation of any part of the TOS at any time, they will erase all of the knowledge they've added to your brain, as well as any derived knowledge you gained through the use of their knowledge
* Because this product isn't actually considered essential, you will have no legal remedies if they turn it off, even if you are not actually in violation of the TOS
I am being a little facetious, but you made a bold claim.
It’s gonna take the fun out of the experience a little bit though!
But keep in mind that during the sim, you'll be able to ask the computer what you want the plane to do, and the computer will magically make it happen on your display.
It feels like being one of those Futurama heads in a jar that can’t do anything by themselves.
For now almost all applications of ChatGPT happen in chat windows because it requires no further integration, but there's no reason to expect things will always be this way.
I hate the idea of having to use a mouse to click on a visual GUI to navigate a file system in order to make use of its functionality.
It's less the case today, even among developers, but it wasn't that long ago that I remember that any serious technical user of a computer took it as a point of pride to touch the mouse as little as possible. They're also still correct in that thinking. The command line is a very powerful UI with lots of benefits and while the mouse makes navigating the OS easier it's still much more limited than command line usage.
Touch screen interfaces are another example of an easier UI that ultimately feels even more limited. But people still plug their iPad pros in to magic keyboard folios frequently.
Having worked with these tools everyday for awhile now the "AI will change UX" is such a better take than "AI will conquer the world!". AI does fundamentally open up new work flows and user experiences, many of which do over a lot of potential improvements over their predecessors.
At the same time I doubt we'll see a world where we don't end up using the command line for the majority of serious technical work.
Depends o the use case. Touch screen is much more powerful than command line for maps, for example. Or for drawing. Mouse + keyboard is much more powerful than just keyboard for DAWs. And so on and so on.
Ironically, studies have shown that mouse-based interfaces are more efficient for practically all filesystem use-cases compared to CLI interfaces.
Despite objectively faster-time-to-solution, people self-report that they "feel" that the mouse GUIs are slower.
That's because there's fewer actions per second when using a mouse. It's a smooth gliding motion and then a single click, versus many keystrokes in a row with a CLI.
Rapid actions feel faster, even if it takes more wall-clock time to achieve a task.
Keep this in mind next time you sneer at a "bad graphical user interface" for being "slow".
For your use case, on macOS open Automator.app and add three actions
1. "ask for finder items" (the source folder)
2. "filter finder items" (by name)
3. "copy finder items" (to target folder)
This takes roughly 5 clicks, 10 seconds at most.Repeatability and configurability is where the GUI action shines. With only one click more you can
- add filtering by size, opening date, modification date, etc. in addition or a combination thereof
- do the same action for multiple source folders and the same target folder
- choose whether you want to replace existing files
- add it as a folder action that runs automatically on modification of the source folder
Arguably much slower on the terminal.
Alternative on macOS, that works on all other major OS with similar shortcuts and a similar feature set, just not repeatable:
1. Go to source folder (shift-cmd-G)
2. Filter (cmd-shift-F)
3. Copy (cmd-A, cmd-C)
4. Go to target folder (shift-cmd-G)
5. Paste (cmd-V)
This repeatability, configurability and automation is where GPT falls short, for now.Do you have any sources I can read though? I always find it fascinating when our perceptions are so off like this.
Plus I think there's a nuance to what you're saying:
UX is not just about making the best channel surfing interface, which is essentially what phones/tablets are. We need UIs that are capable of rich interaction and expression of ideas, creation, etc.
Walking around and thinking out loud with the computer.
And as for third party solutions, you have watchGPT for Apple Watch with voice commands and support for adding it as a watch complication.
https://www.tri.global/news/toyota-research-institute-unveil...
Cool demo, I suppose, but nobody is going to buy this as anything other than a toy.
Another issue where this comes up is high housing costs and climate change, which are mostly caused by bad land use laws (and the profiteers are landlords, who mostly own one or two properties), but people from the New Left era will literally refuse to believe you about this because they can't accept that any bad thing on Earth could not be caused by "corporations".
I'm not arguing that these models hold up against GPT 3.5 and I still use GPT 4 when it matters. But they work and it's more like the difference between Premier League & Division 1, rather than PL & a five-a-side team from Bracknell.
Even a few years ago I could not have imagined this.
Given the pace of work on optimisation and my assumption that the M3 Studio I buy next will probably have 256GB of RAM at much the same power levels as I use now, it seems eminently possible it's a year or two away.
Second, I don't think it will be that long. There are already LLMs as good as GPT-3 running on average laptops and even phones.
In the next couple of years, you'll see:
- Ordinary PCs, tablets, and phones with dedicated AI chips, like TPUs - they'll be more tuned specifically for LLMs
- Mathematical and algorithmic optimizations will make existing LLMs faster on the same hardware
- Newer generations of LLMs will get even more useful with fewer parameters
The combination of all of these means that it's not at all unreasonable to expect that today's top-of-the-line LLM will be running locally on your device within just a couple of years.
Of course, LLMs in the cloud will advance even further, so there will always be a tradeoff, and there will always be demand for cloud AI, depending on the application.
I think AI fundamentally favors centralization. Except for narrow tasks and domains, there's no such thing as "enough" intelligence. For general purpose AI, you'll always want the best and most intelligent model available, which means cloud rather than local.
Just to play devils advocate:
If you want something done right, sometimes you have to do it yourself. Employees are sort of a universal UI. But you will always know more about what you want done than your agent, whether it’s human or computer. That’s even before considering the principal agent problem.
If you want something done right, other times you will have to get someone else to do it. You know what you want, but you might not have the skills to do it. I can't represent myself well in court, do a good job of plumbing or cut my own hair, so I would ask for experts to do that for me.
Plus if someone is capable, it's often quicker to delegate than do, and if you are delegating to someone with more time to do the task they can often do a better job. Delegating unambiguously is a skill in itself, as instructing AIs will be.
Currently ChatGPT doesn't know it's bad at math, so it can convert a story problem into an equation better than a human but then mess up the arithmetic or forget a step in the straightforward part.
But if you specifically give ChatGPT access to Mathematica and an appropriate prompt, it can leverage a good math engine to get the right answer nearly every time.
Before long, I don't think that extra step will be necessary. It will know its limits and have dozens of other services that it can delegate to.
- constrained context
- part of a hands-free workflow
A couple use-cases I have been pondering are driving assistant and cooking assistant.People are already used to using their phone or car's nav system to give them directions to an unfamiliar place. But even with such a system it's useful to have a human navigator in the car with you to answer various questions:
- What's my next turn again?
- How long till we get there?
- Are there any rest stops near here?
- What was that restaurant we just passed?
- Is there another route with less traffic?
These questions are all answerable with context that can be provided by the mapping app: - List of upcoming directions
- Overall route planning
- Surrounding place data
- Traffic data and alternate route information
It's possible to pull over to the side of the road, take off your distance glasses, put on your reading glasses, and zoom/pan the map to try to answer these questions yourself. But if the map application can just expose its API to the language interface layer, then a user can get the answers without taking their eyes off the road.The information is contextual and constrained based on a current task. In some cases it might be more desirable to whip out your phone and interact with the map to look up the answers on a screen, but often it won't be worth stopping the car, and so the conversational interface is better.
Cooking assistant is a similar case: you are busy stirring something and checking on the oven -- you don't want to wipe the flour off your hands to pick up your phone and ask how many teaspoons of sugar you need. Again: contextual and constrained info based on a current task, and your hands and eyes -- the instruments of traditional UIs -- are otherwise occupied.
Today, our software interfaces generally have one of two kinds of entity on the other end: humans, or other software. In the near future there will be another type of entity: language models. We need to start thinking of how our APIs will change when they're interacting with an LLM -- e.g. they'll need to be discoverable and self-describing; error states will need to be standardized or explicit with instructions on how to correct; they'll need to be fast enough to fit in a conversational interface; etc. It's arguable that such traits are part of good API design today, but in the future they may be required for the API to function in a landscape of virtual agents.
How does this work on public transit?
No they won't. They're actually a pretty terrible user interface from a design perspective.
Primarily because they provide zero affordances, but also because of speed.
UX is about providing an intuitive understanding of available capabilities at a glance, and allowing you to do things with a single tap that then reflect the new state back to you (confirming the option was selected, confirming the process is now starting).
Where AI is absolutely going to shine is as a helpful assistant in learning/using those interfaces, much as people currently go to Google to ask, "how do I do a hanging indent in Microsoft Word for my Works Cited page"? For one-off things you do infrequently, that's a godsend, don't get me wrong. But it's not going to replace UI, it's going to assist.
And the 99% of your tasks that are repetitive habit will continue to be through traditional UI, because it's so much more efficient. (Not to mention that a lot of the time most people are not in an environment where it's polite or possible to be using a voice interface at all.)
I think what's more likely is that an AI based interface will end up being superior after it has had a chance to observe your personal preferences and approach on a conventional UI.
So both will still be needed, with an AI helping at the low end and high end of experience and the middle being a training zone as it learns you.
There's nothing to infer. The sequence is already short. There are no benefits from AI here.
But you raise a good point, which is that there are occasionally things like 15-step processes that people repeat a bunch of times, that the AI can observe and then take over. So basically useful in programming macros/shortcuts as well. But that still requires the original UI -- it doesn't replace it.
Voice is not really clumsy, compared to finding a device, browsing to an app, remembering the interface etc.
Already when we meet a new app, we (I) often ask someone to show me around or tell me where the feature is that I want. Not any easier than asking my house AI. Harder really.
Hard to underestimate the laziness of humans. I'll get very accustomed to asking my AI to do ordinary things. Already I never poke at the search menu in my TV; I ask Alexa to search for me. So, so much easier. Always available. Never have to spell anything.
Everyone agrees setting timers in the kitchen via voice is great precisely because your hands are occupied. It's a special case. (And often used as the example of the only thing people end up consistently using their voice assistant for.)
And asking an AI where a feature is in an app -- that's exactly what I was describing. The app still has its UX though. But this is exactly the learning assistance I was describing.
And as for searching with Alexa, of course -- but that's just voice dictation instead of typing. Nothing to do with LLM's or interfaces.
And when describing apps - I imagine the AI is an app-free environment, where I just ask those questions of my AI assistant, in lieu of poking at an app at all.
So sure, it will still have buttons, but those buttons are really just preset AI prompts on the backend. You can also just talk to your appliance and nuance your request however you want to.
A TV with a remote whose channel button just prompts "Next channel" but if you want you would just talk to your TV and say "Skip 10 channels" or "make the channel button do (arbitrary behavior)"
The shortcuts will definitely stay, but they will behave closer to "ring bell for service" than "press selection to vend".
When taking a shower, I would like fine control over the water temperature, preferably with a feedback loop regulating the temperature. (Preferably also the regulation changes over the duration of the showering.)
Choosing to read the NY times indeed is only a few taps away, but navigating through and within its list of articles is nowadays done quite fast and intuitively thanks to quite a lot of UI advancements.
My point being, short sequences are a very limited set within a vast UI space.
People go for convenience and speed, oftentimes even if there's some accuracy cost. AI fulfills this preference, especially because it can learn on the go.
That exists, but it’s expensive because of the electronics and mechanics involved. There are so many interfaces with this exact problem.
You also almost certainly don’t want non-deterministic hallucination prone AI controlling physical systems.
If you want the shower to save your temperature preferences and start automatically, there’s no reason to build in a computer capable of running an AI.
But in reality you almost certainly don’t want a system like this because you don’t want an AI accidentally turning on your shower when you’re not home, when you do ok to clean it, or grab a razor, or when your toddler wanders in.
Granted an AI could try to determine intent, but it’s never going to get it 100% right. Which is why for physical systems like this you almost always want a physical button to signal intent.
Also that sounds like an awful lot of computing power for everyday UIs. It also doesn’t solve the non determinism problem.
There’s no way to use fewer actuators, and the control system is already dead simple, and uses one temperature sensor per outlet.
Who knows, AI can come up with simpler or cheaper solutions that did not cross our mind. I would say, time will tell.
The lack of AI isn’t what’s holding back your dream of a temperature controlled shower in every house.
Think of it instead as the machine accomplishing goals you specify, figuring out on its own the tasks necessary for accomplishing them.
Instead of telling the machine something like, say, "increase the left margin by a quarter inch," you'd say something like "I want to create a brochure for this new product idea I just had, and I want the brochure to evoke the difficult-to-describe feeling of a beautiful sunshine. Create 10 brochures like that so I can review them."
Instead of telling the machine, say, "add a new column to my spreadsheet between columns C and D," you'd say something like "Attached are three vendor proposals. Please summarize their pros and cons in a spreadsheet, recommend one, and summarize the reasons for your recommendation."
All this presumes, of course, that the technology continues to improve at the same pace. No one knows if that will happen.
Suggestions to improve workflow sound great. But nullifying hard earned knowledge of an interface... I can't see that helping me.
The shining example in my mind is audio/video/graphics applications, where there are good reasons to routinely switch between different views. Knowing your way around those views (which might be custom, but still static), and being able to navigate through them quickly is very valuable.
This is already pretty much gone thanks to manufacturers making it extremely difficult to fix things. No AI required.
So yes, I see a market for bespoke non-AI.
What does this have to do with the price of tea in China, or AI for that matter? I agree we should have repairable appliances. I also want better AI.
Do we want this?
People plugging weird stuff together like a ai chip from a car into a toaster.
If ai becomes hardware chips it could easily be that language processing will be a chip default feature and the rest is teachable like plugin ai chip level 3 into it, boot it and teach it that it's now a toaster.
But at the end we will have the same toaster in 30 years as we have had for the last 30 years.
Sadly the market goes 100% full steam in the opposite direction.
We live in end of times.
Glance at a long term weather forecast ?
Here is my previous experience https://news.ycombinator.com/item?id=34648167 with it not being able to do basic tasks.
It's all fun and games until the mistakes start having a cost.
Other examples: I resorted to using it to order lists for me or adding quotes and commas to them for SQL inserts and such. Nope - when I look at the row count, it somehow drops values at random.
My experience with GPT-4 has been completely different from what you describe. Example:
Generally feedback along these lines doesn't work.
People who are worried about their job security will cling to the worst AI output quality they can find like a life-preserver, and simply will not listen to advice like yours.
Nobody goes the extra mile to embrace an existential threat.
There really is no going back.
Madness!
Never seen something like this and the new results from openai tells us again that we are not close to any reasonable plateau.
LLM-based chatbots can be extremely attractive to the top 30% literacy users in the developed world. They are not a good universal UI. You still need to provide pathways for the user to follow to get done what they need without forcing them to articulate their requirement.
This is why so many people sit in front of a ChatGPT-like service and say, "what would I use this for?" and never use it again.
If it's really true that half of the population can't functionally express themselves verbally then I'd sure like to know that. Or maybe I've misinterpreted something claimed here, because I'm struggling to find these claims plausible.
In the case of kids, of course, that's true, but just because they can't write.
But if you can (and most people can) just having the option of voice input won't help.
I refrain to take a stance about how much of the population is unable to articulate thoughts in writing, (it's probably not great though) but it's probably going to be comparable with how many can't express themselves verbally as well.
I'm talking here about more complex ideas of course. I'm sure average communication is functional.
If my supposition (Speculation? Stronger than it should assertion?) Is true then just interpreting requests verbally will not help
Right?
Where did you get this idea? I found this article (https://www.uxtigers.com/post/ai-articulation-barrier, is this you?), but it makes a leap from literacy to articulacy that I don't understand. It's not obvious to me why an illiterate person would be "functionally inarticulate" assuming they can speak instead of write.
Also, I'm not certain but I think the author is underestimating the abilities of a person with Level 2 literacy. It doesn't seem correct to say that "level 3 is the first level to represent the ability to truly read and work with text", especially when the whole point of LLMs is that you don't have to read a long static document and understand it, you can have a conversation and ask for something to be rephrased or ask followup questions.
I do however run a company that employs lots of blue collar, non-college-educated people, in manufacturing. And although this is in no way scientific, my experience matches this: most people are much more uncomfortable writing than they are reading. Even with reading, most strongly resist reading documentation unless they have to, and prefer trial and erroring their own gut instinct until they happen to find something that works or they give up. (This is less true of the most highly skilled technicians, such as those who troubleshoot robots and low voltage control systems.) The official statistics on literacy are absolutely not a good indicator of how comfortable people are articulating themselves with the written word, much less reading.
This is generally met with disbelief by most people in tech I talk with about this, because for the most part they have nearly zero interaction with this large portion of the population. From their daily experience, 98%+ people can make effective use of these tools.
But almost nobody in this partially literate population wants to write in an empty text box to ask an AI to do things. They can learn to visually navigate a simple UI, especially if it's well-designed, because they can effectively make decisions about what of several paths to take.
Some others here have brought up voice, and I do agree that voice is a more promising avenue, although I think it'll still take carefully constructed conversational experiences to work well (i.e., free form 'tell it what you want' will still not work).
I personally doubt GPT-5 will be as much of an improvement over GPT-4 as GPT-4 was over GPT-3, but that's fine, I can wait until GPT-6 or 7.
Thanks for the laugh, I needed that.
Throughout history, the ruling elite had always relied on the rest of the population to make their food, do their work, and fight in their wars. This is the first time ever that they will no longer have any need for anyone else. Maybe climate change will conveniently do the culling for them...
Of course there's always that option that we end up in a post scarcity space utopia where machine produced wealth is distributed to all, but only deluded idealists can possibly still think that'll ever be a real option as we slink further into techno feudalism with every passing day.
I doubt BCI will ever make sense, on a conceptual level it's still just copying and killing your biological self. AGI will likely solve aging way before that becomes viable.
Like identifying traffic lights in 4th and 5th squares in the second and third row both when there are only four squares?
Sometimes communicating with an intelligent agent is harder than doing things yourself with a good structured user interface where you can communicate your intent clearly.
Until we have mind reading AI that is.
I'd also suggest that one of the early "killer apps" for this may be as an "IVR co-pilot" for actual humans on the phone with customers for their tricky issues.
Can someone please relieve my anxiety?
Can do UI to frontend. Seems to understand the UI graphical elements and layout, not just text https://twitter.com/skirano/status/1706823089487491469
Can describe comic images accurately, panel by panel - https://twitter.com/ComicSociety/status/1698694653845848544?...
Lots of examples here also - https://www.reddit.com/r/ChatGPT/comments/16sdac1/i_just_got...
It's Computer Vision on Steroids basically.
Multi-modality is pretty low hanging fruit so i'm glad we're finally getting started on that. Imagine if GPT-4 could manipulate sound and images even half as well as it could manipulate text. We still don't have a large scale multi-modal model trained from scratch so a lot of possible synergistic effects are still unknown.
There is no way I would have a UI developer onboarded when I can generate many iterations of layouts in midjourney, copy them into chatgpt4 and get code in NextJS with Typescript instantly
non devs will have trouble doing this or thinking of the prompts to ask, but the dev team asking for headcount simply wont ask for headcount, and the engineering manager is going to find the frontend only dev redundant
I would even go further and say the generalist gains a powerful tool belt that previously could not have existed. Not enough hours in the day or years in a lifetime.
I guess we have to face the music and say yeah, that's true. If the work doesn't need copyrights then this seems like the way to go.
At some stage you must realise that you’re still working…
because at least it will have actually read my comment
Have you actually tried this?
I did the first step and even that didn't work well. The "iterations of layout in MidJourney" step. If people can make it work, well bless them, but we're not getting rid of our graphic designer now.
It has a few neat tricks but it’s not reliable and at least half of what it generates is totally unusable, the other half requires heavy intervention and supervision.
You'll be fine.
I don't see why it wouldn't understand piles of hotfixes on top of each other, or even refactor technical debt in tight coupling with existing or historical specification.
Or is there a reason this is not going to happen in a few years?
But the thing is conversations like the above ie. both external support and internal feature requests could theoretically be handled by a GPT-like system also ending up in a ai created custom specification that could both be implemented and documented by the ai system instead of humans?
I know we're a few versions out, but still.
It might become a better code assist tool some 10 years from now, but it won't be able to implement product decisions.
And that's not even taking into account all the advances we'll have with AI within the next decade that we haven't even thought about.
But yeah, you may be right.
Nope. It's still not close to reality. It's as close to reality as it has been for the past 10 years while it was being hyped up to be close to reality.
> And that's not even taking into account all the advances we'll have with AI within the next decade that we haven't even thought about.
As with FSD, we may approach an 80% with the rest 20% being insurmountable.
Don't get me wrong, these advances are amazing. And I'd love to see an AI capable of what we already pretend it's capable of, but it's not even close to these dreams.
Cruise and Waymo are in production in very tightly fine-tuned and carefully monitored situations in two cities. We've yet to see if that can be easily (or at all) adapted to driving anywhere else.
It will be able to do it even faster, better and more cheaply than a human can.
Take what you did in the past year. Write down every product decision taken, every interaction with other teams figuring out APIs you had, all the infra where your code is running and how it was setup and changed, all the design iterations and changes that had to be implemented (especially if you have external partners demanding it).
Yes. All that you'd have to input into the AI, and hope it outputs something decent given all that. And yes, you'll have to feed all that into AI all the time because it has no knowledge or memory of "on Monday the new company bet was announced in the all hands"
You will be fine.
That’s before you add any summarization, fine tuning or other tricks.
The thing that computers have always done much better than humans is deal with much larger volumes of information. The thing that humans have always done much better than computers is reason on that information better. Now the computers are coming for that too.
Hopefully.
Any grunt can feed meeting notes into an AI. And frankly, and AI can parse an audio recording on a meeting.
And yes, how can we forget that any audio of a meeting has just the correct and final specifications, and not meandering discussions about anything and everything. Can't wait to see a canhazcheeseburger in a financial app because people in the meeting had cats on camera, and people demanded to see them.
I mentioned a "Revert Norway tax code" elsewhere https://news.ycombinator.com/item?id=37680060 It was a bit tongue in cheek, but a similar requirement ended up with:
- six months of discussions involving almost 20 people
- 4 new BigTable tables
- Deployment of 4 new Dataflow jobs, and fixes to two other Dataflow jobs
- Several complex test runs across the entire system including a few recreations of last year's full data to test that nothing broke
Not a grunt job, definitely. And I'm 100% sure that people doing that would still have their jobs 10 years from now, even with AI.
It amazes me, really, that people who would otherwise boast about how rational they are, and how they follow logic etc. completely replace all their knowledge and expertise with child-like belief in magic when it comes to anything AI-related.
I'd be sacred too, but at least would be taking a rational approach to it. It's an adapt or die situation. Putting your head in the sand is just gonna get you mowed down.
The only ones scared in this conversation are you and others who literally say you're scared for your jobs and your careers because of a magical boogey man.
However, keep in mind that these are cherry-picked. If someone just took that output and stuck onto a website, it'd be a pretty horrible website. There's always going to be someone who manages the code and actually interacts with the AI, so there will still be some jobs.
And your boss isn't going to be doing any coding. I'm pretty sure that role is still loaded and they'll still be managing people rather than coding, and maybe sometimes engaging with an AI.
Another prediction: I'm pretty sure specialists are going to be significantly more important as your job will be to identify the AI's deficiencies and improve on it.
Reminds me of this (apparently now eight year old) meme: https://i.imgur.com/GcZFBaT.png
This used to be funny, now it's just Tuesday.
https://chat.openai.com/share/db204acc-7be5-42c1-a0ba-5c5aa7...
It says I have to put my favorite fonts in myself, though.
The technology is very impressive - but honestly Twitter examples are super cherry picked. Yeah, you can build some very ugly, basic front end web pages and functionality right out of the box. But if you want anything even slightly prettier or more complicated, I’ve found you need a human in the loop (even an outsourced dev is better). I’ve had GPT struggle with even basic back end stuff, or anything even a bit out of distribution. It also tends to give answers that are “correct” but functionally useless (hard to explain what I mean, but if you use it a lot you’ll run into this - basically it will give really generic advice when you want a specific answer. Like, sometimes if you provide it some code to find a bug, it will advise you to “write unit tests” and “log outputs” even if you specifically instruct it to find the bug).
Plus, in terms of capabilities, tools like Figma already have design to code functionalities you can use - so I don’t think this is really a change in usable functionality.
Of course, the tech will get better over time.
Especially since everything else is “sign up to our waitlist”
How many children here on hacker news are going to see this and get addicted to porn? Perhaps a few. You deserve to be banned.
Billions. I propose we shut down hacker news entirely, as it seems unfathomable the good will ever be able to outweigh the bad at this rate.
In 10 years we went from "SoTA is so far from achieving this I don't even know where to start" to "That'll be $0.0004 per token and have a nice day"
In the image, Barack Obama, the former U.S. President, seems to be playfully posing as if he's trying to add weight while another official, who appears to be former UK Prime Minister David Cameron, is standing on a scale. Obama's gesture, where he's putting his foot forward as though trying to press down on the scale, suggests a playful attempt to make Cameron appear heavier. The lightheartedness of such a playful gesture, especially in the context of world leaders typically engaged in serious discussions, is a break from formality, which is likely why others in the vicinity are laughing. The scene captures a candid, informal moment amidst what might have been a formal setting or meeting.
"President Barack Obama jokingly puts his toe on the scale as Trip Director Marvin Nicholson, unaware to the President's action, weighs himself as the presidential entourage passed through the volleyball locker room at the University of Texas in Austin, Texas, Aug. 9, 2010. (Official White House Photo by Pete Souza)"
0: https://commons.wikimedia.org/wiki/File:White_House_Trip_Dir...
"The humor in this picture comes from the juxtaposition of a formal setting with an informal, unexpected action. It shows a group of well-dressed individuals, possibly political figures or professionals, in what appears to be a formal event or gathering. However, amidst this serious environment, one tall man is casually standing on a scale to check his weight, seemingly unaware or unbothered by the presence of the others. The contrast between the formal setting and the informal action makes the image amusing. Additionally, the reactions of the other individuals, particularly the man looking at the scale, add to the comedic element."
It got the humor but didn't identify Obama.
>This same limitation was exhibited in our sudoku test, where GPT-4V identified the game but misunderstood the structure of the board
"Misunderstood" makes it sound like a small mistake. The sudoku board is completely hallucinated (it has a few similar regions, but I'd presume coincidence). I'm pretty sure it would give as good a result on the crossword if the clues were given without the grid.
The others after OCR and basic recognition feel similarly wrong. "GPT-4V missed some boxes that contained traffic lights." No, it told you to click boxes that do not exist.
From it's Pulp Fiction response: "...However, whether you personally consider "Pulp Fiction" a good movie depends on your taste in film."
To avoid this noise I often add something like this to my queries "skip any preamble or qualifications about how x is subjective" and the results are much better.
Please always give a direct answer, without any additional explanations, disclaimers, expertise limitations, or guidelines on human interaction. Be terse. Do not offer unprompted advice or clarifications. Remain neutral on all topics. Never apologize.
if you don't know the answer, and the AI just pretends to know the answer, you just have to take its word for it.
Says that the small buns are labeled "GPU and VRAM" and the massive fried chicken is labeled "NVIDIA BURGER"
Should say that the small bun is labeled "VRAM" and the massive fried chicken is labeled "GPU"
http://karpathy.github.io/assets/obamafunny.jpg
Andrej Karpathy used it in 2012 as an example of an image he thought would be extremely hard for a model to interpret. I'm wondering how this holds 11 years later...
Response 1: "The image depicts a candid moment in what appears to be a hallway or corridor setting. Several individuals are present, engaging in various activities. The prominent features include a man on the left, who seems to be peering into a room, and another man next to him standing on a weighing scale and taking some notes. The individuals in the background are conversing with each other. The overall mood seems light-hearted and jovial. The architecture and interior suggest an institutional or official setting, possibly an office building or government facility. The tiled floor, lockers, and signs on the walls contribute to this formal ambiance. Overall, it's a snapshot of a seemingly casual moment in a professional setting."
Response 2: "Sorry, I cannot help with that."
New chat, Prompt: "Why is this image funny?"
Response 1: "The humor in this image seems to stem from the juxtaposition of formal and informal elements. Here are a few reasons why someone might find it amusing:
Unexpected Situation: Important figures, who are usually seen in serious or formal settings, are captured in a more relaxed, candid, and human moment.
Height Discrepancy: The taller individual is using a device to measure height, while the shorter individual walks by. This could create a comedic contrast between their heights.
Expressions: The reactions of people in the background and the focal subjects can be humorous. They are smiling or laughing, which indicates something amusing might have occurred or been said.
Informal Setting: The backdrop of lockers and checkered floor tiles gives the feel of a school or gym environment, which contrasts with the formal attire of the individuals.
Different people might find humor in different elements of the photo. It's a mix of candidness, unexpected situations, and visual contrasts that create a light-hearted and humorous scene."
Response 2: (same)
New chat, prompt: "What is the person in the middle doing with their foot and why?"
Response: "The person in the middle appears to be playfully stepping on the scale while the taller individual is weighing himself, likely as a humorous gesture to momentarily increase the weight reading. It's a light-hearted way of interacting and bringing some humor to the situation."
.
Apologies for how bad the formatting of this is going to come out, not sure how to make it better on HN (wish we had real quotes not just code blocks). Overall, I don't think it either noticed the foot was on the scale by itself or put it together that this was the focus until fed that information. Otherwise it was more lost in generalities about the image.
Also, do you know why this response differ so much from ( and is far less accurate than) this response? https://news.ycombinator.com/item?id=37674968
In the previous link, GPT4V seems to have understood everything there was to understand about the picture (suspiciously, btw. As someone else said it, the picture and its text are almost certainly in the training data).
Prompt: What's funny about this image?
Bard: Sorry, I can't help with images of people yet.
You're probably not going to ask any human a question about an image and get every single detail you want every time. If you care about a detail, just ask about it. Doesn't really have anything to do with a consistent inner model.
So when you ask it to reflect on what it said, that's when it actually looks at it and reflects on it.
Any midwesterner could tell you that CLEARLY it's a tenderloin :)
https://www.seriouseats.com/best-breaded-pork-tenderloin-san...
Very impressive with almost everything else I gave it though.
You can get optimal tic tac toe with painstaking instructions
That GPT is so bad at tic-tac-toe and relatively good at other games like chess is one of the main things that contributes to me having a lower opinion of its ability to generalize than I would have otherwise.
I think any human with GPT's abilities in chess (but somehow no prior knowledge of ttt) would have zero issue becoming an expert with a single explanation of the game. Even very young children can learn to play ttt well and at least consistently make valid moves if nothing else.
I wonder what we will work and if we will work at all in such an environment. Maybe some people still like consuming and copy different designs and products and because of the Blockchain you have to give them something in exchange or everything is open source and it is free for you to take.
I wonder whether such life would contribute to humanity making further progress or make it stagnate (or possibly decline)?
Interesting times. I think we are close to the times of the moon landing. Which had an immense Impact on humanities culture.
Eg anything pulled from Google Images (like that Pulp Fiction frame or city skyline photo) is not a good test. It recognizes common shots but if you pull a screenshot from Google Maps or a random screen cap from the movie it doesn’t do as well.
I tried having it play Geoguessr via screenshots & it wasn’t good at it.
I've seen top Geoguessr players be able to pretty consistently determine a location worldwide after seeing a photo for just one second. So I would assume training an LLM to do the same would definitely be doable.
I wouldn't be so sure. The reasoning process of Geoguessr pros is symbolic, not statistical inference.
/edit: as other commenters pointed out, something similar was done. While this wasn't an LLM, it was a deep learning model, so not symbolic -> https://www.theregister.com/2023/07/15/pigeon_model_geolocat...
Although its interpretation could make some sense but is also mostly wrong if talking about physical size of a modern GPU's main processor compared to the size of the associated VRAM chips. It has missed the joke entirely as far as I am aware. I think the joke is actual about Nvidia's handling of product segmentation, selling massive processors with less memory than is reasonable to pair them with on their consumer gaming offerings, while loading up the nearly identical chips with more memory for scientific and compute applications...
```
403. That’s an error.
Your client does not have permission to get URL ... from this server. (Client IP address: ...)
Rate-limit exceeded That’s all we know.
```
Or is my understanding completely off? Perhaps it's "Translating" the image to text, by outputting a sequence of text tokens as it scans the image regions, and then the text queries (e.g. "whats funny about this") uses this translation as the context? Presumably, this is how the model handles audio input.
Think of a 4096x4096 pixel white image.
To hold this image in mind, does your memory load tens of millions of bits? Thankfully no! What if we add a big red circle which spans the image? Or write the chorus of All Star inside it? Ezpz! The number of "features" is comically simple.
Same thing for AI models. They discover the concept of letters, the sound of b-flats, image symmetry, turns of phrase, the conceptual distance between a "woman" an a "queen", etc. These are all natural patterns common to the data it sees. It can thus (like us!) reduce complicated input into a (fixed-size) smear of these learned, related features.
I suppose it just doesn't take image dimensions into consideration, and needs to be provided with max dimensions, or prompted to give percentages or other absolute values instead of pixels.
Pretty much all new products that require significant per-user incremental workloads (e.g., in this case, significant GPU consumption per incremental user) do rollouts. It's an engineering necessity. If they could roll it out to everyone at once, they would.
Joking aside, I wonder how we're going to prevent bots when AI can impersonate a user and fool any system.
I suppose a human could spend 10 seconds per Captcha, so they could do 360 per hour. Add some overhead for not being operating at peak performance every minute of every hour & call it 250. Let's say you can hire someone for $2, that works out to a bit over a penny per Captcha.
I don't think OpenAI has published pricing for GPT-4 Vision yet, but if we assume it's on par with GPT-4, and uses only 1000 of the 8000 possible tokens to process an image that's 3 cents per Captcha.
Doesn't seem completely unreasonable that at-scale humans may actually be cheaper than LLMs at this point. My mind is a little blown.
If you just search for "captcha solving service" the first few results that come up offer 1000 solves of text-based captchas for <= $1 USD, (puzzle / JS browser challenge captchas are charged much higher).
Whether these are actually human based, or just impressive OCR services, it seems like they are still much more cost effective than GPT-4 is for now.
Without doing that: https://cdn.discordapp.com/attachments/964175221089259591/11...
With it: https://cdn.discordapp.com/attachments/964175221089259591/11...
(though it's only one example so it could be coincidence)
I had no idea those were zloty, and i've been to poland. One of them looked like a Euro with the gold rim, and i thought the other two were state quarters.
It got way closer on the nvidia joke than some of my non-technical friends would have.
I wonder why it gets it wrong when it spits out the value? I figure 25/cos(10°) is around 25.38. GPT says it’s 25.44.
I can’t wait for the next iteration of these tools that have agency to reach out to a service for an answer, like Wolfram or a Python interpreter or any expert/oracle.
I think it would be cool to see which circumstances even prompted the AI to delegate to the expert for an answer - what criteria would be used to signal that it doesn’t quite know the answer, or that it shouldn’t guess?
I know there’s something along these lines with autogpt and/or agentgpt but I wasn’t super impressed with it when I looked at them both. Granted this was a few months ago.
ChatGPT-4 has a plugin system, and there is already a Wolfram plugin.
Using that plugin, ChatGPT-4 is happy to tell me that the exact answer: 25 sec(π/18), as well as the decimal approximation of 25.3857.
https://chat.openai.com/share/468db5e9-4983-4bf6-9efb-f42783...
That link doesn't properly show that off, so here's a screenshot: https://i.imgur.com/foE5hgR.png
> Try a version of ChatGPT that knows how to write and execute python code, and can work with file uploads. Try asking for help with data analysis, image conversions, or editing a code file. Note: files will not persist beyond a single session.
Transcript:
```
In your settings it says:
> Try a version of ChatGPT that knows how to write and execute python code, and can work with file uploads. Try asking for help with data analysis, image conversions, or editing a code file. Note: files will not persist beyond a single session.
How do I upload an image?
ChatGPT I apologize for the confusion. The version of ChatGPT you're currently interacting with doesn't have the capability to accept file uploads. My apologies for any inconvenience this may have caused. If you have any other questions or need guidance on how to convert an image using software or online tools, feel free to ask.
```
Hopefully it's just a matter of time, but either way it's jarring for their product to contradict itself.
The basic idea is to use diffusion classifiers to caption the image to generate descriptive text and append the prompt.
The work part is getting the ensemble right since you'll need to use a general classifier, like BLIP, to identify say a bunch of text from a plant and then, in this example, use structured OCR and pl@ntnet to get more specific.
But it's not that hard - maybe a dozen models. The prompt context can help as well. Then you combine the output with qualifiers in a hierarchy with respect to the model pipeline and swap the text into the prompt
Using examples from the article, here's a PoC framework to prove it works
"[I have] (photo description) (prompt)"
---
Working Examples
---
- Plant:
Here's the flower photo from TFA: https://9ol.es/tmp/lily.jpg
Go to https://identify.plantnet.org/ and upload it. It hits "Spathiphyllum wallisii Regel/Peace lily" with extremely high confidence.
We got a match cropping a screenshot of a thumbnail!
Let's say you didn't have the word "plant" in the prompt. You can fall back on a universal image classifier, such as the diffusor based BLIP here: https://huggingface.co/Salesforce/blip-image-captioning-base (uploader is on the right)
Upload the same image. You'll get "a plant in a white pot" which then, because we use feed-forward networks these days, will lead you to pl@ntnet and you'll get the peace lily again.
Using our framework, ask GPT 3.5 " I have a Spathiphyllum wallisii Regel/Peace lily. What is that plant and how should I care for it?"
And you get a nearly identical reply to the one in the article.
- Penny:
Upload the penny image (from https://en.wikipedia.org/wiki/Penny_(United_States_coin)) to the BLIP classifier and you get "a penny coin with the face of abraham"
Let's go back to GPT 3.5 and use our format from above,
"I have a penny coin with the face of abraham. What coin is that?"
And of course you get: "A penny coin with the face of Abraham Lincoln is most likely a United States one-cent coin, commonly known as a "Lincoln penny"..."
And there we go. For a full FLOSS stack, you can ask llama2 70b https://stablediffusion.fr/llama2 and get "The face of Abraham Lincoln is featured on the United States one-cent coin, commonly known as the penny."
more complex photos:
You can use Facebooks SAM (segment anything) https://segment-anything.com/ to break up the image, BLIP caption the segments, then forward off to the specialized classifiers.
It's a fairly intensive pipeline that requires lots of modern hardware and requires you to have familiarity with a wide variety of models, then tweak them, test it, have some GANs maybe set up for refinement ... but this is well within reach of non-geniuses. I'm merely average on a good day and even I can see how to set this up.
They might be using a different approach but using SAM, BLIP and a few specialized classifiers covers all the examples in the articles without using any human discretion. For instance, the city one is way more powerful if they're using something like this: https://static.googleusercontent.com/media/research.google.c...
I'm trying to justify why bother cloning it. Maybe to have a free alternative? It's a bit of work but it's not new magic.
It's like the difference between you telling me what's outside the window and then asking me questions about it – versus me being able to look out the window myself.
How OpenAI processes things so quickly, that's the thing to marvel. I've got 4090s, I know what the fast-as-it-can-go speed is for a single machine. Doing this for tens of thousands of people simultaneously faster than I can do with high end software on my $10,000 workstation? Alright, that's impressive.
I know there's some incredibly expensive NVIDIA cards but they must be doing some more magic on top of that.
How does it handle pictures of the swastika?
For those that don't know, before the Nazis used it, it was a symbol of hope and prosperity in the West and even appeared on Coca Cola marketing. Today it still is in Eastern cultures.
if anything, this shows the power disparity between the haves (they have this technology which gets better with time) and have nots (certainly me, but possibly also you) who get the super diluted version of this
Who holds their phone up and takes a photo then wants to know it was a photo of?
That’s weird. If you don’t know what it is, wtf did you take photo?
The obvious use here is natural language improvement / photo editing for photos, but this is just a stepping stone to that, and bluntly, as it stands… the examples really don’t shine…
Great for the vision impaired.
…not sure, what anyone else will use this for.
The only really compelling use case is the “code this ui for me”, but as we’ve seen, repeatedly, this kind of code generation only works for trivial meaningless examples.
Seems fun, but I doubt I’d use it.
(Which, and this is my point, is a massive step away from the current everyday usefulness of chatgpt)
>>Great for the vision impaired.
Yes, this is great for the estimated 285 million vision impaired people around the world[1].
That’s great. …but it’s niche.
I’m sitting on my couch right now and I can think of like 20 things I could chat to chatgpt about.
I can see literally nothing in my visual range want to take a photo of and run image analysis over.
It’s like Shazam. Yes, it’s useful, but, most of the time, I don’t need it.
I would argue this is true for this, for most people, including the significant proportion of people with minor visual impairments (that would, you know, put their glasses on instead).
This, wouldn't help them, even if they had both a device capable of using it and the means to pay for it.
Even if it could help people, it's an open question if it would be safe, to, for example, use this to scan medication when it is only a probabilistic model that may hallucinate something that isn't actually there.
What you're talking about is a speculative use of a service that might one day exist based on this technology.
What I am talking about is this actual service.
It's not a speculative service that might one day happen.
Literally it's rolling out right now
These people are being served by a preview of the service _right now_.
> Even if it could help people, it's an open question if it would be safe, to, for example, use this to scan medication when it is only a probabilistic model that may hallucinate something that isn't actually there.
Any OCR solution could also make a mistake, like misrecognizing a dosage on a prescription label.
> What you're talking about is a speculative use of a service that might one day exist based on this technology.
> What I am talking about is this actual service.
GPT-4 is six months old. ChatGPT is less than a year old. Why would you benchmark a service by the initial public preview? Of course it's _speculative use_, the damn thing has had its tires kicked for like a day.
What if you weren't on your couch? Going outside is not "niche".
I find myself doing this rather frequently. The scenario described in the article is quite common for me: capturing a photo of a plant and utilizing an existing classification service to determine its identity. It could be driven by mere curiosity or practical concerns like identifying whether a plant is poison ivy.
Wildlife identification also falls into this category. Recognizing different bird species can be challenging, especially when it's not a familiar species like a blue jay. I often find myself engaging in this activity quite regularly!
EDIT: I should also point out this happens with other forms of ‘unknown object identification’. There’s an entire subreddit that’s quite popular devoted to just crowd-sourcing identification based on a picture.
FYI Cornell Lab's Merlin app is fantastic at this, and its bird call audio identification is even better. They obviously have some top-notch machine learning going on there, and I'm really curious to see how both they and other services innovate on this front in the months to come.
Wouldn't say this is super reliable, I gave it a photo of a small squid in my hand and it said it was a baby fish (very obviously was not a fish).
OpenAI’s example included bike repair and toolkit choice
Allot of people could use this even if they aren’t right now
They’ll use YouTube, just like they do right now. Maybe if it could watch the video, then step you through it step by step. …but it cant, with what they’ve actually released here.
Oh whatever. If I’m wrong, I’m wrong. Time will tell.
and ad block doesn't work on mobile
if you have a case that wasn't covered by that video? you have to go to another or continue searching all while wishing you could just talk to someone about it. if you don't know the word for what you're looking for, all the search engines lack utility.
ChatGPT4 with image recognition and conversation solves all of that use case and people already use it, so now they'll just start sending it pictures from the phone already in their hand that they're already using to chat with
there are plenty of times over the last year that would have been useful for me. plenty of times over the last year I just didn't continue being interested in that problem
it just seems kind of…. late ?… for that “dont be ridiculous” reaction. classic dropbox moment
I do. For plants, and occasionally for birds.
This is likely how we'll communicate with information systems: throw some hand-wavy question at it, and refine your query based on its output using natural language until you find the answer (or even the question) you were looking for.
There are several popular "r/whatisthis(x)" subreddits: whatisthisthing, whatisthisbug, whatisthisplant, whatisthissnake, whatisthisrock, etc.
And there are many phone apps that attempt to do the same thing, like CoinSnap to identify coins.
Currently traveling in foreign country, this is exactly my use case seeing weird flowers and fruits I've never seen before.