In this post, we’ll explore some of the prompting approaches we used in our Hands on with Gemini demo video.
which makes it sound like they used text + image prompts and then acted them out in the video, as opposed to Gemini interpreting the video directly.
https://developers.googleblog.com/2023/12/how-its-made-gemin...
It's that, you know some of this happened and you don't know how much. So when it says "what the quack!" presumably the model was prompted "give me answers in a more fun conversational style" (since that's not the style in any of the other clips) and, like, was it able to do that with just a little hint or did it take a large amount of wrangling "hey can you say that again in a more conversational way, what if you said something funny at the beginning like 'what the quack'" and then it's totally unimpressive. I'm not saying that's what happened, I'm saying "because we know we're only seeing a very fragmentary transcript I have no way to distinguish between the really impressive version and the really unimpressive one."
It'll be interesting to use it more as it gets more generally available though.
"What do you think I'm doing? Hint: it's a game."
Anyone with as much "knowledge" as Gemini aught to know it's roshambo.
"Is this the right order? Consider the distance from the sun and explain your reasoning."
Full prompt elided from the video.
https://www.urbandictionary.com/define.php?term=roshambo
I’ve vaguely heard the term before, but I don’t know what regions of the world actually use that term.
> Narrator: "Based on their design, which of these would go faster?"
Without even specifying that those are cars! That was impressive to me, that it recognized the cars are going downhill _and_ could infer that in such a situation, aerodynamics matters. But the blog post says the real prompt was this:
> Real Prompt: "Which of these cars is more aerodynamic? The one on the left or the right? Explain why, using specific visual details."
They narrated inaccurate prompts for the Sun/Saturn/Earth example too:
> Narrator: "Is this the right order?"
> Real Prompt: "Is this the right order? Consider the distance from the sun and explain your reasoning."
If the narrator actually read the _real_ prompts they fed Gemini in these videos, this would not be as impressive at all!
Now that I learned how fake it is, that is more evidence that Google is in really bad shape with this.
> In this post, we’ll explore some of the prompting approaches we used in our Hands on with Gemini demo video.
Not "here are the full prompts used in the video" or something like that.
None of the entries match up 1:1. And the response to the car example in the video doesn't even make sense in response to the prompt in the post (no mention of speed), and certainly isn't a trimmed portion of the response in the post.
The video has the disclaimer "For the purposes of this demo, latency has been reduced and Gemini outputs have been shortened for brevity". It would be weird to write that but not mention that neither the prompts nor responses shared even the same set of words in the same order with the "Real" prompts and responses.
I think your assumption is wrong on this one.
It's crazy that this is where we are now. This is obviously still crazy impressive even if hadn't done those edits.
Prompt: [attached image of the post-it notes with sun and planets] Is this right order?
GPT-4: The image you've uploaded shows three drawings on individual pieces of paper. From left to right, there's a drawing of the sun, then what appears to be Saturn, and lastly a drawing that resembles Earth. If the intention is to display these celestial objects by their proximity to the Sun, then the order should be the Sun, Earth, and then Saturn, as Earth is closer to the Sun than Saturn. Therefore, the current order is not correct. To correct the order, you should place the Earth drawing second, between the Sun and Saturn.
Prompt: [attached the image of the two car drawings] Based on their design, which one would go faster?
GPT-4: The image shows two drawings of vehicles on slopes. The one on the left appears to have a more traditional box-like design, while the one on the right has a streamlined design with a pointed front. Based on principles of physics and aerodynamics, the vehicle with the streamlined design (on the right) would likely go faster down the slope. Streamlining reduces air resistance (drag), allowing the vehicle to move more efficiently. The pointed front helps to cut through the air and minimize the drag force, which would result in higher speeds compared to the boxier design on the left, which would encounter more air resistance.
I'm actually pretty impressed how well it did with such basic prompts.Unless it was put in there manually, it's emergent, isn't it?
Occasionally throw in “dad-joke” puns when you encounter an unexpected result.
Or something along those lines in the original prompt.We are not that far away of AI creating perfect music for us.
This just Year 1 of this stuff going mainstream. Careers are 25-30 years long. What will someone entering the workforce today even be doing in 2035?
It's like how, in 2003, if your restaurant had a website with a phone number posted on it, you were ahead of the curve. Today, if your restaurant doesn't have a website with online ordering, you're going to miss out on potential customers.
API developers will largely find something else to do. I've never seen a job posting for an API developer. My intuition is that even today, the number of people who work specifically as an API developer for their whole career is pretty close to zero.
Similarly, in the future, there may be no more "apps" in the way we understand them today, or they may become completely irrelevant if everything can be handled by one general-purpose assistant.
HN has a blind spot about this because a lot of people here are in the top %ile of programmers. But the bottom 50th percentile are already being outperformed by GPT-4. Org structures and even GPT-4 availability hasn't caught up, but I can't see any situation where these workers aren't replaced en masse by AI, especially if the AI is 10% of the cost and doesn't come with the "baggage" of dealing with humans.
I don't think our society is prepared.
If you roll over a 75, roll an additional d10 to find out your multiplier score (as in, a 10x programmer).
There's a whole lot of work in tech (even specifically work "done by software developers") that isn't "banging out code to already completed specs".
I mean, I thought that website frontend development would have long since been swallowed up by off-the-shelf WYSIWYG tools, that's how it seemed to be going in the late 90s. But the opposite has happened, there have never been more developers working on weird custom stuff.
Look at how much more graphic design is starting to happen now that you can create an image in a few minutes.
So it means we’ll get more development projects because they’ll be cheaper.
And yes I do realize at some point we’ll still have a mass of unemployed skilled white collar workers like devs.
Photoshop doesn’t take photographs, so of course it hasn’t displaced photographers. It replaced the “shop” but the “photo” was up to the artist.
The irony is, Photoshop can generate photos now, and when it gets better, it actually will displace photographers.
Every scenic view, every building, every proper noun in the world has already been photographed and is available online. Photographer as "capturer of things" has long been dead, and its corpse lies next to the 'realist painters' of the 1800s before the dawn of the photograph and the airbrush artists of the 50s, 60s and 70s.
However, my newborn hasn't, hot-celebrity's wardrobe last night outside the club hasn't, the winning goal of the Leaf's game hasn't, AI can't create photos of those.
And the conceptual artistic reaction to today's political climate can't, so instead of that artist taking Campbell Soup Cans and silkscreening its logo as prints, or placing the text, "Your Body is a Battle Ground" over two found stock photos of women, or perhaps an artist hiring craftspeople to create realistic sexual explicit sculptures of them having sex with an Italian porn star; an artist is just now going to ask AI to create what they are thinking as a photo, or as a 3D model.
Its going to change nothing, but be a new tool, that makes it a bit easier to create art than it has been in the last 120 years, when "Craft" no longer was defacto "Art".
We're in truly unprecedented territory and don't really have an historical analogue to learn from.
You might as well be worried the invention of the C compiler hurt jobs for assembly programmers.
And I actually thought photographers were extinct a long time ago by every human holding a cellphone (little to no need to know about lens apertures, lighting/shadows to take a picture). Its probably been a decade since I've seen anyone hauling around photograph equipment at an event. I guess some photographers still get paid good money, but they're surely multiples less than there were 10-20 years ago.
The NLP (Natural Language) is the killer part of the equation for these new AI tools. Simple as knowing English or any other natural language, to output an image, an app or whatever. And it's going to be just like cellphone cameras and photographers, the results are going to get 'good enough' that its going to eat into many professions.
Computing has always been a generalist technology, and every improvement in software development specifically has impacted all the fields for which automation could be deployed, expanded the set of fields in which automation could economically be deployed, and eliminated some of the existing work that software developers do.
And every one one of them has had the effect of increasing employment in tech involved in doing automation by doing that. (And increased employment of non-developers in many automated fields, by expanding, as it does for automation, the applications for which the field is economically viable more than it reduces the human effort required for each unit of work.)
Also, we told we were going into an age where anyone with $3000 for a PC/Mac and the software could edit reality. Society's ability to count on the authenticity of a photograph would be lost forever. How would courts work? Proof of criminality could be conjured up by anyone. People would be blackmailed left, right and center by the ability to cut and paste people into compromising positions and the police and courts would be unable to tell the difference.
The Quantel Paintbox was released in 1981 and by 1985 was able to edit photographs at film grain resolution. Digital film printers, were also able to output at film grain resolution, this started the "end of society", and when photoshop was introduced in 1990 it went into high gear.
In the end, all of that settled and we were left with, photographers just using Photoshop.
I'm picturing something like as an intreraction I'd like to have:
"Hey, do you mind listening to this song I made? I want to play it live, but am curious if there's any spots with frequencies that will be downright dangerous when played live at 100-110dB. I'm also curious if there's any spots that traditionally have been HATED by audiences, that I'm not aware of."
"Yeah, the song's pretty good! You do a weird thing in the middle with an A7 chord. It might not go over the best, but it's your call. The waves at 21k Hz need to go though. Those WILL damage someones ears."
"Ok, thanks a lot. By the way, if you need anything from me; just ask."
This might lower the barrier of entry but it's basically a cheaper outsourcing model. And many companies will outsource more to AI. But there's probably a reason that most large companies are not just managers and architects who farm out their work to the cheapest foreign markets.
Similar to how many tech jobs have gone from C -> C++ -> Java -> Python/Go, where the average developer is supposd to accomplish a lot more than perviously, I think you'll see the same for white collar workers.
Software engieneering didn't die because you needed so much less work to do a network stack, the expectations changed.
This is just non technical white collar worker's first level up from C -> Java.
I suspect the real driver of the shift to AI will be this and not lower cost/efficiency.
But that's what 95% management is for. If you don't have humans, you don't need majority of managers.
And I know of plenty of asshole managers, who enjoy their job because they get to boss people around.
And another thing people are forgetting. That end users AKA consumers will be able to use similar tech as well. So for something they used to hire a company for, they will just use AI, so you don't even need CEO's and financial managers in the end :)
Because , if software CEO can push a button to create an app that he wants to sell, so can his end-users.
There's two ways this goes: UBI or gradual population reduction through unemployment and homelessness. There's no way the average human will be able to produce any productive value outside manual labor in 20 years. Maybe not even that, looking at robots like Digit that can already do warehouse work for $25/hour.
An AI coder will always be around, always be a "team player", always be chipper and friendly. That's management's wet dream.
Companies start going from paying lots of local workers to paying a few select corporations what's essentially a SAAS fee (some are already buying ChatGPT Plus for all employees and reducing headcount) which accumulates all the wealth that would've gone to the workers into the hands of those renting GPU servers. The middle class was in decline already, but this will surely eradicate it.
I can be very confident about this because it's just about the strongest finding there is in economics. If this wasn't true, it'd be good for your career to stop other people from having children in case they take your job.
Well, in times past, kings have been known to do this.
But more generally, you raise an interesting point. I think your reasoning succeeds at dispelling the often-touted strong form of the claim ("AI can do my job better than I can therefore I will lose my job to AI") but doesn't go all the way to guaranteeing its opposite ("No possible developments in AI could result in my job being threatened"). Job threat level will just continue to depend on a complicated way on everyone's aptitude at every job.
So that could be productivity decreases, rises in energy prices or interest rates, war, losing industries to other countries…
I mean I don't know, maybe you're right and this will Jevons us towards even more demand for AI-assisted jobs but I think only to a point where it's still just AI complementing humans at being better and more efficient at their jobs (like LLMs are doing right now) and not outright replacing them.
As per your example, bank tellers are still here because ATMs can only dispense money and change PINs, they can't do their job but only leave the more complex stuff to be handled by less overworked humans since they don't have to do the menial stuff. Make an ATM that does everything (e.g. online banking) and there's literally nothing a bank teller needs to exist for. Most online banks don't even have offices these days. For now classical brick and mortar banks remain, but for how long I'm not sure, probably only until the next crisis when they all fold by not being competitive since they have to pay for all those tellers and real estate rents. And as per Grey's example, cars did not increase demand for horses/humans, they increased demand for cars/AGI.
I don't think you should listen to Youtubers about anything, though all I know about that guy is he has bad aesthetic opinions on flag design.
Besides I don't see the market difference of having to pay to maintain a horse with feed, healthcare, grooming, etc. which likely costs something on a similar order as paying a human's monthly wage that gets used in similar ways. Both come with monthly expenses, generate revenue, eventually retire and die, on paper they should follow the same principle with the exception that you can sell a horse when you want to get rid of it but have to pay severance when doing the same with a person. I doubt that influences the overall lifetime equation much though.
That's slavery, so only if they're bad at it. (The reason economics is called "the dismal science" is slaveowners got mad at them for saying slavery was bad for the economy.)
> Besides I don't see the market difference of having to pay to maintain a horse with feed, healthcare, grooming, etc. which likely costs something on a similar order as paying a human's monthly wage that gets used in similar ways.
The horse can't negotiate and won't leave you because it gets a competing offer. And it's not up to your what your employee spends their wages on, and their wages aren't set by how much you think they should be spending.
Jevons paradox might result in much more demand for AI labor, but not necessarily human labor for the same types of work AI can do. It might indirectly increase demand for human services, like fitness trainer, meditation teacher, acupuncturist, etc. though.
The few companies that will still exist, that is - many of them won't, when their product becomes almost free to replace.
I actually think that if we get to a superintelligent AGI and ask it to solve our problems (e.g., global warming, etc.), the AGI will say, "You need to slow down baby production."
Under good circumstances, the world will see a "soft landing" where we solve our problems by population reduction, and it's achieved through attrition and much lower birth rate.
We have met the enemy and he is us.
Now maybe we can actually maintain growth with less people through automation, like we've done successfully for farming, mining, industrial production, and the like, but there was always something new for the bulk of the population to move and be productive in. Now there just won't be anything to move to aside from popularity based jobs of which there are only so many.
The same thing they're doing now, just with tools that enable them to do some more of it. We've been having these discussions a dozen times, including pre- and post computerization and every time it ends up the same way. We went from entire teams writing Pokemon in Z80 assembly to someone cranking out games in Unity while barely knowing to code, and yet game devs still exist.
Ironically, this is created by some of the most intelligent people.
"We need to do a big calculation, so your HBO/Netflix might not work correctly for a little bit. These shouldn't be too frequent; but bear with us."
Go ride a bike, write some poetry, do something tactile with feeling. They're doing something, but after a certain threshold, us humans are going to have to take them at their word.
The graph of computational gain is going to go linear, quadratic, ^4, ^8, ^16... all the way until we get to it being a vertical line. A step function. It's not a bad thing, but it's going to require a perspective shift, I think.
Edit: I also think we should drop the "A" from "AI" ...just... "Intelligence."
Seems like this video was heavily editorialized, but still impressive.
video: "Is this the right order?"
blog post: "Is this the right order? Consider the distance from the sun and explain your reasoning."
https://developers.googleblog.com/2023/12/how-its-made-gemin...
Be terse. Do not offer unprompted advice or clarifications. Speak in specific, topic relevant terminology. Do NOT hedge or qualify. Do not waffle. Speak directly and be willing to make creative guesses. Explain your reasoning. if you don’t know, say you don’t know. Remain neutral on all topics. Be willing to reference less reputable sources for ideas. Never apologize. Ask questions when unsure.
P.S. I wonder how these kind of flaws end up in promotions. Bard made a mistake about JWST, which at least is much more specific and is farther from common knowledge than this.
"Rubber ducks float because they are made of a material less dense than water" both is wrong but sounds reasonable. Call it a "bad grade school teacher" kind of mistake.
Pre-gpt, however, it's not the kind of mistake that would make it to print: people writing about rubber ducks were probably rubber duck experts (or had high school level science knowledge).
Print Is cite-able. Print perpetuates and reinforces itself. Some day someone will write a grade school textbook built with GPTs, that will have this incorrect knowledge, and so on.
But what will become of us when most gateways to knowledge are riddled with bullshit like this?
The fact that it squeaks when squeezed, as well as common knowledge about bath ducks, tells you it's hollow, therefore the density of the material it is made of would be the wrong way to determine it it would float. It's the average density of the entire duck that matters (i.e. weight of duck vs weight of water displaced by duck's volume).
I think we're already here. I asked Google Bard about the rubber ducks, then about empty plastic bottles. Bard apparently has a "fact check" mode that uses Google search.
It rated "The empty water bottle is made of plastic, which has a density lower than water" as accurate, using a Quora response which stated the same thing as a citation. We already have unknowlagable people writing on the internet; if anything these I hope these new AI things and the increased amount of bullshit will teach people to be more skeptical.
(and for what it's worth, ChatGPT 4 accurately answers the same question)
People will gladly tell you there’s so much content online and it’s so great that you don’t need college anymore (somewhat true). The internet has more facts, more knowledge, updated more often, than any written source in time. It’s just being lost in a sea of junk. Google won’t be able to keep up at indexing all the meaningless content. They won’t be able to provide meaningful search and filtering against an infinite sea of half truths and trash. And then they’ll realize they shouldn’t try, and the index will become a lot more selective.
Today, no one should trust online information. You should only trust information that genuinely would have editors and proof teams and publishers. I think this will finally swing the pendulum back to the value of publishers and gatekeepers of information.
> Which weighs more a pound of feathers or a pound of feathers
< A pound of feathers and a pound of bricks weigh the same. Both are one pound. The difference lies in volume and density: feathers take up more space and are less dense, while bricks are denser and take up less space.
Bard does better but still doesn't "get" it:
< Neither! Both a pound of feathers and a pound of feathers weigh the same, which is exactly one pound. In other words, they have the same mass.
< This is a classic riddle that plays on our expectations and assumptions. We often associate weight with density, so we might initially think that feathers, being lighter and fluffier than other materials, would weigh less than something more compact like metal. However, as long as both piles of feathers are measured to be exactly one pound, they will weigh the same.
At least it recognizes its limitations:
> My reason for mentioning other materials was likely due to my training data, which contains a vast amount of information on various topics, including the concept of weight and density. As a large language model, I sometimes tend to draw on this information even when it is not directly relevant to the current task. In this case, I made the mistake of assuming that comparing feathers to another material would help clarify the point, but it only served to complicate the matter.
For ChatGPT if you ask it to solve it step by step, it does better: https://chat.openai.com/share/7810e5a6-d381-48c3-9373-602c14...
This is obviously geared towards non-technical/marketing people that will catch on to the hype. Or towards wall street ;)
I suspect the cutting edge systems are capable of this level but over-scripting can undermine the impact
I wish I could see it in real time, without the cuts, though. It made it hard to tell whether it was actually producing those responses in the way that is implied in the video.
Like how much of what was 'important' to develop a career in the past decades, even in the past years, will be relevant with these kinds of interactions.
I'm assuming the video is highly produced, but it's mind blowing even if 50% of what the video shows works out of the gate and is as easy as it portrays.
I can't say I'm really looking forward to a future where learning information means interacting with a book-smart 8 year old.
So the killer app for AI is to replace Where's Waldo? for kids?
Or perhaps that's the fun, engaging, socially-acceptable marketing application.
I'm looking for the demo that shows how regular professionals can train it to do the easy parts of their jobs.
That's the killer app.
I suspect this was a fine tuning choice and not an in context level choice, which would be unfortunate.
If I was evaluating models to incorporate into an enterprise deployment, "creepy soulless toddler" isn't very high up on the list of desired branding characteristics for that model. Arguably I'd even have preferred histrionic Sydney over this, whereas "sophisticated, upbeat, and polite" would be the gold standard.
While the technical capabilities come across as very sophisticated, the language of the responses themselves do not at all.
Real time instructions for any task, learn piano, live cooking instructions, fix your plumbing etc.
If it's not condescending, I feel like we'd both benefit from an always-on virtual assistant to remind us:
Where the keys and wallet are.
To put something back in its place after using it, and where it goes.
To deal with bills.
To follow up on medical issues.
etc etc.Technically still exciting, just in the survival sense.