Veo
deepmind.google
deepmind.google
Wider bandwidth isn’t always better.
How would you know?
“I can’t see anything” “Maybe that’s something over there?” “What’s everyone looking at?”
Someone shows their phone.
“Ooh!” “How do you turn on night mode?” “Wow it’s so much clearer on the phone!”
So I can’t know what their eyes see or what they really think, I could hear what came out of their mouths.
I don’t think this is an instance that warrants deep philosophical skepticism about the nature of truth or the impossibility of knowledge.
see: https://theconversation.com/what-causes-the-different-colour...
> Thus, the human eye primarily views the Northern Lights in faint colors and shades of gray and white. DSLR camera sensors don't have that limitation. Couple that fact with the long exposure times and high ISO settings of modern cameras and it becomes clear that the camera sensor has a much higher dynamic range of vision in the dark than people do.
https://www.space.com/23707-only-photos-reveal-aurora-true-c...
This aligns with my experiences.
The brightest ones I saw in Northern Canada I even saw hints of reds - but no real greens - until I looked at it through my phone, and it looked just like the simulated video.
If I looked up and saw them the way they appear in the simulation, in real life, I'd run for a pair of leaded undies.
And yes, they can be as green to the naked eye in that AI video. I've seen aurora shows that fill the entire night sky from horizon to horizon, way more impressive than that AI video with my own eyes.
(Edit: it was still super-cool even if grey-ish, and there was absolutely beautiful colors in there if you could find your way out of the direct city lights)
I've also seen the northern lights with my own eyes. Way up in the arctic circle in Sweden. Their color changes along with activity. Grey looking sometimes? Sure. But also colors that are so vivid that it feels like it envelopes your body.
Then again, the average viewer in Seattle this past weekend is hardly representative of what the northern lights look like.
The H in HN stands for Hubris.
Phone cameras have a Bayer filter which means they only have RGB color-sensing. The Bayer filter cuts out some incoming light and dims the received image, compared with what a monochrome camera would see. But that's how you get color photos.
To compensate for a lack of light, the phone boosts the gain and exposure time until it gets enough signal to make an image. When it eventually does get an image, it's getting a color image. This comes at the cost of some noise and motion-blur, but it's that or no image at all.
If phone cameras had a mix of RGB and monochrome sensors like the human eye does, low-light aurora photos might end up closer to matching our own perception.
But what sort of eyes are those?
Priming the opsins in your retina is a continuous process, and primed opsins are depleted rapidly by light. Fully adapting your eye to darkness takes a great deal of darkness and a great deal of time - on the order of an hour should set you up.
Most human beings in arctic regions live in places and engage in lifestyles where it's impossible to even come close to attaining the full light sensitivity of the human retina in perfect darkness. The sky never gets dark enough in a city or even a small town to get the full experience, and if you saw your smart watch five minutes ago you still haven't fully recovered your night vision. Even a sliver of moon makes remote dark-sky-sites dramatically brighter.
Everybody is going to have different degrees of the experience because they'll have eyes with different degrees of dark adaptation. And their brains are going to shift around the ~10^3x dynamic range of the eye up or down the light intensity scale by a factor ~10^6, without making it obvious to them.
The golden standard for rendering has always been cameras. It’s always photo-realistic rendering. Maybe this won’t be true for VR, but so far most effort is to be as good as video, not as good as the human eye.
Any sort of video generation AI is likely to have the same goal. Be as good as top notch cameras, not as eyes.
I echo what some other posters here have said: they're certainly not gray.
The fact that some things are too dark to be seen by humans but can be captured accurately with cameras doesn't mean that the camera, or the AI, is "making things up" or whatever.
Finally, nobody wants to see a video or a photo of a dark, gray, and barely visible aurora.
Except those who want to see an accurate representation of what it looks like to the naked eye.
I highly recommend checking them out if you're nearby, the recent auoras have been quite astonishing
I work at Andøya Space where perhaps most of the space research on Aurora had been done by sending scientific rockets into space for the last 60 yrs.
> Prompt: Timelapse of the northern lights dancing across the Arctic sky, stars twinkling, snow-covered landscape
because completeness isnt really what we are going for
https://www.reddit.com/r/dalle2/comments/1afhemf/is_it_possi...
https://www.reddit.com/r/dalle2/comments/1cdks71/a_hand_with...
This is what dall-e came up with after trying to correct many previous iterations: https://imgur.com/Ss4TwNC
Conventionally this term means the opposite -- problems that AI unlocks that conventional computing could not do. Conventional computing can render a very wide range of different stylized chess boards, but when an ML technique like diffusion is applied to this mundane problem, it falls apart.
It seems that the SynthID is not only for AI generated video but for image, text and audio.
For that it needs a "director" to say: "turn the horse's head 90˚ the other way, trot 20 feet, and dismount the rider" and "give me additional camera angles" of the same scene. Otherwise this is mostly b-roll content.
I'm sure this is coming.
And when we want it to match exactly in an animatic or whatever, it needs to be far more precise than this, matching real locations etc.
I could see this being really useful for exploring tone, movement, shot sequences or cut timing, etc..
Right now you scrape together "kinda close enough" stock footage for this kind of exploration, and this could get you "much closer enough" footage..
So if it’s an optional tool, great, but some people would be fine with it, some would not.
I've worked with other developers that want to build high fidelity wire frames, sometimes in the actual UI framework, probably because they can (and it's "easy"). I always push back against that, in favor of using whiteboard or Sharpies. The low-fidelity brings better feedback and discussion: focused on layout and flow, not spacing and colors. Psychologically it also feels temporary, giving permission for others to suggest a completely different approach without thinking they're tossing out more than a few minutes of work.
I think in the artistic context it extends further, too: if you show something too detailed it can anchor it in people's minds and stifle their creativity. Most people experience this in an ironically similar way: consider how you picture the characters of a book differently depending on if you watched the movie first or not.
That said, I personally think the solution will not be coming that soon, but at the same time, we'll be seeing a LOT more content that can be done using current tools, even if that means a dip in quality (severely) due to the cost it might save.
Because camera angles/lighting/collision detection/etc. at that point would be almost trivial.
I guess with the "2D only" approach that is based on actual, acquired video you get way more impressive shots.
But the obvious application is for games. Content generation in the form of modeling and animation is actually one the biggest cost centers for most studios these days.
And let the robot tween?
Vs an imperative for "tween this by turning the horse's head left"
And there are a lot more degrees of freedom to get something wrong in film than in a single still image.
Maybe this works for ads for duner place or shisha bar in some developing country. I’ve seen generated images used for menus in such places.
But I doubt a serious filmography can be done this way. And if it can - it’d be again thanks to some smart concept on behalf of humans.
The problem is that these video clips are very unimpressive compared to the Sora demonstration which came out three months ago. If this demo was announced by some scrappy startup it would be worth taking note. Coming from Google, the inventor of the Transformer and owner of the largest collection of videos in the world, these sample videos are underwhelming.
Having said that, Sora isn't publicly available yet, and maybe Veo will have more to offer than what we see in those short clips when it gets a full release.
Google the company known to launch way too many products? What other big company launches more stuff early than them? What people complain about Google is that they launch too much and then shut them down, not that they don't launch things.
I’ve switched to Opus from GPT-4 for coding and it was non-trivially easy
wow the speed at which we can be blasé is terrifying. 6 months ago this was not possible, and felt this was years away!
They're not underwhelming to me, they're beyond anything I thought would ever be possible.
are you genuinely unimpressed? or maybe trying to play it cool?
Just a few years ago, I would have been absolutely blown away by these demo videos. Six months ago, I would have been very impressed. Today, Google is rolling a product that seems second best. They're playing catch-up in a game where they should be leading.
I will still be very impressed to see videos of that quality generated on consumer grade hardware. I'll also be extremely impressed if Google manages to roll out public access to this capability without major gaffes or embarrassments.
This is very cool tech, and the developers and engineers that produced it should be proud of what they've achieved. But Google's management needs to be asking itself how they've allowed themselves to be surpassed.
Even search, in and of itself, is incredibly amazing but fairly commoditized at this point. They should've highlighted more unique footage.
I’ll just go back to living under a rock.
I would generally agree though, it is not normal they didn’t show more human
That is, AI humans can look "creepy" whereas AI animals may not. The cowboy looks pretty good precisely because it's all shadow.
CGI animators can probably explain this better than I can ... they have to spend way more time on certain areas and certain motions, and all the other times it makes sense to "cheat" ...
It explains why CGI characters look a certain way too -- they have to be economical to animate
By comparison, the shots here are only a few seconds long and almost all look like slow motion or slow panning shots cherrypicked because they don't have that much movement. Compare that to Sora's videos of people walking in real speed.
The only shot they had that can compare was the cyberpunk video they linked to, and it looks crazy inconsistent. Real shame.
It's impressive as hell though. Even if it would only be used to extrapolate existing video.
Sora videos ran at 1 beat per second, so everything in the image moved at the same beat and often too slow or too fast to keep the pace.
It is very obvious when you inspect the images and notice that there are keyframes at every whole second mark and everything on the screen suddenly goes in their next animation step.
That really limits the kind of videos you can generate.
My point is that it isn't obvious at all that Soras way actually is closer to the end goal, it might look better today to have those 1 second beats for every video but where do you go from there?
I think comparing them now is probably not that useful outside of this AI hype train. Like comparing two children. A lot can happen.
The bigger message I am getting from this is it's clear OpenAI won't have a super AI monopoly.
Video models are interesting, and to some extent trying to imagine which company is gonna eat the other’s lunch is kind of interesting, but sometimes that’s all people are interested in and I can see my girlfriend's reasoning for being disinterested in such discussion.
Citation needed?
Although I did not work in AI, I did work at Google X robotics on a robot they often use for AI research.
Maybe some people felt like it was a competition, but I don’t have much reason to believe that feeling is common. AI researchers are literally in collaboration with other people in the field, publishing papers and reading the work of others to learn and build upon it.
When OpenAI suddenly stopped publishing their stuff I bet that many researchers now started feeling like it started to be a competition.
OpenAI is no longer cooperating, they are just competing. They still haven't said anything about how gpt-4 works.
I think this is arguably better than the alternative. With slow-mo generated videos, you can always speed them up in editing. It's much harder to take a fast-paced video and slow it down without terrible loss in quality.
being cautious often puts a dent in innovation
The most impressive Sora demo was heavily edited.
It's quite early to race to the conclusion that one is better than the other when not only they are both unreleased, but especially when the demos can be edited, faked or altered to look great for optics and distortion.
EDIT: It appears there is at least one commenter who replied below that is upset with this fact above.
It is OK to cope, but the truth really doesn't care especially when the competition (Google) came out much stronger than expected with their announcements.
Distortion is easiest when the products really work. :)
https://www.youtube.com/watch?v=KFzXwBZgB88 (posted the day after the short debuted)
https://openai.com/index/sora-first-impressions (no mention of editing, nor do they link to the above making-of video)
>The videos below were edited by the artists, who creatively integrated Sora into their work, and had the freedom to modify the content Sora generated.
https://web.archive.org/web/20240513050023/https://openai.co...
They also just added a link to the making-of video.
The intention wasn't to show "This is what Sora can generate from start to end" but rather "This is what a video production team can do with Sora instead of shooting their own raw footage."
Maybe not so obvious to others, but for me it was clear from how the other demo videos looked.
I can see more real-world impact from this (and/or Sora) than most other AI tools
And I'm someone who is fine playing fast action video games. Can't imagine what it's like if you're older or have sensory processing issues.
> Some of y'all may find how awful this editing gets pretty interesting: I did an Average Shot Length (ASL) for many movies for a recent project, and just to illustrate bad overediting in action movies, I looked at Taken 3 (2014) in its extended cut.
> The longest shot in the movie is the last shot, an aerial shot of a pier at sunset ending the movie as the end credits start rolling over them. It clocks in at a runtime of 41 seconds and is, BY FAR, the longest shot in the movie.
> The next longest is a helicopter establishing shot of the daughter's college after the "action scene" there a little over an hour in, at 5 seconds.
> Otherwise, the ASL for Taken 3 (minus the end credits/opening logos), which has a runtime of 1:49:40, 4,561 shots in all (!!!), is 1.38 SECONDS . For comparison, Zack Snyder's Justice League (2021) (minus end credits/opening logos) is 3:50:59, with 3163 shots overall, giving it an ASL of 4.40 seconds, and this movie, at 1 hour 50 minutes, has north of 4,561 for an ASL of 1.38 seconds?!?! Taken 3 has more shots in it than Zack Snyder's Justice League, a movie more than double its length...
> To further illustrate how ridiculous this editing gets, the ASL for Taken 3's non-action scenes is 2.27 seconds. To reiterate, this is the non-action scenes. The "slow scenes." The character stuff. Dialogue scenes. The stuff where any other movie would know to slow down. 2.27 SECONDS For comparison, Mad Max: Fury Road (minus end credits/opening logos) has a runtime of 1:51:58, with 2646 shots overall, for an ASL of 2.54 seconds. TAKEN 3'S "SLOW SCENES" ARE EDITED MORE AGGRESSIVELY THAN MAD MAX: FURY ROAD!
> And Taken 3's action scenes? Their ASL is 0.68 seconds!
> If it weren't for the sound people on the movie, Taken 3 wouldn't be an "action movie". It'd be abstract art.
"He's 68. I'm guessing they stitched it together like this because "geriatric spends 30 seconds scaling chainlink fence then breaks a hip" doesn't exactly make for riveting action flick fare."
Lingering shots are horrible for obscuring things.
And Neeson was only 60 when filming Taken 3.
There ways to shoot an action scene with an aging star that doesn't involve 14 cuts in 4 seconds. You just have to care about your craft.
I can tell what's going on, but I always end up feeling agitated.
Maybe it's easy, and you feed continuity stills into the prompt. Maybe it's not, and this will always remain just a more advanced storyboarding technique.
But then again, storyboards are always less about details and more about mood, dialog, and framing.
Just worth keeping that in mind. You could not just switch between multiple shots like you can today.
Google and Apple have the ecosystem advantage.
Apple in particular has the deeper stack integration advantage.
Both Apple and Google have a somewhat poor software innovation reputation.
How does it all net out? I suspect ecosystem play wins in this case because they can personalize more deeply.
https://www.crn.com/news/networking/2024/google-cloud-posts-...
They are embedding their models not only widely across their platforms suite of internal products and devices, but also computationally via API for 3rd party development.
Those are all free from any perceived golden handcuffs that AdWords would impose.
[1] https://www.cnbc.com/2021/05/18/how-does-google-make-money-a...?
[2] https://aag-it.com/the-latest-cloud-computing-statistics/?t
Google (and to a lesser extend also Microsoft and Meta) also have a data advantage, they've been building search engines for years, and presumably have a lot more in-house expertise on crawling the web and filtering the scraped content. Google can also require websites which wish to appear in Google search to also consent to appearing in their LLM datasets. That decision would even make sense from a technical perspective, it's easier and cheaper to scrape once and maintain one dataset than to have two separate scrapers for different purposes.
Then there's the bias problem, all of the major AI companies (except for Mistral) are based in California and have mostly left-leaning employees, some of them quite radical and many of them very passionate about identity politics. That worldview is inconsistent with a half of all Americans and the large majority of people in other countries. This particularly applies to the identity politics part, which just isn't a concern outside of the English-speaking world. That might also have some impact on which AI companies people choose, although I suspect far less so than the previous two points.
X is not going to sit quietly as well.
There is also the rest of us.
Also engineering wise, currently every tweet is followed by a reply "my nudes in profile" and X seems unable to detect it as trivial spam, I doubt they have the chops to compete in this arena, especially after the mass layoffs they experienced.
I'm assuming you mean reputation as in general opinion among developers? Because Google's probably been the most innovative company of the 21st century so far.
See 1:11 in this video https://www.youtube.com/watch?v=MAvid5fzWnY
Incidentally that was one of the early uses of computer graphics in a movie, supposedly those short scenes took many hours to render and had to be done three times to achieve a colorized image.
But the parallel you made between android Brynner's vision and the generated imagery is fun to consider!
And we're supposed to believe that this is resilient against prompt injection?
How do you prevent state actors from creating "proof" that their enemies engaged in acts of war, and they are only engaging in "self-defense"?
It's not just technology though. Globalization has added so many layers between us and the objects we interact with.
I think Etsy was a bit ahead of their time. It's no longer a marketplace for handcrafted goods - it got overrun by mass produced goods masquerading as something artisan. I think the trend is continuing and in 5-10 years we'll be tired of cheap and plentiful goods.
I never heard HN claiming that Copilot will replace programmers. Why do so many people believe generative AI will replace artists?
"Hey guys big artist says this is fine so we're good"
This is what movie pre-vis is actually like, it doesn't need to be pretty, it needs to be precise:
https://www.youtube.com/watch?v=KFzXwBZgB88
https://www.fxguide.com/fxfeatured/actually-using-sora/
> While all the imagery was generated in SORA, the balloon still required a lot of post-work. In addition to isolating the balloon so it could be re-coloured, it would sometimes have a face on Sonny, as if his face was drawn on with a marker, and this would be removed in AfterEffects. similar other artifacts were often removed.
Why does this have to be so confusing? Is the name "Veo" or "VideoFX"? Why is the waitlist for VideoFX telling me something about public availability of ImageFX and MusicFX? Why is everything US only, again? Sigh..
Censored models are not going to work and we need someone to charge for an explicit model already that we can trust.
(this post is sarcastic)
How is this achieved? Is there temporal memory between frames?
Kudos to Google for if not foregrounding, being entirely transparent, about this.
https://findthatmeme.com/blog/2023/01/08/image-stacks-and-ip...
discussed on HN:
For straight OCR, it does work really well but at the end of the day its still not 100%
If I don't know better I'd think you just cherry-picked the prompts with the best-looking results.
1. safety measures lead to huge quality reductions
2. the devil's in the details. you can make me 1 million videos which look 99% realistic, but it's useless. consumers can pick it instantly, and it's a gigantic turn-off for any brand
No matter how good AI gets, it will never be the highest budget. Hell, even technically more accurate quartz watches cannot compete price wise with mechanical masterpiece watches of lower accuracy
Turns out if you open the video in a new tab the smoothness is much more impressive.
Why are they working there then ?
Huge grammar error on front page too.
We really are about this close to infinite jest. Imagine TikTok's algorithm with on demand video generation to suit your exact tastes. It may erase the social aspect, but for many users I doubt that would matter too much. "Lurking" into oblivion.
But those still require some human input. I'm imagining a sort of genetic algorithm for video prompts, no human editing, input, or curation required.
In fact I would say the comments are too good. They clearly have something ranking them for "niceness" but it makes them impossibly sentimental. Like I watched a bunch of videos about 70s rock recently and every single comment was about how someone's family member just died of cancer and how much they loved listening to it.
Walter Reuther: Henry, how are you going to get them to buy your cars?
So...you're not cynical, it's an explicit product goal.
Reminds me of the competition in tech in the late 80's early 90's between Microsoft and Borland, Microsoft and IBM, AMD and Intel, Word vs Wordperfect, etc.
It's a two horse race between Google and OpenAI.
Even at $1 per 5-second video, I think some use cases (including fun/non-business ones) would still overwhelm capacity.
The experiments often/usually fail, but they do experiment.
If Google were as focused on ads as you seem to think we'd at least see some sort of coherent org-wide strategy instead of a complete lack of direction.
I'm referring to this article that was posted here recently:
Nonetheless, she was a good engineer and a good manager, back when we crossed path many moons ago.
https://searchengineland.com/liz-reid-google-new-head-of-sea...
Google had blinders on. They didn't relentlessly focus on reinventing their domain. They just milked what they had. Gradually losing site of the user experience[1] to focus on monetization above all else.
The cyberpunk video seems better in that aspect, but I wish there were more.
How apropos...
How did I know I would see this message before clicking "Sign up to try"?
1. Generating an image of "a group of catgirls activating a summoning circle". Anything related to catgirls tends to get tagged as sexual or NSFW so it's censored. Unsurprising.
2. The lamb described in Book of Revelation. Asking for it directly or pasting in the passage where the lamb is described both fail to generate any images. Normally this fails because there's not much art of the lamb from Book of Revelation from which the model can steal. If I gave the worst of artists a description of this, they'd be able to come up with something even if it's not great.
Overall, a very disappointing release. It's surprising that despite having effectively infinite money this is the best that Google is able to ship at the moment.
Sigh. I still have hopes for VEO though
DeepMind people: AI can do it.
First, randomly selected 'feeling lucky' prompt got rejected, because it did not meet some criteria and pop-up helpfully listed FAQ to explain to me how I should be more sensitive to the program. I found it amusing.
Then I played with a couple of images, but it was nothing really exciting one way or another.
I guess you can color me disappointed overall. And no, I don't consider videos on repeat sufficient.
So of course you are going to see snarky comments and straight up denial in the competition. We saw that yesterday in the comments with the release of GPT4o in anticipation of Gemini 2.0 (GPT-5 basically) release being announced today at Google I/O
I'm SORA to say Veo looks much more polished without jank.
Big congratulations to Google and their excellent AI team for not editing their AI generated videos like SORA
I don’t know if there is a sentiment analysis tool for HN, but I’m pretty sure it’s been dead negative for Altman since at least Worldcoin.
You have to be pretty deep inside your own little bubble to think that even more than a 0.001% of HN has "stakes in YC" or "secondary shares in OpenAI".
I wouldn't discard.
Yes, there is a noticeable negative response from HN towards Google, and there has always been especially when speaking about their weird product management practices and incentives. Google hasn't launched any notable (and still surviving, Stadia being a sad example of this) consumer product or service in the last 10 years.
But to suggest there is a Sam Altman / OpenAI bias is delusional. In most posts about them there is at least some kind of skepticism or criticism towards Altman (his participation in Worldcoin and his accelerationist stance towards AGI) or his companies (OpenAI not being really open).
PS: I would say most people lurking here are just hackers (of many kinds, but still hackers), not investors with shady motives.
None of this is surprising to me and shouldn't shock you. You are literally on a site called Ycombinator. Had this been another platform without ties to investments or drawing from crowd that actively seeks to enrich themselves through participation in a narrative, this wouldn't even be a thing.
Large number of people who read my comment seems to agree and this whole worldcoin thing seems to me just another distraction (We've already been through why that was shady but we are talking about something different here).
Google Photos is less than 10 years old and I think a lot of people use it.
In fact, they already did. What OpenAI announced was nothing that Google could not do already.
The top comments around Sora vs Veo suggesting that Google was falling behind, given the fact that both are still unavailable to use wasn't even a point to make in the first place, but just typical HN nonsense.
I don’t think I’ve seen serious criticism of Google’s abilities. Apple didn’t release anything that Xerox or IBM couldn’t do. The difference is they didn’t.
Google’s problem has always been in product follow through. In this case, I fault them for having the sole action item be a buried waitlist request and two new brands (Veo and VideoFX) for one unreleased product.
Serious or not, that criticism existed on HN - and still does. I've seen many comments claiming Google has "fallen behind" on AI, sometimes with the insinuation the Google won't ever catch up due to OpenAI's apparent insurmountable lead
Google is large enough to not care about small opportunities. It ends up focusing on bigger opportunities that only it can execute well. Google's ability to shut down products that dont work is an insult to user but a very good corporate strategy and they deserve kudos for that.
Now, coming back to the "follow through". Google Search, Gmail, Chrome, Android, Photos, Drive, Cloud etc. all are excellent examples of Google's long term commitment to the product and constantly making things better and keeping them relevant for the market. Many companies like Yahoo! had a head start but could not keep up with their mail service.
Sure it has shut down many small products but that is because they were unlikely to turn into bigger opportunities. They often integrated the best aspect of those products into their other well established products such as Google Trips became part of search and Google Shopping became part of search.
Do you have any examples of something they launched in the last decade?
that result in shittier products overall. For example, just a few months ago they cut 17 features from Google Assistant because they couldn't monetize them, sorry, because these were "small opportunities": https://techcrunch.com/2024/01/11/google-is-removing-17-unde...
> all are excellent examples of Google's long term commitment to the product and constantly making things better and keeping them relevant for the market.
And here's a long list of excellent examples of Google killing products right and left because small opportunities or something: https://killedbygoogle.com/
And don't get me started on the whole Hangouts/Meet/Alo/Duo/whatever fiasco
> Sure it has shut down many small products but that is because they were unlikely to turn into bigger opportunities.
Translation: because they couldn't find ways to monetize the last cent out of them
---
Edit: don't forget: The absolute vast majority of Google's money comes from selling ads. There's nothing else it is capable of doing at any significant scale. The only reason it doesn't "chase small opportunities" is because Google doesn't know how. There are a few smaller cash cows that it can keep chugging along, but they are dwarfed by the single driving force that mars everything at Google: the need to sell more and more ads and monetize the shit out of everything.
Where did SORA get all its training videos from again and why won't the executives answer a simple Yes/No question to "Did you scrape Youtube to train SORA?"
Google attorneys want to know.
Wait, really? Could you point to proof for this? I'm very curious where this is coming from
In terms of software that's actually been released Google is still at best in third place when it comes to AI products.
I don't care what they can demo, I care what they've shipped. So far the only thing they've shipped for Veo is a waitlist.
I always found HN contrarian but as I say it’s really tiring. I’ve no idea what the negative commenters are working on on a daily basis to be so dismissive of everybody else’s work, including work that leaves 90% of the population in a combination of awe and fear. Also people sometimes forget that behind big corp names there are actual people. People who might be reading this thread.
...why am I feeling to urge to point out that I am only making a joke here and not trying to make an actual counter point, even if one can be made...?
If you’re negative and you get it wrong, nobody cares, get it and right you look like a damn genius. Conversely, if you’re positive and get it wrong, you look like an idiot and if you’re right you’re praised for a good call once. The rational “game theory” choice is to predict calamity.
I am still talking to a lot of people who say, “what can any of this AI stuff even do?” It’s like, robots you could hold a conversation with effectively didn’t exist 3 years ago and you’re already upset that it’s not a money tree?
I think that peoples expectation horizon narrowing down may be the clearest evidence that we’re in the singularity.
From my own perspective the critique is usually a counter balance to extreme hype, so maybe let's just agree it's ok to have both types of comments, you know "checks and balances".
Or why the existence of the UK hasn't, since they have a lot of English speaking programmers paid in peanuts.
Not saying the latter was never true, it's just interesting to see how people have reframed their work in the wake of breakneck AI progress.
If you’re cynical and you get it right that everything “sucks” you look like a genius, if you get it wrong there is no penalty.
If you aren’t cynical and you talk about how great something is going to be and it flops you look like an idiot. The social penalty is much higher.
I don’t know if this is true.
>It diminishes human value/creativity
I don’t see this at all, I see it as enhancing creativity and human value.
>and will be owned and controlled by the wealthiest people.
There are a lot of open source models being created, even if they are being released by Meta…
>It's not like the horse being replaced by the tractor. This time it's different there is no place to move to but doing nothing on a UBI (best case).
So, like, you wouldn’t do anything if you could just chill on UBI all day? If anything I’d get more creative.
> That same power also opens the door to dystopian levels of censorship and surveillance.
I don’t disagree with this at all, but I think we can fight back here and overcome this, but we have to lean into the tech to do that.
> I see more of the Black Mirror scenarios coming true rather than breakthroughs that benefit society.
I think this is basically wrong historically. Things are very seldom permanently dystopian if they’re dystopian at all. Things are demonstrably better than they were 100 years ago, and if you think back even a couple decades things are often a lot better.
The medical applications alone will save a lot of lives.
> Nobody is denying that it's impressive but the question is more whether it's good overall. Unfortunately the toothpaste seems to be out of the tube.
There are going to be annoyances, but I would bet serious cash that things continue to get better.
There is a lot of empirical research on UBI and all of it shows that it has very little effect on employment either way. That is, nothing will change here.
(This is probably because 1. positional goods exist 2. romantic prospects don't like it when you're unemployed even if you're rich.)
"When you go to an art gallery, you are simply a tourist looking at the trophy cabinet of a few millionaires" - Banksy
There's genuinely impressive progress being made, but there are also a lot of new models coming out promising way more than they can deliver. Even the Google AI announcements, which used to be carefully tailored to keep expectations low and show off their own limitations, now read more like marketing puff pieces.
I'm sure a lot of the HN crowd likes to pretend we're all perfectly discerning arbiters of the tech future with our thumbs on the pulse of the times or whatever, but realistically nobody is going to sift through a mountain of announcements ranging from "states it's revolutionary, is marginal improvement" to "states it's revolutionary, is merely an impressive step" to "states it's revolutionary, is bullshit" without resorting to vibes-based analysis.
Companies can either get peopled hyped or have never-ending georestricted waitlists, they can't have their cake and eat it too.
What is a user interface which can move from color, light, shade, and space to images or text? Could there be an architecture that takes blueprints and produces text or images?
ChatGPT already allows this workflow to some extent. You should try it out. I just talked to ChatGPT on my phone to test it. I think I will not go back to text for these purposes. It's much more creative to just say what you don't like about a picture.
If you speech is also affected rough sketches and other interfaces will/are also be available (see https://openart.ai/apps/sketch-to-image). What kind of expression do you prefer?
If you use tablets or screens, I would imagine a two screen/tablet setup, where on one screen there is a variant gallery with AI output and on the other screen there is the drawing area. The drawing constantly refreshes the gallery.
One can click on images in the gallery to move the whole image or parts of it into the drawing area. Additionally voice input leads to a conversation in the background that affects the variants as well. The process would be a mix of sketching, overpainting and voice-controlled image manipulation.
Automatic image segmentation that is automatically applied to all variants would make it easy to move objects/parts from the variants easily. The pulled parts would be stitched automatically into the drawing area, as some kind of super charged collage technique.
Maybe the variant gallery would be more like an idea board. You would say things like: "Can you make a variant with clinkers", "Please add garden furniture near the pond." etc. In the gallery these images would pop up and you can pick what you like from it.
“The engineers of the future will be poets.”
It is not a given that everyone can or should be enabled to do everything possible at any cost; people in wheelchairs can't be firefighters and we don't make all old subway lines fully accessible because it is incredibly expensive.
Disadvantaging a huge number of people for the benefit of very few has a societal cost.
You can feed AI an image and ask it to describe. Kind of the inverse process.
Do you have a source for this stat? I can't seem to find anything to support it.
> Sign up to try VisionFX
Is it Veo or VisionFX? Is it a sign up, a trial, or a waitlist?
How hard can it be to write a clear message? In the words of Don Miller, if you confuse, you lose.
This landing page feels as haphazardly put together as the Coinbase downtime page last night.
Kidding... I signed up for the waitlist. I have ideas for videos I'd like to use to explain things that I have no hope of creating myself.
Veo is the name of a video model. VideoFX is the name of a new experimental tool at labs.google.com, which uses Veo and lets you make videos.
Thanks for the feedback though, I see how it's confusing for users.
Imagen 3 is awesome though, generates nice logos :D
Agree this is totally confusing.
Imagine how bad the model must be if this is the best way Google can think of selling it.
Personally, I liked the above link, even as a Google skeptic, but the videos aren't helping their case.
In any case - no details on compute needed. Curious if this ever can be cheap. Even Midjourney still requires a lot.
I’m also surprised there hasn’t been some attempt at creating benchmarks for this. One example could be color accuracy.
Not to mention the whole AGI topic is forever doomed from SciFi fans, just remember what happened with that room-temperature superconductivity.
GPT-4o: out
Veo: waitlist
Admittedly this is impressive and the direct comp would be Sora, which isn't out, but sometimes the caricature is very close to the truth.
"We have no moat" swings both ways.
All my apps are updated in the App Store too.
Website also only has toggle for 3.5 and 4 with the Plus upgrade. Not sure if it's cause I'm in Canada?
Edit - it does list the new model in my app at least
The bigger issue is that 4o without the multi-modal, new speech capabilities or desktop app isn't that different to GPT-4. And those things aren't yet launched.
look at their linkedin pages, that will tell you why they are desperate
(hint: they bought OpenAI bags on the secondary market)
It's been there for three months and still isn't even close to being released and available.
Essentially Google has already caught up to OpenAI with their recent responses and it's clear that there are private OpenAI investors pushing such nonsense around Google struggling to compete.
Users: Lol it wont even tell me how to draw a picture of a human because its inappropriate.
Google flipped like a switch a few years ago. Instead of going for product quality, it seems they went full Apple Marketing and control the narrative of top social media.
I keep trying thinking: "well its Google, they will be the best right?" No, I'm at giving up on Google, they are not as powerful as I once thought... Hmm seems like a good time to get into Lobbying and Marketing...
Is it? I can't use it yet at least
I don't know what's wrong with GPT-4o, but the answers I'm getting are much worse than before yesterday. It's constantly repeating the entire content required to provide a seemingly "full" answer, but if it passes me the same but slightly modified Python code for the fifth time even if it has become irrelevant to the current conversation, it really gets on my nerves.
I had so well tuned custom instructions which worked beautifully and now it's as if it is ignoring most of them.
It's causing me frustration and really wasting my time when I have to wait for the unnecessary long answers to finish.
Here's my 10 minutes to 12:09 album debut:
Google missed this train, big time.