Gemini Omni 1.1 Flash
blog.google
blog.google
For example: https://sites.suffolk.edu/jhtl/2025/10/30/game-over-for-unau...
Software developers, unfortunately, have been convinced that they don't need to unionize, so have no collective bargaining power for dealing with situations like this.
I think too many just expect that they can get by like before because they're a 10x engineer or whatever
In my experience this is what almost everyone thinks, right up to the point where it starts happening to them.
And it isn't just an individual blind spot, organizations suffer from the same thing in the sense that many companies will be happy to automate away all their labor to avoid paying the "human tax" without giving much thought to the fact that the need for the company itself will also be automated away soon after.
If you can replace nearly all of your workers with some cheap tokens and prompts, everyone else who previously would have been a customer can replace your whole company with the same thing.
To be fair, most of us spent our entire career trying our best to automate ourselves away one way or another, and always seen that as our job description.
Some people dislike automation altogether and like the computational aspect. Most people seem to enjoy constructing virtual worlds and Rube Goldberg machines.
Computation IS automation
Inexplicably for the current moment, AI has so far actually meant I have even more to do instead of less. I expect that to change eventually... but man, is it a bit of whiplash to go from wondering if my job will be around in 10 years to starting the work day and being more backed up than ever. Doubly so since the rate of change does not seem to be very evenly distributed by tech role.
Spoiler alert: unions can help with that.
It's kind of hypocritical to change just because it's now a different industry being impacted.
https://www.computerweekly.com/news/252481961/Amazons-Whole-...
https://www.sciencedirect.com/science/article/abs/pii/S01762...
Sweden just has an "office worker" union.
https://www.unionen.se/in-english/this-is-unionen
Now, they cannot prevent offshoring or AI use or replacements. Very very rarely if ever.
But they can negotiate a large layoff and give better pay/conditions during work etc. This view that unions can prevent anything... when so many jobs have gone to east EU (all amazing and good at what they do but lets face it, it is due to pay) so yea there wont be any union preventing AI...
That’ll be somewhat detrimental to me perhaps (which isn’t certain), but that’s okay. I’d rather attempt to adapt to a changing world.
It is if it makes it worse for everyone else.
A union doesn't give you unilateral power; it just gives you a better seat at the bargaining table. Capital pools its resources to negotiate better as a single bloc; why shouldn't labor as well?
For your unique approach to the problems at hand. Yes, 90% of the code written is CRUD, but when it's not, how it's solved matters.
This is why we have grown with id Software games. They solved the problems, created new paths and methods for limited hardware for their time.
This is why we praised DivX and XviD back in the day. They solved problems again which are caused by the limits of the computers.
The way one solves problems matter, makes one valuable. This is why we do research and create new ways to do things.
In short, your creativity is being sold/sprayed to masses.
Until you actually remove the need for a human, you have to pay full price for human labor. Which if you used AI is never really going to happen. Someone has to evaluate what AI creates to make sure it does what was asked. This trend of not prosecuting crimes conducted by AI cannot last forever. The CEO would be the only one arrestable as the CEO is the only person with decision power in the company.
Yes, exactly. Do not advocate for yourself and your colleagues by joining them in a united front. Your voice does not matter. You will continue to adapt to widening wage disparities.
You will accept a pay cut and increased productivity goals and be happy to continue adapting as your cost of living keeps on climbing.
localized gains, distributed loss.
a computer can never be sad or horny or have taste. it can only make product. i also think artists should be paid living wages with good working conditions.
thank you SAG-AFTRA, WGA, IATSE, AEMI and Teamsters.
So far ;-)
Neither can a record player. Nevertheless, we consume recorded music.
Who made those?
What's wrong with Automated Intelligence (TM) making things for me to enjoy?
Was your question supposed to be profound? Surely someone has explained to you that record players don't compose music.
Perhaps you could be a little more charitable and less toxic here and answer my obviously non-rhetorical question: Why shouldn't I enjoy things that were created through an automatic process? Hmm?
I guess if your goal is just to get any reaction, it was good, since it definitely stirred up annoyance in me? Unfortunately it also showed that you're fundamentally misunderstanding the arguments brought against AI media, be it willfully or accidentally. Either way, it gives me zero reason to want to discuss the topic with you.
> Was your question supposed to be profound? Surely someone has explained to you that record players don't compose music.
No, it's a stupidly simple question. GP is equivocating two things that obviously show fundamental qualitative differences, but did not give any reasons why this equivocation is supposed to make sense.
When I see comments where the only interpretation I can see has major and obvious logical gaps, I sometimes ask the obvious question in the hopes that they can show me a perspective I haven't properly encountered before. Because surely all of us are capable of recognizing how weak arguments based on "two topics look vaguely similar once you strip all nuance", when you can do the same with almost any two things?
Do you think that my goal is only to get a reaction and not to try and make any other point? I think someone who was "just" trying to get a reaction probably would just curse at you or do something else. Reactions are signals though and annoyance in particular is a very useful signal, hopeful to the one who experiences it. More importantly though, why, when you responded to me the first time, did you not address the very obvious non-rhetorical variant of my question though and instead, decide to appeal to me with sarcasm about the profundity of the rhetorical variant?
I'm trying to engage here and not put my trauma on display so I'll try to be as straightforward as I can. I am sincerely trying to respond in kind to you by telling you the truth and I will be glad to explain, in detail, my point of view. So, let me do that:
> GP is equivocating two things...but did not give any reasons why...
Yes. And I had hoped to correct that equivocation by pointing out the many random and/or automated processes that produce beautiful or otherwise enjoyable things. I wasn't agreeing with GP, I was improving the position. And I did, because very many useful things are created without any human interaction.
> fundamentally misunderstanding the arguments brought against AI media
Okay and you haven't engaged with the point that I'm making at all. It's different than the GP's. You also haven't given me any reasoning to backup your point that I am fundamentally misunderstanding the arguments brought against AI media. You want reasons why and I want reasons why. Let's engage in a civil manner.
The point that we're engaging on, really, is not the GP's. It's what computerliker said about how "a computer can never be sad or horny or have taste. it can only make product" and that "having less media [created] by a computer is a loss".
Nothing to misunderstand there and I disagree strongly on both the implications of the first and the conclusion of the second quote. That's what I'm attempting to address and not every possible argument ever made about AI creations.
So, do you still think that I'm misunderstanding something? Also, what about my point? Do we really want fewer things created by random and/or automated processes simply because a human didn't create them? I don't think so.
Unions are very much an in-group vs out-group phenomenon (and in many cases, the benefits are specifically to the more senior union members vis a vis the less senior ones.)
If you're talking about AB 1751 (Missing Middle Townhome Ownership Act) I think you have this exactly backwards.
> Unions are very much an in-group vs out-group phenomenon
As opposed to what? You and the GGP phrase this as a dichotomy, but I'm really curious what the other side of this is, because "not being in a union" doesn't erase your self interest or all the many many groups of people doing the same that don't feel like they're apparently being overly self interested by forming corporations, governments, non-profits, advocacy groups, etc, etc, etc
In the US. In the UK all the major unions I know don't require membership and fight for all employees in the industries they represent. E.g. as a software developer on a research project at a University I got all the negotiated benefits of the academic Union.
Unions can be a deterrent to this. If you have sectoral bargaining, you can’t get a better deal anywhere because everyone pays the same. Or even if it’s within a single company, you’re just stuck with whatever scheme the union bosses agreed to.
But also, unions are a deterrent to new employers since fewer people want to start companies in industries where unions have taken over. They’re a nightmare to deal with. Fewer employers means less competition means worse services.
We have a dynamic economy with lots of competition. It’s easy to switch jobs and find a better place to work. We should not want to give this up.
It was cynically selfish and short-sighted. The US from 1980–2020 is a perfect example of how much more complex systems, economies, worker dynamics, and second/third/nth order effects are.
Competition is generally good for the consumer, but Americans need to wake up to the fact that it's not built into the system, it's to be regulated into it, and competing is only good so far as it doesn't turn into a race to the bottom (which again, is a matter of a regulator setting the bar).
I don’t have any objection to voluntary unions. Meaning people can join them for support or collective bargaining, but that citizens can just opt to not employ anyone in the union. Unions could be support structures for workers, instead they are mafia bosses.
When have you not been able to get timely medical treatment at the fault of nurses ?
I read it more like they're against trying to stop, hinder, or prevent the use of new technology, simply because it's personally threatening.
And I feel the same way. If I was part of a union, I would believe in the collective bargaining power, but I wouldn't necessarily believe we should use that to pursue every single possible agenda that might benefit us. And I've personally pooled my capital with others in the past in order to achieve certain outcomes. And in those situations I also didn't attempt to reach so far as to limit the totally valid freedoms of others just because I knew it might be better for me.
Acting in the interest of a collective doesn't mean one must also abandon the ethics they usually employ when acting in their solo self interest.
That's not to say unions, or self-interest itself, are bad.
Caring about your ability to pay your bills, and support your family is "self-interested" in the same way that choosing to not drive into oncoming traffic is "self-interested".
The former is obviously unfair to the person whose identity is stolen. The latter is sour grapes from people who want to control the future.
I also can't bring myself to get incensed over someone using a robot to write code. It's literally the modern equivalent of smashing knitting looms.
Union = a group of people defending their rights and that includes the right to a life that isn't at the utter mercy of perpetual enshittification due to the infinite mechanical plodding of heartless (by definition) megacorporations that would never hesitate to eradicate peoples' livelihoods and treat humans as material for the orphan grinder if they could get away with it.
And by the way, they do, if you haven't noticed.
What is “your class of people”, and why is it threatened by what other people do with coding robots?
You can try to dress it up as an intellectual argument, but at root it’s the same thing that the luddites were trying to do: defending your preferred lifestyle from technological advances.
We, the ditch diggers of the world, support your unqualified resistance to the use of steam shovels! Fewer holes, slower!
Thank you for outing yourself as a low-empathy, capital-aligned individual who makes bad-faith arguments, I'll know not to engage you from now on.
You are part of a whole. You are not alone.
Saying you don’t mind to adapt is insulting.
It is possible. But it's quite uncommon in this industry in the US.
And given the current administration, it is hard to trust you'd get fair enforcement of labor laws if the company did illegal things to block unionization.
>California’s legislation passed AB 2602 and AB 1836 in September 2024, prohibiting media companies from using AI to replicate actors’ performances without their consent.
it's a very reasonable law, but it is not the win you seem to believe it to be. the actors who refuse to consent will simply be passed over in favor of those who don't. no law will ever be passed to force companies to employ humans over machines, and if it were, the industry would move elsewhere. this is not without precedent :)
speech models have got so good so quickly that you can already replace a VA -- even an AAA prima donna -- with a teenager from Fiverr, who will simply bruteforce the right inflection.
Seems like a reasonable take
e.g a lot of video game studios even the AAA ones could be co-ops. same as a lot of SAAS software companies. Linear - just announced a tender offer. I don't see a reason - why linear couldn't work as a co-op.
In any case, the overwhelming majority of jobs is non-unionized and I don't think that software developers can stop technological changes by unionizing.
They settled for a large undisclosed sum out of court. But largely won unfortunately as unionisation efforts failed
If your students are in Oklahoma, your teachers also have to be in Oklahoma. If the Oklahoma teachers unionize, you have no option but to hire those unionized teachers if you want your students taught. If your bank is in Oklahoma, you can hire programmers in London or buy a SaaS from a company based in Geneva.
Programmers also have much less of a reason to unionize, as "quit and go work somewhere else" is often a realistic option. This is usually not the case for teachers, where the only employer seeking their skills in the area is the government, whose idea of what teacher salaries should be is not based on market forces.
It’s brutal for these people. The creative industry was always hard, but this is just plain brutal.
I was bored yesterday and started watching a stupid documentary series about UFOs/UAP on Netflix. It was immediately obvious that the voice over was AI slop bullshit. All I could do is stop watching and thumb down it, but they still have my money. :(
And jobs where quality is paramount and or there is greater chance of risk / bodily harm
Society seems to be very bad at that, and we end up with high frequency traders making millions while teachers are paid peanuts.
It has nothing to do with skill. Both were very skilled jobs. Draftsman as well despite many people going to it straight from school. But computers and CAD mean that it is now necessary for someone to do a STEM degree to be a draftsman. Recorded audio made many stage musicians redundant. It is cheaper to do it this way and gets superior results, that is all, there is no further agenda.
Now too, the next generation of voice actors and many other knowledge workers will have to go up the value chain one step and operate or potentially build these tools (in whatever form they mature to in a decades time).
The current generation of voice actors will face the same situation as many before in the performance industry - stage musicians/performers for example that were made redundant by recorded audio. The reality is that most of them just left and dispersed into the economy doing completely unrelated jobs.
For software and generally computer engineers, this new primitive happens to itself be software, so it's less of a transition and an easier upskilling path to learn to build it. And building it is one step higher in the value chain than simply using it. That is a structural advantage.
But now that we are automating white collar work... where will people go? I'm a 36 year-old veteran who has returned to college and so many of the younger students seem to be filled with despair.
I have hope for the future, but I think there will be an uncomfortable period of time.
This slow transition period will also help answer the "where will they go" questions. We can't answer them right now.
Ultimately, everything we do is in service of social political and personal human incentives, and I think the effect of that is discounted when people make these takeoff predictions for AI and "AGI"
All these demos are focused on how easy it makes everything. Easy is great, but if everyone is able to make instant cute cat videos or whatever it just devalues it. I want to see turning a photo into a rigged 3d model, letting the artist animate and then generate the video. This technology could be used to increase creative expression, but instead it's being used to squeeze out creative expression
There could easily be at least some time period of low skilled ugly people acting in approximate but shitty ways in cheap sets just to give an input reference to a model and then describing the differences in text, yielding gorgeous people speaking with prestigious accents doing stuff in fancy locations in the output.
The "slop" phase is early AI, like chunky ugly very early 3D graphics where you can see the triangles.
I'm already seeing images generated by first-generation diffusion models like Stable Diffusion 1.5 (which is small enough to run on a phone) used ironically as meme generators due to the now-retro silliness of what they generate.
one youtuber used irl footage, with hands and stuff - so I know there's human behind the camera
the other was a letsplay that reacted to events just fine emotionally
and yet the uncanny valley of the sound is in full force. Maybe youtube has done something with the codecs?
I feel sad
* https://i.ibb.co/ccqKZ71L/keenlore.png
I don't know how the next generation of beloved actors comes about and how we don't descend into a pit of neverending photocopies of things people once loved in the 1990s/2000s.
* $250 per finished hour for the narrator (lowest professional rate).
* 200k to 270k words.
* 22 hours, 20 hours, and 25 hours.
* Books 1 to 3 cost $5,500, $5,000, and $6,250, respectively.
My novel has 8 major characters, including the 3 narrators, and 25+ minor characters. That price tag is daunting, presumably in USD. This would mean investing $7,750 CAD in crafting an audiobook, which may not even sell, much less recoup the investment.
In contrast, a locally hosted solution has a wallet cost of pennies for electricity plus my time to develop the system.
To me it's been obvious for a while now: there won't be one
it always sucks at everything else.
Articles and discussions pop up from time here: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
>Firefox makes up about 90 percent of Mozilla’s revenue, according to Muhlheim, the finance chief for the organization’s for-profit arm — which in turn helps fund the nonprofit Mozilla Foundation. About 85 percent of that revenue comes from its deal with Google, he added.
https://www.theverge.com/news/660548/firefox-google-search-r...
Billion with a B, as in 10 zeroes, not 7.
If anything FF gets left out because usage is so low.
I've been following this for a long time. They were leaving Firefox out when its usage wasn't low.
Probably the typical backdoor executive mandate that led to death by "sprint prioritization":
Yes, we will for sure work on the Firefox compatibility bug, Dave-Open-Source-Enthusiast-Google-Dev.
But we can only pick up 10 bugfixing tickets this sprint and the ticket you highlighted, as the entire team agrees, is priority #12.
<repeat every sprint, where during the sprint 9-10 new higher priority items magically appear just in time for the next sprint>
Death by slow asphyxiation.
The page instant scrolled suddenly back to previous video examples after I scrolled down. It was disorienting.
Maybe because they see video generation as key to developing "world models"?
And, Veo and omni simply were better than Sora
And, Youtube is huge both as a place where video contents goes and where can be trained from. Microdramas are starting to become a real category--14 Billion USD, 90% of it made with AI.
Chinese video models can be more immediately impressive, but none of them come close to beat the value of Google's Flow. Especially when you are throwing away a lot of generations as part of the creative process. Which is what you have to do to make longer content with any video model.
OpenAI needed to be able to focus. Google can walk and chew gum, and they're not going to run out of money to buy chewing gum.
Almost all concrete english-language info I can find about microdramas is astroturfed to hell by consultants and "independent" industry publications. Wikipedia's citations for 2025 revenue are 'Reel Reel' and 'Duanju News'. Neither cite their source, though they are likely just regurgitating predictions from Omdia, a media consultancy.
> Real Reel™ works with companies across entertainment, technology and the creator economy to build relevant industry conversations around mobile-first storytelling.
Duanju News is published by 'Studio Phocéen', who run their own microdrama production house.
The Omdia revenue estimates, which is where the projected $14 billion comes from, are completely unsubstantiated as far as I can tell. They even go so far as to make "according to new research" a link that when followed sends you to their generic "/advance-your-business/media-and-entertainment" sales pitch.
https://omdia.tech.informa.com/pr/2025/oct/microdramas-to-ge...
I don't have any special insights here, it all just smells a bit like McKinsey's "The metaverse will be worth $5 trillion by 2030, you better not miss out!!! Hire our 22 year old slide deck experts today"
Previously when making video ads you'd need to actually create the video. Actors, cameramen, editors - you name it. Now a new video ads is just a prompt away, directly inside the ad-spend web UI too no doubt.
People say Google have lost and that they're having their lunch eaten by anthropic, but I am not so sure...
Google is selling shovels, leasing mines, buy stakes in "competitors" and doing it's own exploration/mining. When you look at the full picture, it kinda doesn't even look like Gemini matters that much to them overall.
Gemma 4 I am way more impressed by than whatever the top ranking model is
That's old news. People are saying the chinese models/labs are eating anthropic lunch.
I never personally got the least bit excited about Sora or nanobanana or whatever video/audio generation thing. But I guess I'm just not their customer. I do love the read-side of it though.
I work in advertising and some days I spent a lot of money using these models. The amount and rapidity of prototyping using them has changed everything about advertising pre production.
Quote under the video of a short Argentinian footballer wearing no 10 with "RESSC" on his back. Can't make it up.
Pro models are mainly for coding agent work; it doesn't necessarily make them any money.
The capital infusion the frontier labs have received has gotten to a size where many believe it may not be possible to recoup this investment without some very unrealistic things happening.
I think it's reasonable to not completely drain one's cash reserves trying to stay ahead in a race where participants may very clearly be about to run straight off of a cliff.
For example, PDF's and powerpoints can be generated using agentic coding by things like https://bento.page or other ways of generating them in an agentic coding fashion.
A lot of browser automation could/is also done by agentic coding.
It can also help them set up and configure self hosted software with the help of LLM's and debugging if its working or not.
You can create videos using Manim and remotion.dev and also excalidraw-animate and generate excalidraw files agentically if what you need is more vector style graphics (which surprisingly can fit into many ideas) rather than say a real life human waving video/more photo-realistic video (but I must say that this has certainly its own pros/use-cases as well).
It might sound self-explainatory but turns out that coding can represent a wide range of problems!
I personally wish to get more smaller models (like the recent qwen model) and other open source models like GLM 5.3 and the glm flash model.
> No one is suddenly missing out on some giant competitive edge because their model is a few months behind
Sure I can agree with that. The competitive edge might still exist but I do get the underlying sense of what you are trying to suggest.
> things other than how well your model can write code will matter more and more in 12 to 24 months.
What are the things then which you feel like could be more differentiative factor? For example, I personally think multi modal is still quite preferrable in AI models. I use GLM 5.2 and it doesn't have vision and I can certainly imagine time/use-cases where multi-modality would've helped coding and even other use cases as well. So what are some other use cases that you are thinking? Video generation models like Veo/Sora?
Sure downside would be not learning from people using your model for coding, if we're on the cusp of huge leaps in self-improvement. But there is a reasonable case for avoiding desperate scramble, especially if other parts of the business can also create value with the compute.
They never gave an official answer as to why, so I'll let you draw your own conclusions.
They did not decide it wasn't worth spending the money to train.
They absolutely spent the money.
Also, look at Flash 3.5 to 3.7. Flash 3.7 is a genuinely decent Sonnet 5 class model. Flash 3.7 is quite efficient too. Also, whatever was spent training 3.5 pro is probably not wasted. However, as a strategy, when I see models like Kimi K3, Fable, Sol. If you discard "because the model sucked" what other alternatives or potential options might exist?
I thought of a quite a few and they are far more compelling and interesting to me.
(Also Gemini models tend to be pretty decent at more than just programming. Enterprise AI use is more than just software eng / programming)
If you'd told me at the end of Cloud Next 2025 that by now Google still wouldn't have a competitive offering to agentic coding offerings from Anthropic (Claude Code + Fable) or OpenAI (Codex + Sol), I wouldn't have believed you.
In our non-coding use cases where we're embedding models in our product, we're also not reaching for GCP stuff. Because Anthropic has the mindshare of our engineers and product folks, since it's what they use every day.
They mentioned that they have already started pretraining Gemini 4, which will be the full ground up rip-your-face-off-expensive training that is often discussed.
Version 3.1 has plenty of room for improvement, yet they don't seem to be giving the attention it deserves or at least communicating accordingly.
It may be that they wish to slow their cadence of releases, or develop their models to focus more in a different direction, etc. No matter what the actual reasoning, they have chosen to not compete in the same race, and I cannot say I fault them.
In the real world out there, Google and Microsoft are absolutely dominating enterprise customers.
Every single non-tech office worker I know is writing Gemini "gems" (sort of claude prompts/skills) or prompting Copilot to help drafting board meeting notes, insurance contracts updates that reflect changes in regulations, make quick loan feasibility assessments before passing them to the relevant office, presentations, etc, etc.
I'm talking insurance, banking, consultancy, manufacturing, etc, etc.
Why? Because Google and Microsoft already were in these companies, all they had to do is "oh, you also have AI now with your plans". Procurement and data compliance are the first thing businesses have to sort out. They were already sorted out.
Google doesn't need to have the best coding model or triumph in meaningless benchmarks, it only needs their models to get better and cheaper while serving them to their existing customer base.
They are playing a different game.
And Microsoft, doesn't even need to care about models at all, they can provide whatever open or closed AI with their services and have to focus on the harness in Excel or Github/Azure Copilot or whatever.
E.g. while developers in most of my clients use whatever they prefer or the company pays for, the remaining 90% uses either Google or Microsoft products.
Not a single one has incentives into venturing into OpenAI or Anthropic or Z.Ai lands because they might be better at some benchmark that is completely irrelevant to their tasks of updating powerpoints or summarizing incoming emails.
Google is an advertising company with an enterprise SaaS branch. I bet they have chosen to focus on running the most efficient "everybody" model instead of running a heavy model for coders.
When using Gemini for other tasks than coding, it is actually pretty good. It grounds well with Google search and gives mostly correct, well written answers to many niche questions.
I feel like I should be excited about being able to generate almost perfect videos but, I just don't care anymore.
Meanwhile, I'm happily using Minimax H3 locally on my 12Gb 4070RTX to finally finish the lip syncing to recorded dialog on my abandoned 20 year old Flash animation hobby projects.
I agree though. My issue is the cost for using AI video models is way too high for anyone not building anything serious with them, at the same time they are too restricted for actually using professionally. Prompting them with text to get something generated is cute, but then you just end up creating slop that everyone hates, ultimately devaluing the power of these things.
Minimax H3 is about 240Gb alone, how do you do? How much quantised is it, and how good are the results?
Most people are running the stock release of Minimax H3 using the INT8 quant and it's about ~20GB.
https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/diffus...
Generate lightweight previews in 360p resolution up to 60% faster
and at a third of the cost compared to Omni 1.1’s standard 720p resolution. This is helpful for rapid prototyping, storyboard iteration, and quick rendering in developer platforms."
This is a great idea, to have a low-resolution mode for additional speed to create previews, do test runs, create rapid prototypes, etc.
My curiousity is, what's the absolute useable minimum that this could be?
That is, would/could 240p resolution work? If so, what about 144p? How about lower? Then, could those images be upscaled quickly (and is the result still usable?) with a faster image upscaling-only neural network?
The reason why knowing such lower numbers / lower bounds -- is because they could be important for additional cost/time savings and/or running derived LLM's on local resource-constrained hardware...
Anyway, great post, great idea, and we welcome Gemnini Omni 1.1 Flash to the ever-expanding list of LLM/AI's!
The amazing videos and photos carried the promise that I could experience that for real. They were aspirational.
Today I suspect every cool shot is made of pixels arranged on a 2D screen by an algorithm. It doesn't do it for me.
That makes me sad...
Why do we look at art, watch videos/movies? Is that replicable as a function of text, other existing media, and 3-30 cents of compute per second? I'm pretty functionalist about these things, and at some point it probably won't be possible to tell the difference. But until then, at which point we might just say 'death of the author', it seems like a category error.
I do work with artists that use video and image generation models to create stuff, but from what I can tell they're interested in faster iteration and controlling a lot of intermediate steps (their graphs can get pretty labyrinthine).
I enjoy making short films with AI. When my latest short screens at a festival in Ocotber, alongside traditional and AI films, hopefully the audience will like it too.
The quick "one shot" video generation might be slop to you, or I. But if someone wants to send it as birthday greeting to their aunt, and they both enjoy it, what business is it of ours?
I have a young boy, and whenever he builds an impressive "scene" from LEGO (like a diorama or whatever), I take a couple of reference pictures with my phone and make it into a "real" movie scene, cartoon, or whatever. He loves it, and this motivates him to build more and bigger things out of LEGO.
If he builds something really special, I might actually fork over the $5 to use Omni to turn his LEGO creation into a 10-second video instead of a still image. It'll blow his mind!
PS: There also are cheap and even free phone apps that make stop-motion animation trivial. We've already made a couple of videos of his toys moving around that way.
I like to generate songs from obscure poems
I have been pondering this and have come out with a single statement. It allows you to create the missing piece in the art you want to create. Assets for games, music for the lyrics you wrote, or just a whole song to justify a crazy dance you want to do. Whenever AI art is discussed, people tend to be so purist about art. One person cannot do everything in a project they take on to express themselves. Usually people come back with then get someone else to do it for you. I think that is an economic argument as people are trying to protect artist pay. I get that. But is it fair to just not let something they want to create because they don't posses every talent needed to do this?
While it sounds great you're quickly disappointed after you run the same prompt at standard resolution only to get a different result because it's non deterministic.
Their changelog suggests using their new upscaler but I've had nothing but disappointment from upscalers.
Upscale when ready: Users on paid tiers can seamlessly upgrade their favorite clips using our new 360p to 720p upscaler.
"Omni" -> everwhere/all
"The final chapter will end with a choice. The choice between holding on to the past, or letting go. What will you choose?"
When it comes to fish swimming around I don't think I would be able to reliably tell what was real vs generated even with deep inspection.
I mean I have not tinkered that much, but trying to even get a video to 30 seconds (I just want a cartoon AI avatar to narrate tutorials) has been incredibly difficult. They drift so easily.
Many AI videos you can tell just stitch short clips together. I just want a continuous scene for like 30 seconds to 2 minutes.
x There are so-called “SOTA” or “frontier” models that are more effective than the other ones (independent of harnessing and routing)
x OpenAI and Anthropic have all the SOTA models and lead all the innovation
x Google’s moat is its search bread/butter (it’s the only reason they’re relevant)
All 3 operating assumptions are - I think - false.
What Google has done that the “cuter products” (Claude, ChatGPT) haven’t is connected relatively standard LLMs to an externally valuable live service.
As more companies realize that is where all the value is (the service) and not in the AI capability, then products (and humans) become important again.
Google should just be Google again, and Gemini should be Gemini, off to the side. Omni confuses everyone (and angers some iykyk), they should resolve “AI mode”, rename Gemma? and consolidate the brand overall so it’s clear what Google is.
Google is search.
It helps people on all sides of the market find what they’re looking for.
I don’t really see how repeatedly reinventing and rebranding the same AI chat UX is accomplishing anything toward that goal.
Implicit to that is "find". Their AI integration into search has really hit its stride for me. They have that search box (or speech prompt) hard wired into people and they are finally iterating and crafting AI into that experience. They really failed hard initially.
I know others have worse experiences than me but Google knows a lot about me so maybe that affects my results. YMMV
It does. Very commonly it will get acronyms wrong or assume I mean something else, even when it should be in my “ad profile” or whatever it bases it off.
Often I think the search engine is working perfectly then the LLM is ruining it in delivery.
There are other UIs besides chat
You say "ruining the delivery". I believe I understand. It is not traditional www text search.
I do have a problem with most chat ai's that always ask a question at the end of their answer. This is where configurability is more important to me than absolute performance. e.g. "don't end an answer with a question unless it is directly relevant to the current conversation.
Yes there are many other UIs besides chat. The default chat UI really bothered me at first. Seemed cheesy. However the majority of human interaction is "chat".
Something that is really really cool in this tech is the vision and audio and text models that naturally cluster similar things in massive multi-dimensional arrays. I have a hard time mapping nut~squirrel~food~imageofnut~dogsaspredators~whateves. How does that condense?
We have come so far on the audio transcription dimension. In this discussion context yes it is still chat but we've moved to a different place in the brain. Not radically different (humans without visual sight come to mind in how the process goes from photons to perception (also see hank green's vid on how the eyes are part of the brain)).
Anyway hope this isn't a throwaway account. That username doesn't inspire confidence in that regard.
What the hell am I doing?
[edit: made things more clear about my un-clear thoughts! I did not shift things in the intent space]