Gemini 3.8 Flash and 3.8 Flash Cyber
blog.google
blog.google
Here's what I got for 1.8 cents and 13 seconds from the prompt "make me a cool thing in html":
https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f
Transcript here (it was part of a chat): https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992...
Since this transcript has HTML in it, I decided to upgrade that tool to also render HTML.
I set Gemini 3.8 Flash the task, using my own VERY shonky coding agent tool (llm-coding-agent) - and it did a solid job.
So now you can see the "cool thing in html" rendered within the Markdown document using code that Gemini 3.8 Flash also wrote: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Transcript where it built that is here: https://gist.github.com/simonw/3e36b98292dfdc1b3baff158faa74...
I noticed it felt a little janky on my PC despite being "60 FPS"...then I noticed the "60 FPS" is hard-coded into the HTML.
I find Ling 3.0 tiny particularly interesting as it looks really nice for a tiny model with 7.9B total parameters, with only 1.3B parameters activated per token. Here is the result https://coolthing-ling-3-tiny.tiiny.site (sorry for the weird hosting, first I found that worked)
(it cost me almost 0 cents and done in 49 seconds)
> Aside from reading identically forwards and backwards down to the letter
No it doesn't.
Thought processs: "Oh, simonw is asking me to make something cool, I think I know what he really wants..."
i don't know if Gemini models per se are fully is in line with that purpose, but the results we see keep seeming to be in-line with that split-of-focus.
I would hope the people who make one of the most used JS engines in the world are capable of making a model good at JavaScript ;)
For comparison's sake, I tried something similar with a couple other cheap models I've used lately, with the prompt "Impress me. Make something cool in HTML. Ensure that it is mobile friendly." (Added the mobile condition as I was on my phone when I did it).
Mimo-2.5 created something similar, only a bit less complex than Gemini's (though, at least the FPS counter is real!), in a minute or two for about 1/3 of a cent: https://gisthost.github.io/?740c325c21e9bfbee59c4f94d9aab0af
GLM-5.3-Flash, currently my workhorse model, spent 12 minutes (ouch) thinking about the prompt. Didn't cost me anything directly because I have a GLM sub, but I did the math and it would have cost about 1.1 cents through the API. Turned out nicely in my opinion (though in reality, it still isn't really anything special): https://gisthost.github.io/?9ef050e16cec2561e6504e725a3f0bcc
Side note: thanks for setting up that Gist Host tool, it's very convenient!
---
Editing to add this bonus from Mercury-2.5-Preview, which I just learned released a couple days ago. It's much less impressive-looking than any of the above, but it cost less than 1/20th of a cent, and the response was generated effectively instantly: https://gisthost.github.io/?02f40b50aa891bf396bfaaa3a7998203
User: use them both
Made me giggle.
- Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order.
- Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the view from it.
- Document parsing (extracting the relevant trip info from PDFs).
If you use LLMs for anything other than coding, I definitely recommend not discounting Gemini like I did just because other models are more popular.
Beginning to think Google is a dark horse in this race and some of Anthropic's "everything feels janky and rushed" karma is going to catch up.
For awhile now I've found Gemini will use Google search for pretty much any real world knowledge, which is a huge plus IMO. It's basically Google with a much better frontend and no ads/seo nonsense.
Why you would rely on the model's weights to know opening hours, instead of having the model call a web search tool to verify it on the official site?
Never had a bad suggestion.
My only wish is it were somewhat cheaper, as it tends to balloon pretty quickly when I'm using it in Opencode. I'm currently trying to offload a lot of work to subagents to stop the context expanding so rapidly. But on the upside, I rarely have to correct it - I've spent far less time arguing with this than with anything else so far.
this has to be stong suit of ai agents any model
Think 2023 style ChatGPT. Something like “to open a document on your Mac click File > Open docurrrar” - like it suddenly forgot it had to produce actual words.
Overall I enjoyed its speed and comprehensiveness. But those occurrences of nonsense just made it feel like a great car that once a month just stops in the middle of the highway.
https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium!
Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.
With a score of 59, Gemini 3.8 Flash is in eighth place, falling behind even Grok 4.6, Kimi k3, and GLM 5.3.
Further, opus 5 medium outputs 4x fewer tokens to achieve the same result, negating a lot of the speed difference.
It's almost across the board better than Terra at less than half the price. 3.9 is likely to approach Sol at the 1/10th the price.
Hopefully OpenAI releases Astra first, and it's not only better than Sol but significantly cheaper, too.
The result is the same as the previous Gemini 3.6/3.7 Flash days: Claude could always note much more problems in Codex's plan and implementation than Gemini could - the ratio is like 10:1.
I occasionally switch the roles between Codex and Claude, and result is the same, Codex could always catch much more problems in Claude's plan and implementation, than Gemini could.
So I am guessing in a relatedly complex codebase, Gemini is much less effective in acting as a guardrail (or a senior engineer/team lead) than the other SOTA models.
...on Medium reasoning. Claude Opus 5 (high) is the default in e.g. Claude Code and scores 61. Still very impressive.
For example, it's not even close to Opus 5 on Terminal-bench 4.0, 19.1% vs. 51.8%.
Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents
(I think thinking level low is a regression on 3.8 compared to 3.7.)
> https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
> Took just under 14 minutes to generate, and at 65927 output tokens cost me a hefty $3.30!
So 50x cheaper - and how much faster?
Edit: scrolled down to medium effort, its better but also has a weird clipping issue with the fish in the beak.
Gemini Flash is also pretty cheap, so it's a great family for performing media analysis, like extracting structured data from images and video.
https://blog.google/innovation-and-ai/models-and-research/ge...
We transcode everything to 480p before we send it to Gemini batch api. Works great
True multimodal support would be way better, but I have no issues pasting in full screen recordings while QA'ing games and having Claude identify and fix issues in the video.
These sort of fast and cheap models are great for tasks that are verifiable and can be retried infinitely (like coding), you can basically get frontier results with a good harness (at a fraction of the time and money).
I might be wrong about this, but obviously Google would like to provide inferencing at the lowest cost to themselves, so perhaps their slow ‘pro’ releases and rapid ‘flash’ releases is an attempt to guide people to use more profitable models?
Lots of models seem to just allow the model to "bloatmax" tokens in order to get bumps at high/max reasoning levels. Many of the max reasoning levels allow models to use up to double or more the tokens the next lowest reasoning level uses. Its basically only useful for people who have no cost or time stipulations on anything.
I think I actually preferred it when we had models that either had reasoning enabled or didn't.
I eagerly wait more info but sounds like Deepmind without Demis calling the shots has been unleashed and are operating at full speed? Shocker!
At this point it is a meme of course, but where is 3.5 Pro :)
Kind of wild that they haven't (successfully) pretrained a base model since Jan-25.
But, also... Sol crushes Flash 3.7 at writing code in a codebase of any size beyond "tiny".
Flash is my go-to for prototyping, and basically anything that isn't writing production code.
I have been testing 3.7 flash against 3.5 flash and it seems to lose every time in overall latency. Every benchmark I've seen seems to suggest the opposite[1] - that 3.7 flash is significantly (at times 2x) faster than 3.5 flash - but I have never been able to prove this out in real world use cases.
Has anyone found their latency numbers to actually be accurate? Is this why they've toned it down in this release? For context, I'm testing larger generation payloads that take 8-10 seconds in 3.5 flash and 15-25 seconds in 3.7 flash. Lowest reasoning settings in both cases.
1: https://artificialanalysis.ai/?speed=intelligence-vs-speed&m...
One aspect of model releases that don't get discussed as much are the cache invalidation (changes in underlying architecture, weights, or tokenizers); I assess Google seems to be squeezing the maximum out of the last 'Pro' version they released with 3.1 back in February.
Small models cataching up with their bigger siblings are fantastic news.
A SOC/IR or AppSec team doesn't need a generalized model that knows when Chaucer lived but it absolutely needs a model that can efficiently, quickly, and accurately prioritize vulnerability severity or validate patches.
Unless they have an even more powerful Gemini Pro in the oven...?
3.0 flash -> 3.8 flash is all post training which is pretty impressive.
I always enjoy interacting with Gemini . When i ask it questions about designing a new model etc - it's always the most helpful and encouraging . I really want to thank the Google team for this and their happy positive models they generate.
I use gemini flash after a round of deliberation between Sol and Kimi these days on the main plan . Kimi 2.7 paired with Gemini 3.5/4.6/3.7 flash has been my implementer - Kimi k3 and Sol 5.6 have been my planners and code reviewers.
I would have ideally like Anthropic and was on their USD 200 plan - but after they didnt sign the letter for Open source models - i dropped my subscription . Wont make a difference to their lives.
But Sol is great . Combined with Kimi K3 for adversarial plan reviews - you get robust plans . And Sol as a reviewer for Gemini/kimi 2.7 - you get great edge case handling and robust code.
I also integrated muse 1.2 - on the contributor tier - and I Just saw facebook release 1.3 muse . this is great news. The training im permitting is my thanks to FB for releasing the open source models of the past ! Thank you !
I've complained plenty on here about Claude verbosity and TED-talk phrasing, and it seems by contrast Gemini has already arrived at the dream end-state of Claude from a prose standpoint.
Sometimes I ask for feedback, and I get back a list of multiple-choice options as if it's already ready to go. If I indicate I'm thinking about doing something, sometimes it'll just...do it. (Not in an annoying way.)
It seems very geared toward action in a way that's completely refreshing coming from months steeped in Claude essays.
My experience is that antigravity is awful and reckless - but that the model itself isn't.
"available to trusted defenders through our new Fairwind Program"
Then why even bother announcing this? Ordinary people can use K3 and GLM 5.3 or whatever drops next and avoid all this hassle."Valgrind is only available to trusted defenders in our new UnfairAdvantage program"
Is this weakness in their training regimen the impact of operating under regulatory frameworks for too long?
And even when searching for internet, it still cannot suggest a up-to-date approach to the problem.
For example I'm using crystal, it recently revamped the concurrency/parallel model. Even using web search, gemini still does not aware of the new feature and still give the outdated code.
I'm sure my crystal usage is not the unique case here.
Re: Chinese models, even if the model itself isn't censored, some of the big model providers have guardrails now that you can't exceed, which somewhat defeats the purpose.
If I had to pay per token I would probably consider using this (they seem to be on the pareto of performance) but not being able to use opencode with a subscription is not really something I'm realistically going to do when claude and codex are around. Also never gotten along well with gemini-cli / antigravity-cli.
I also used it for a an app for my Garmin watch, and it wasn't good. The code was compiling, but functionality was totally broken and even with a lot of steering it wasn't able to make it work. GLM 5.3-flash instead was up for it and the code wasn't bad at all. I am curious to see if 3.8 is an improvement in this use case.
I tested Gemini CLI while ago, and it was awful tbh.
- Flash-Lite
- 3.6 Flash [new]
- 3.1 Pro
The above is why i don't use LLM products from Google. If the model is not available right this minute (heck, hours before the release!), then I'm not gonna bother getting back to it tomorrow, because tomorrow I'll be playing with the new model from OAI/Anthropic.
(It also shows that the internet isn't dead. Even people who are not aware of Google AI Studio can express their valuable opinions on LLMs!)
So frustrating and confusing.
Meanwhile Anthropic and OpenAI simply release a model everywhere (Fable on Pro only as a somewhat mild exception).
But yeah, they really dgaf about gemini.google.com -- I dropped that sub in April when it was clear OAI and Anthropic had lapped them
I have a weird vibe from all the comments in this thread, they feel like a script rather a real experience.
It's available in antigravity which I started using again (for small things until I can trust gemini for coding again).
(I guess the only relevance here is that if your problem matches a model's strengths, then you can do fine with a model that is several generations out of date.)
Do they make it intentionally to look previous model less superior than current models? or these are the real numbers when re-ran the benchmark..?
(By high volume I mean things like "main app just updated with XYZ commits, please scan XYZ plugins and surface any compatibility issues")
3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash
Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...
3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash
Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...
Flash 3.8 seems like where I can specify Flash3.8 as the coding model as part of agent workflow.
The video recognition is especially impressive as they got all of Youtube to train from.
- Def people who has to queue video recognition jobs to use the model.
Serving more models also adds a significant ops burden on the SREs and trust& safety teams.
And in my other app I was debugging and using OpenAI to optimize some path it cut me off numerous times because it did not like JIT functionality (this is my commercial business rule evaluation engine that compiles rules to executable code inside the app to increase performance using asmjit library)
I am basically paying for them to waste my tokens and time on these 2 tasks
In all seriousness, gemini has the best interactive planning document/orchestration. Tell it to create a plan document and work through it with it and it will preform really well(in antigravity products). But this is the case with plan modes with every model, I just think the interactive document that antigravity uses is really well thought out.
Yesterday I asked for food stop on my road trip 45 minutes from the current time and it gave me some options, but then I changed my mind and specifically asked for Asian restaurants and it completely forgot about the 45 minutes and gave me the closest Asian restaurant to me.
Once the system prompt complexity goes up, Flash starts to write very dense english. it might be fine for tasks like coding, but not for user-facing text meant to be digested by the average person.
I haven't tested 3.8 on my workload yet.
[0]: https://aibenchy.com/compare/google-gemini-3-6-flash-high/go...
I have an eval harness that runs every Thursday to determine which models are the current best for a few different client workflows. And since May(?) flash has slowly been taking over more and more stuff to the point it is now 100% on 8 out of 11 document extraction flows with the other 3 being a Flash / Opus 4.8 mix for high value stuff where cost is less of a factor.
What's the simplest explanation?
It may be that these flash models are simply post trained larger older models.
https://developers.googleblog.com/an-important-update-transi...
Gemini is great via the Chat interface and decent via Github Copilot.
I honestly hate it via Antigravity CLI because their sandboxing system frankly doesn't work. Every other harness has mastered "don't ask me if you're working in this one directory and using common commands". Agy instead either tries to pull a global elevation or wants every tedious variation of a command string whitelisted. Madness - circa 2023.
Agy _really_ needs to make the out-of-the-box experience cleaner and hassle-free. Heck, even Grok CLI "just works".
This may reflect a global mind-shift from "approve and validate everything" to "just do the stuff and only ask permission if it's outside the folder or a command that actually requires elevation". Maybe that's not for everyone, but for those that do want to perform unattended agentic work -- Agy is painful.
I just ran two light tasks on my codebase and got 100% of the weekly limits of a Pro plan blown away. Is Ultra plan any different? Because on Codex it wouldn't affect my Max plan at all, I think it would have been below 1% othese usage.
Anthropic and OpenAI are in a different game of spending investor money to buy market share.
Similarly Qwen3.8-Max was updated in just 30 days (to the 0902 release) and Muse Spark in just 28 days (to the 1.3 release).
A year ago iterative releases were every 3-6 months. At what point will they reach nightly candidates?
But it's bad at code reviews (maybe it's the harness agy cli?). Could not get it to same quality level on reviews like Opus, GPT 5.6, Grok. Even tried special code review skills but no luck.
I switched my subagent-swarm skill(https://github.com/bazelment/yoloswe/blob/main/.claude/skill...) to the 3.8 model and has been very happy so far. It has been driving the swarm to produce steady outcome, and more importantly, the exec communication is also crystal clear, instead of filling with jargons and long sentences.
Nah, we'll just get the 4.0 Pro Preview.
As much as I like the speed and interactivity, I really don’t trust it
[1] For tone and instruction following, a positive percentage increase represents an improvement in the tone of the model on sensitive topics and the model’s ability to follow instructions while remaining safe compared to Gemini 3 Flash. We mark improvements in green and regressions in red.
Gemini 3 Flash?! So is Gemini 3.8 Flash less safe than 3.7 Flash in all areas besides Text to Text Safety (and identical on Image to Text Safety)?Why bother with a column “Gemini 3.8 Flash vs. Gemini 3.7 Flash” when you’re going to disregard the label for 20% of it? Also is the “Tone” label short for “Tone and Instruction Following”?
Chartcrime, the major AI lab tradition.
I think the worst thing we can do is have loyalty towards models. I used to be loyal towards Claude, and my viewpoint changed dramatically when I used codex.
I highly recommend that if you are someone who only used one model so far, that you really give another model a shot and see how it goes. It's very eye opening and gives you a more holistic perspective.
Vendor locking is a big problem when it comes to models, and I hope the software world doesn't do this blindly.
Almost suspect that the rate of improvement to post-training is so fast that small models have an advantage - it takes much more compute to train a bigger model, so the flash models are just running in circles (well, not exactly of course) around the larger models right now.
1 what is tesla cybercab plan to address legal implications of accident that will happen? who is going to be responsible for them when they happen? are they covered by tesla insurance or some other insurance? are there any official plan/statements around that?
2 what was the name of the experiment they started in san antonio tx when some cars didn't have a driver? what was the results of it? did they expand the operations? it was much smaller than waymo, is it growing? how it is related to robotaxi?
It is not able to connect the dots that I keep asking about Tesla in 2nd prompt and spit out some unrelated stuff. Really? How it can be that bad? Gemini 3.1 Pro model works fine in this case btw. I thought maybe it is about knowledge cut over date and it doesn't know about those events from 2025 but it seems it has the knowledge up to March 2025. Top 10 in Intelligence on artificialanalysis ladies and gentlemen.
They're apparently evolving slower than most SOTA models but "slow and steady wins the race" is probably still a thing.
And since Google doesn't depend exclusively on AI models, they can probably afford to "wait and see" where all this craze is heading.
I don't think slow and steady will win this race, but I think anyone can still win--especially Google.
It still uses 3.6 Flash for example.
The Google brand remains powerful on HN!
I’m shocked.
I am continually impressed with Gemini's chat responses, which encourages me to test their agentic capabilities and... no... no... and no... every single time.
It's terrifying watching it, really.
so its fast sure and decent at non coding usage but for developers nothing can really top sol or fable.
even grok 4.6 is so so and i would not choose 3.8 flash over it.
So just another mid-tier flash model, nothing exciting, but Antigravity has very generous quotas so it's a good workhorse model when your Claude/OpenAI subs run out.
And whilst it's a fast model, having to baby sit through and approve prompts every few seconds ends up making it slower than the Auto approve modes of Claude/ChatGPT - they definitely need an auto approve mode.
A recent example - I searched for how to unsubscribe from Pearson emails. Google Search "AI Mode" confidently gave me a sequence of steps along the lines of Settings > Profile > Email preferences > Unsubscribe.
Of course, I looked for an unsubscribe link before asking Google. None of those options existed. The correct answer was there is no way to unsubscribe through the account, so I just blockthe emails instead.
I've run into this pattern quite a few times. AI Mode seems to make up things all the time.
Are you people truly this excited about this crap? I mean I guess if you work for Google or Anthropic or whatever I could see it??? Otherwise, are these just bot comments?