GPT-6 Astra on OpenRouter
openrouter.ai
openrouter.ai
Pelicans from Astra, plus 5.6 Sol, Terra, Luna for comparison: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...
I think this is a genuinely interesting comparison grid. Astra may be more expensive, but if you have a budget of 10 cents for a Pelican Astra low gives you something SO much better than the other models.
Astra uses less tokens overall too, for better results.
Astra transcript here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
I'm wondering if this is being trained on by the models today.
https://openai.com/index/advancing-the-price-performance-fro...
Sol discount is until November 21, 2026 according to https://developers.openai.com/api/docs/changelog
Luna — costs in cents
+--------+--------+---------+
| Effort | Before | After |
+--------+--------+---------+
| max | 7.83 | 1.57 |
| xhigh | 4.24 | 0.85 |
| high | 2.46 | 0.49 |
| medium | 1.26 | 0.25 |
| low | 0.76 | 0.15 |
| none | 0.71 | 0.14 |
+--------+--------+---------+
Per million tokens:
Before: $1 input / $6 output
After: $0.20 input / $1.20 output
Sol — costs in cents
+--------+--------+---------+
| Effort | Before | After |
+--------+--------+---------+
| max | 48.55 | 32.37 |
| xhigh | 24.11 | 16.08 |
| high | 10.38 | 6.92 |
| medium | 10.55 | 7.03 |
| low | 8.33 | 5.55 |
| none | 5.90 | 3.93 |
+--------+--------+---------+
Per million tokens:
Before: $5 input / $30 output
After: $4 input / $20 output
Terra — costs in cents
+--------+--------+---------+
| Effort | Before | After |
+--------+--------+---------+
| max | 32.09 | 25.67 |
| xhigh | 14.67 | 11.74 |
| high | 3.74 | 2.99 |
| medium | 3.46 | 2.77 |
| low | 3.47 | 2.78 |
| none | 2.60 | 2.08 |
+--------+--------+---------+
Per million tokens:
Before: $2.50 input / $15 output
After: $2 input / $12 outputReminded me of https://clocks.brianmoore.com/
You’ve been doing this public service for so long (well, for so long in “AI hype” years anyway) that it’d be fascinating to see the evolution of this artifact across time.
I'm running out of excuses not to build a proper comparison site though. Maybe I'll have Astra do rhat.
I used Light mode and it used up all my limits for the day and had to continue the following day.
Missing GLM-5.3, though.
Loving the pelican silliness. My ChatGPT is over the moon about it. Gonna miss it when it really is done. Though maybe a pelican riding a bike game could become the benchmark in a year.
Isnt Netherlands the leader in bike riders and they dont wear helmets.
If that was the case, the models would have been producing near perfect outputs for it a year ago.
Instead they are just training on general SVG generation, which in no way should be viewed as "benchmaxxing".
Yes
I'd be shocked if they didn't myself.
You really have to go out of your way to completely misrepresent what's being claimed here in order to make such a wildly off-topic reply.
Treat it like a bit as is
Quote from the thinking trace:
> I’m thinking about how a helmet would obscure lemur ears, but using an electric scooter helmet seems responsible.
It's pretty solid - face is a little wonky but excellent tail and scooter.
Did even better from the front. What's surprising is that it used the exact same colors as the simonw example, despite my prompt only being
> Generate an SVG of a ring-tailed lemur riding an electric scooter. Front view
GPT-6 Astra Extra High
Apparently the crowd agrees because they keep upvoting these.
The improved front fork design mentioned by Threatripper is about the only thing Astra is doing better, IMO.
I also think there's an extra level that I would hope an AI would nail which maybe an amateur artist would also fail at, such as thinking about what position a pelican would actually ride a bike in (maybe angling the beak down for aerodynamics etc.), but we are far away from this.
Coz who knows if astra low will produce max like output if tried once more.
It feels silly to say that about making a pelican image but it really shows the difference in output and cost in an easy to understand way.
I heard a rumor that Luna is a slightly different architecture from Sol and Terra, which makes me wonder if Luna and Astra might be more related to each other than to Sol and Terra.
Bit of a big leap to make from a token count though!
I just checked the tiktoken library and couldn't see any changes relating to Luna: https://github.com/openai/tiktoken
Here's an image design source of truth: https://image.non.io/78f4cd8b-2560-4643-9a51-96a89171f994.we...
And here's the page it build from it: https://image.non.io/e7d3a9e5-f9df-4fd8-b79f-1f90280f978f.we...
Note the flowing svg lines, and how accurately it recreated them. Here's Opus 5 for comparison - you can really see how while Astra really recreated the flow that was in the original design, opus only got the general vibe: https://image.non.io/dfe13de0-4487-431f-8b69-544ff3030dac.we...
One thing I will say is you are paying for quality. That site build cost $24 - extremely non-trivial for a simple frontend.
I would say that $24 is trivial IF that's the final design. The truth is that the cost doesn't leave much room for error or experimentation.
Compared to what?
Yeah if you are solo developer without budget.
For any business this is nothing, the ROI is massive.
Everything costs more now than it did a year ago... except for THIS, and we are still complaining that a 90-99% reduction in cost is STILL too expensive. And a 50-75% reduction in time is STILL too long.
We used to have to wait for weeks for a design like that when I worked at a consultancy, and that is a week of salary. For the design, then it got handed off to a front end developer to slice it and get built so the back end developer can hook it up to a CRM. We are talking a month turn around with design, revisions, development, testing, and bug fixing.
It can now be done in a couple of hours for less than a single hour's cost. If it were 10x slower and 10x more expensive, it would STILL be "good deal".
It sounds completely trivial and likely I'm wrong here, but could it be that opus saw the reference image squished? That might explain the sharper horizontal curvature
Another way to think about it is you could prompt Astra to put the building back in. You couldn't prompt Opus 5 to get the correct curvature of the line / cutout. That part has always been a huge struggle for models.
I would recommend staying away from OpenRouter. No matter how good the service is, if anything does go wrong, you have no recourse and you lose every credit in your account. Ironically some of the few responses I actually saw in the Discord were doubling down on their “no refunds no matter what” policy.
https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...
It took a while to test it, initially OpenRouter was giving Not Found errors for this model ID.
Good point about the mouths, I just noticed, lol
Imo, it's still better than most models, I personally like the stylized perspective.
You can view here all generations for all models: https://aibenchy.com/showcase/
It's more "correct" but looks a lot worse in my opinion:
https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...
https://aibenchy.com/showcase/?page=2#showcase=67fc6d6c8e4c3...
https://aibenchy.com/showcase/?page=3#showcase=c215b5c915da6...
What panel of judges are you using for scoring/ranking this? Seems subjective enough to not be able to be ranked/scored at all
I was thinking to manually grade/rank the SVGs, but I decided against it, as it is indeed subjective.
I was thinking it could have at least a simple objective check (hamster doesn't have extra or missing parts, table has 2 sides, and net is in the middle, etc.).
>Playing ping pong from the side of the table
I think marketing might be getting a bit absurd at this point
I even keep seeing obvious stealth marketing like this: "<topic> and how do I use it with <product> in <product>"
The entire mainstream media and political establishment, and every normie I meet is convinced Skynet is already upon us
I think they’re really going to struggle selling these models long-term. My company is already massively cutting down on access because they’ve realised most people don’t actually produce any value using it. All the tokenmaxers have ruined it for the rest of us now that accounting have seen the costs.
It must be super interesting working there rn :D
Edit: nevermind it JUST gave me a notification to use it!
Yeah, when I saw that Tweet I knew the person was saying it because they knew it'll be available within 24h.
almost had the feelin it was watching its little brother fail and had to 'step in' for a moment :').
time to go play outside...
It's terse, like all GPT models by default, but the sentences feel less obscurantist.
It's also more pro-active about problem solving.
Where do we go from here??
Edit:
GPT-6 Astra (low): 57 Intelligence Index, $7.70/M tokens
GPT-5.6 Sol (high): 57 Intelligence Index, $3.08/M tokens
So for the same measured intelligence, Sol costs only 40% as much — i.e. ~60% cheaper, while Astra is ~2.5× more expensive.
Why is the burden of proof on me tho!?
astra high is also 3x cheaper than opus max at basically the same intelligence.
astra high is also about as expensive as sol max while being more intelligent.
astra medium is cheaper than sol max while also being cheaper and roughly same intelligence.
im going to replace my sol usage with astra high/medium i think
caveat: benchmarks are really fuzzy with llms
Then I threw $100 for a Codex Max sub and it included Astra and it did it for me.
Sure seems like Astra is expensive AF.
Do Azure offer something that simply hitting the OpenAI endpoint doesn't provide?
The only reason most of my customers would use Azure Foundry instead of OpenAI directly is the ZDR assurance but it is so incredibly difficult to extract out of their model menu.
There is no trivial way to block non-ZDR models either so every customer has to “vet” and individually approve models.
If anyone from Microsoft is reading this: get your act together! You’re failing at the one thing people might want to pay you to do!
but again, seems like there's no word from Azure if this applies to them.
very confusing.
I'm a Business plan user with Cyber verification enabled, FWIW.
Has there been anything published about if Astra uses different amount of usage from your subscription plan compared to Sol? Don't recall coming across that in the press releases.
I've created with Sol a skill called Low Quota Mode that intends to reduce the use of tokens usages by the frontier (intelligent model) and delegate the use of bulk reading of docs/code and implementation to a sub-agent running Luna Max. Sol is asked to supervise, read the diffs and approves the commit/pr.
The skill might need some iterations while you use it, for example at the end of a rough session you can ask Sol how did it went, which were the points of conflict with Luna and try to iron them little by little by editing the skill.
Also in difficult tasks, ask to babysit the sub-agent model, I've seen it makes more effort into communication between frontier and sub-agent to guide the task with more care.
So far it has reduced my tokens usage a lot (have not quantified but the quota lasts more).