6,679 karma · joined October 14, 2022
Did we forget what Flash was meant for? I have tried Flash-lite, in my tests, almost all the recent models we have today are good and enough for most daily tasks.
Flash-lite is one of the few models available that is easy to notice its flaws. Its good to summarize data (for example search results), that’s it.
https://github.com/streetcomplete/StreetComplete/issues/5421...
Wow, just wow. We are racing to the bottom with these prices.
One thing I wish was better communicated is the mileage we get for our subscriptions. I do not fully understand how much usage I get with each model and their reasoning effort on 5h and weekly limit in Codex. I am asking because I know switching to Astra would consume my 5h usage limit quite rapidly, so I avoid it. If I knew how much mileage I would get from each model and respective reasoning effort, then I would be able to plan my workflow better and know when to upgrade model for a task. In almost all cases, GPT-6 Luna (XHigh) have been enough. That's why I appreciate its discount, because its dirt cheap, yet highly capable.
In other news:
> In the coming days, we’ll also offer GPT‑6.1 Sol Ultrafast , with up to 8x faster token generation compared to its standard speed in Codex.
8k diffs PR is incredible, is it even sane for a human to comprehend that much?
There has to be a name for such scenario were the handoff from Agent to Developer is so difficult, its not even worth pursuing but to instead wait until the Agent comes back online.
No matter how good of a developer you are, we are still constrained to our minds working memory. Lots of people I have spoken to are aware of this. Agents produce lots of diffs in a fast pace, our minds don't get the time to grasp the changed logic. And then, we an Agent is unavailable all of the sudden, say in the case of it being an outage or usage limit reached, the human now have to go through massive amounts of changes. They cannot continue as usual, there is a spike, they have to learn the new changes and grasp everything before they can continue towards the goal.
1. https://www.google.com/maps/@31.2963743,34.2449952,1665m
2. https://maps.apple.com/frame?map=satellite¢er=31.296508%...
> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.
https://community.openai.com/t/experimental-context-manageme...
I would also like to point out that it was quite predictable that Terra got discontinued, it didn’t make sense to have it when both Sol and Luna overlapped it.
Lunas insane discount is a game changer, OpenAI knows what they are doing here. Luna at max reasoning effort, even though its not optimal for long conversations, its incredibly intelligent while dirty cheap. Its not even competition anymore.
Whats even crazier is that I’ve underestimated how good Luna actually is. I’ve seen colleges create fantastic things with just Luna medium. This basically means you never have to think about your Codex usage anymore. You can run all day and not
have to worry about your 5h or weekly usage limit. To me, the discounts OpenAI is offering with Sol and Luna is truly a new milestone.
Title seems baity to get customers. Maybe it should be in ”Show HN”?
Oh interesting, I can assume what the benefits is for including the Encoder, but whats the downside? I’m thinking GPT (which is decoder only) ruled out Encoder for a reason?
Whataboutism?
I am just thinking loudly here but, it seems like even though Astra is pricier than Sol, you might actually get more usage out of it? I did some digging myself and looking at FrontierCode and DeepSWE, Astra (low) seem to perform better than Sol (medium) and on par with Luna (max) while being somewhat on the same price range to Sol? [2][3].
And now we have four models to chose from, each with their varied reasoning efforts: Astra, Sol, Terra and Luna. Personally, I feel like Terra have turned into this middle child in a weird spot that's neither the option as cheap model because Luna is, yet it is not an good option for complex tasks because Sol is already good at it.
For background, I use Luna (xhigh) daily, I think its a fantastic and underrated model. Especially Luna (max). It is way more capable than what it looks like, I think people underestimate it because OpenAI described it as "roughly corresponds to the nano model tier used in earlier GPT-5 families" [4]. I also like Luna because it barely consumes my weekly usage. Last week, it only ate ~15% of my weekly usage. So usage is not an issue anymore. I never have to worry. It may not be the fastest model because, well, it reasons as max effort, but it does the job way better than I expect. Also, considering how much one saves on the weekly usage, one can probably turn on "fast mode". Haven't done it myself though.
Also, another thing that caught my eyes is this:
> Historically, models have used compaction to summarize work during long sessions, such as when debugging complex issues or tackling large refactors. Each compaction can leave out details about why a fix failed or how a component behaves.
> In Codex, Astra can keep notes across context windows, preserving accumulated details without repeatedly compressing them into a single summary. Earlier context windows remain searchable, so Astra can find requirements or test results from previous messages and tool outputs—even if that information wasn’t captured in its notes. You can enable this experimental feature in your Codex config.toml [5].
I was curious about this, because I know Luna (max) spews out tokens which can trigger compaction quite often. If you go to the config reference [6] and search for "features.context_management.experimental_mode", you will find this:
> Enable experimental context management. Rather than repeatedly compressing context into a single summary, it uses notes and searchable history to preserve accumulated details.
This is a very interesting feature and perhaps very useful during long horizon work in a thread where the conversation context window grows and compacts often.
1. https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...
2. https://deepswe.datacurve.ai, Astra (low) got 67% $2.19, Sol (medium) 61% $1.42 and Luna (max) 67% $0.61
3. https://cognition.com/frontiercode, Astra (low) 45.3% $1.60, Sol (medium) 39.9% $3.12, Luna (max) 39.8% $0.36
4. https://developers.openai.com/api/docs/models/gpt-5.6-luna
5. https://openai.com/index/gpt-6-astra
6. https://learn.chatgpt.com/docs/config-file/config-reference