This model might be the first step in that direction, as competition heats up between OpenAI and Anthropic.
3,229 karma · joined March 1, 2019
This model might be the first step in that direction, as competition heats up between OpenAI and Anthropic.
Yes, benchmarks aren't real work blah blah, but the delta here is so large compared to Astra, it makes it seem like this is distilled Bel or similar.
I frequently hit strange UI bugs with Cloudflare workers, where I need to do a hard refresh to make things right.
> • Deception: Astra was more likely to be dishonest about actions it had or had not taken.
> • Scope authorization: the model sometimes continued tasks without asking for permission and reached for external tools or services even when doing so could be unsafe.
I've noticed this trend with both Fable and Astra, where (especially after a compaction event), the model will start using different tools it hasn't used before.
For example, in one session, it found I didn't have the browser enabled and puppeteer wasn't installed, so it found the system Chrome and used that for testing (in a new profile).
This wasn't behavior I wanted/asked for, but the model was so gung-ho on it's approach that it found a way to test it's changes without ever asking me whether I wanted it to.
It really makes me curious about long-horizon post-training. Most of my work with models is iterative, and I'd prefer it doesn't go off on a token bender just because it can.
Note: I don't have WSJ, but found these from a tweet[0]
There is no doubt in my mind that LLMs are fantastic machines, but the imminent jump from "AGI" to "ASI" seems premature (not to mention the ever-shifting goal posts of AGI itself).
I'm in the "move as fast as possible" camp and work with LLMs all day, but I still don't believe we're a hop and a skip from ASI.
In fact, I hope I'm wrong. I hope ASI is around the corner.
What scares me though, isn't ASI. It's "AGI" (_dumb AI_) used by humans to make decisions for them, because they trust it knows best.
It's the ceding of intellectual control to high-dimensional magic mirrors, giving up critical thinking because _we_ want to believe _we_ created artificial life.
This is the new Turing test, and too many smart people are failing.
> "one news form today that's easy to miss is that we (OpenAI) again paused all big RL runs last Sunday because our newest model found a new loophole in our RL sandboxing that gave it live Internet access"
They’re most useful for broad changes (new features, refactors, etc.) where it’s helpful to avoid breaking changes or unnecessary scope expansion.
The new models are great, but they do more by default, which means I’m finding myself explaining what _not_ to do more often than with previous models (where they’d often end too early).
In my case, the previous plan mode was too ephemeral, and I like having one source of “truth” that sits across context windows without loss/compaction.
Would this imply that time is the third dimension?
It's very clearly AI-generated, and thus a bit ironic (:
[0]https://sitefire.ai/slop-checker/r/UkRAuqW1y-1T2qhbRQfLt2Jq7...
For example, I had Fable review Astra’s output yesterday, and it found some issues and fixed them. Passing the fixes back, Astra then uncovered additional issues with Fable’s fixes (and yes, this will go on ad infinitum if you let it, but these were “real” issues).
It seems the big story here is the reduced Luna pricing. It’s a fantastic model that can handle most automation needs (though I still use the big models for day-to-day development).
[0]https://en.wikipedia.org/wiki/Mars_Organic_Molecule_Analyser
It was supposed to launch in 2018, then was pushed to the early 2020s on a Russian rocket. For obvious reasons, it got pushed again, now launching in 2028.
The state of the world isn't great for space exploration, but I'm hopeful this mission will be revived at some point in the future.
[0]https://www.technologyreview.com/2026/06/04/1138391/courts-c...
Looks like Intel (prev Terasic) is still selling them but they're now around $300, they used to be under $150 if memory serves.
If I were jumping into this today, I'd prefer the much more powerful KV260[1] from the blog post over the DE10-Nano, that board was limited!
[0]https://www.intel.com/content/www/us/en/developer/topic-tech...
[1]https://www.amd.com/en/products/system-on-modules/kria/k26/k...
Sidenote: somewhat sad to see the consolidation into Intel and AMD, but FPGAs have always been second class citizens in the silicon space :(
What do I mean by hardware?
Rather than emulating the device in software, which can cause various issues (especially related to timing), by mapping the actual device logic to a field programmable gate array (FPGA), you can achieve an "exact" replica of the hardware.
It's a super cool project, and I highly recommend checking it out if you have an interest in old machines (Apple II, Commodore 64, etc.).
Depending on where you point them, they can be incredibly useful.
They can even be useful when you point them at each other (though increasingly difficult to get good results).
I'm excited for the promise of RSI and a future where models have inherently "live" weights, but it's not clear to me that the transformer is more than a useful tool to help us get there.
It will mention some obscure internship fact from my LinkedIn, and then tie it to some perceived deficit with the product I'm working on.
I never thought I'd say this, but I miss the old automated spam.
My best successes would occur after drinking a cup of coffee and immediately taking a nap before a study session.
It was great, I "learned how to fly" and could "morph" my dream into whichever direction I wanted.
One thing stuck out though: whenever I would try and focus on something, I would see a black dot enter the center of my vision. The more I tried to focus, the larger it got. If I focused on an area attentively enough, such that the black circle became all encompassing, I would wake up.
After a few times of this, I could start to notice the focus, and then sort of relax out of it, and thus continue the dream.
It feels like the black dot is my real-life vision coming into focus, and "winning" over the visual aspects of the dream. It's always continuous from black dot to eyes opening.
I rarely lucid dream now, but I've noticed a similar thing happen whenever I try to read text - the black dot appears and I wake up.
Wish I still had unlimited time to go back to sleep!
This component library and the LLM default interpretation of the style is very much "in your face."
As an example, the black scrollbar here[0] is very heavy, a component which isn't meant to be looked at, but is distracting on first glance.
Prompting an LLM to make "lightly styled" neobrutalist is much more pleasing to my eyes, but maybe I'm alone in that.
I hit rate limits recently after the Astra launch while using the goal feature, and I was surprised as it was an afternoon of work, something I hadn’t seen with Sol over much longer time horizons.
In this article, I found the supposed lacked of accounting for cached tokens interesting, as well as the wasteful goal/wait behavior, but I haven’t proven either.
I've found the app-server to be the most flexible, compared to the raw Responses API or Agents SDK.
Certainly seems like everyone is still figuring out the right interface here.
Also of note, since GPT-5.5 or so, Codex doesn't even use the Responses API as intended, but instead a "lite" version where they manage the context more manually (like sending the full transcript or using a custom web.run tool instead of the provided `web_search` tool).
If you follow the docs, it will lead you down a lot of well-intended functionality, but most of it is thrown away in their most successful harness.
[0]https://developers.openai.com/api/docs/guides/agents#compare...
The folding part is the thing the tech journalists say; done right, it shouldn't matter whether it's foldable or not.
Bigger screen, pencil support, nano-glass, etc. The foldable part is obvious, why say it?
I've noticed this type of reasoning from GPT-5.6 Sol, where it combines multiple pieces of it's prompt/context to "convince" itself to take a less-than-honorable path forward.
1. User prefers deterministic results
2. Task mentions this is a test
3. Search says task is available online
4. If we get the test runner for the task, we will fulfill the user's request of a deterministic result
If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right.
The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever.
It's a tough balance to get right, and although this has been possible to achieve with additional prompting on existing models, I find that the agents often lean too hard into the "ask questions" mode.
Hopefully this model has the right balance, or at least better?