4,814 karma · joined March 23, 2011
Building pervasive.app
hn @ <username> . com
On GitHub, Twitter, etc. as tekacs.
If you've replied to me and I've not got back to you... I'm probably busy. I'm trying to change that.
ac5ecc4e73f4852ebd856bd3959cc2b611802b6330356a61088dabc826d60c86
meet.hn/city/us-New York
For some reason they're comparing their harness running on Astra xhigh to Codex with Astra max?
---
Also worth noting that OpenAI just added support for async tool calling to their harness, which isn't 1:1 with this approach, but is slowly ramping up in being able to provide something similar.
A big part of why Codex uses so many tokens is that it basically hot loops on polling tasks it starts for... absolutely no good reason: https://www.reddit.com/r/codex/comments/1wdlp7q/weve_discove...
I fixed it on my fork of Codex too, also back in Jan/Feb – I keep this patch rebased, for anyone who wants it: https://github.com/tekacs/codex/commit/9ffcf8db9078eae43d411...
It results in token savings similar in scale to those displayed here by Unreal.
---
My harness has used a slightly fancier version of the approach that Unreal is using since ~Feb, and... it definitely works excellently, but it's also assuredly smoother with Astra and other recent models that are more aware of async tool calling.
The docs page has a bunch more interesting details, including for example async tool calling!
Reuters announced at 2.03pm and at 2.40pm still no blog post.
All the news articles say that OpenAI announced it in a blog post, of course.
All the love to the folks at OpenAI scrambling to get this out right now!
Edit: HN user codergautam mirrored the launch post, below: https://astratest.codergautam.workers.dev/GPT-6%20Astra_%20A...
Edit 3.31pm: Live now! https://openai.com/index/gpt-6-astra/
https://www.reddit.com/r/codex/comments/1w69ff6/whats_this_e...
(See the video at the bottom)
And I think they suffer a little a little bit from their infrastructure being _comparatively_ unique. As providers go, they're on the harder end to replace with something else if you were to ever want to retarget your workloads.
I have no doubt that further work was required to enable this, but it's still very cool to be possible to say that.
On the HN frontpage is Superlogical (I have no affiliation) right now, and I used the shell provided by the Kedge News demo to `ssh superlogical.jobs` from my browser on a phone on an instance started seconds earlier. :)
I didn't immediately spot the durability contract for volumes -- local NVMe like Fly, or sitting on some underlying system like EBS?
For starters, a defender gets to pick the surface area, an attacker has to work with what they're given.
This always engenders a solid amount of distaste from me, because much like Google and Chrome, it creates the incentive for you to treat yourself better than others. Especially coupled with the trust stuff. Of course, Cloudflare is always going to trust their own platform.
It's kind of exhausting seeing Cloudflare playing both sides of the arms race.
I just can't imagine bringing myself to use their technology to build agents and build AI products when they're also doing things like this.
> This also lines up the incentive model we want to foster. Losing trusted status across the more than 20% of web domains that sit behind Cloudflare is a deterrent with teeth. Trust becomes something you can carry with you, and something you can lose.
And even more so, LLM language aside, fun and fascinating to see them flagrantly calling out their position here as if it's a positive.
I totally understand why the author would love to see this fixed. But I'd honestly continue to go the other direction, (software freedom), and say that I rather wish that we would instead make this whole architecture less rigid. Apple seem to have looked at Windows Vista era UAC... and taken all the worst parts of it.
I tend to think that if we were to give the user a lot more information: people aren't as dumb as Apple treats them. Absolutely many users just do not care, but I think that making the UX as terse as it is is half of what's doing the damage.
---
Imagine:
A new app downloaded from (signal.org)[1] would like to launch.
The app is Signal by the (signal.org)[2] team, [which MATCHES where you downloaded it from][3].
[Signal and signal.org are extremely popular, having been seen N million times this month][4]
You may wish to check all of this information before launching it, but this is a [COMMON][5] app.
---[1]: this would be the quarantine attr's source
[2]: this would be from the signing identity – but no need to explain that
[3]: I would literally check this and display the prose
[4]: yes, this involves a bunch of localization strings
[5]: This is where you'd have to decide whether you want to provide an opinion or completely disclaim liability, like Apple has had a tendency to do.
This would leave the door open to incrementally more relaxed versions of this posture, including allowing an unsigned app, or less common apps with more alarming language, up to and including color, to communicate how different that is. Even just a (localized to region) traffic light style appearance would go a long way. I totally acknowledge that, for example, the Chrome folks decided that they have to hide and bury this, because users can't be trusted not to proceed anyway when advised to do so. But this is why I think spelling it out in more detail would go a long way, which Chrome never (competently) did. Even with HTTPS certificates, I'd rather write an EXTREMELY high-quality explainer paragraph to the user about WHAT they're accepting.
I'm very excited for the impact of this effect in science and medicine and other disciplines too.
So... at least as of a week ago or so, I don't believe so.