HNHacker News
TopNewBestAskShowJobs

tekacs

4,814 karma · joined March 23, 2011

Amar Sood

Building pervasive.app

hn @ <username> . com

On GitHub, Twitter, etc. as tekacs.

If you've replied to me and I've not got back to you... I'm probably busy. I'm trying to change that.

ac5ecc4e73f4852ebd856bd3959cc2b611802b6330356a61088dabc826d60c86

meet.hn/city/us-New York

submissionscomments
tekacs··on Claude partial outage
They don't, no (last that I looked) – if you switch providers, encrypted reasoning no longer works.
tekacs··on Unreal Agent
The headline graph is kind of bizarre.

For some reason they're comparing their harness running on Astra xhigh to Codex with Astra max?

---

Also worth noting that OpenAI just added support for async tool calling to their harness, which isn't 1:1 with this approach, but is slowly ramping up in being able to provide something similar.

A big part of why Codex uses so many tokens is that it basically hot loops on polling tasks it starts for... absolutely no good reason: https://www.reddit.com/r/codex/comments/1wdlp7q/weve_discove...

I fixed it on my fork of Codex too, also back in Jan/Feb – I keep this patch rebased, for anyone who wants it: https://github.com/tekacs/codex/commit/9ffcf8db9078eae43d411...

It results in token savings similar in scale to those displayed here by Unreal.

---

My harness has used a slightly fancier version of the approach that Unreal is using since ~Feb, and... it definitely works excellently, but it's also assuredly smoother with Astra and other recent models that are more aware of async tool calling.

tekacs··on Claude Cowork and chat are now one Claude
The OpenAI folks have been clear that they are merging them by the end of the year, obviously TBD what they do with quota.
tekacs··on GPT-6 Astra
https://developers.openai.com/api/docs/guides/latest-model

The docs page has a bunch more interesting details, including for example async tool calling!

tekacs··on OpenAI begins rolling out GPT-6 Astra
I think they embargoed the news, and then they failed to put up their own blog post synchronized to the scheduled news releases, probably because of the outages they're having today.

Reuters announced at 2.03pm and at 2.40pm still no blog post.

All the news articles say that OpenAI announced it in a blog post, of course.

All the love to the folks at OpenAI scrambling to get this out right now!

Edit: HN user codergautam mirrored the launch post, below: https://astratest.codergautam.workers.dev/GPT-6%20Astra_%20A...

Edit 3.31pm: Live now! https://openai.com/index/gpt-6-astra/

tekacs··on Bernie Ban Artificial Superintelligence Act
Posting this in opposition to it for awareness
tekacs··on Elevated Errors for Multiple Models
Quite possibly, yeah! Although in the past, whenever Anthropic has been down, I've always been able to rely on using OpenAI.
tekacs··on Claude outage – Resolved
Amusingly and frustratingly, this coincides with Codex also being down right now:

https://www.reddit.com/r/codex/comments/1w69ff6/whats_this_e...

tekacs··on Dwarf Fortress is getting the mother of all magic updates
Sadly no – the current 'classic' is just the Steam game without the steam features AIUI – i.e. still the mouse-oriented UI. 0.47.05 is the last keyboard version AFAIK.
tekacs··on Tencent Releases and Open-Sources Tencent Hy4 Preview
Yes, lots – I think that folks will hopefully discover more of these as they scale up their ambition, now that LLMs make a lot of previously difficult things far easier.
tekacs··on A new ceiling for Λ: the de Bruijn–Newman constant
It's really fun seeing Manim [1] used to illustrate a proof!

(See the video at the bottom)

[1]: https://github.com/3b1b/manim

tekacs··on Cloudflare OS: an open platform for agents, apps, and work
Yeah, with love the words "open source" mean a lot less here (de facto, of course not de jure) when the resulting open source product is entirely contingent on Cloudflare's PaaS.

And I think they suffer a little a little bit from their infrastructure being _comparatively_ unique. As providers go, they're on the harder end to replace with something else if you were to ever want to retarget your workloads.

tekacs··on Advancing the price-performance frontier with GPT‑5.6
In this case, and I don't mean this critically, I guess it would technically be, "Instructed model to find efficiencies... reducing inference cost by 20% saving company x billion dollars per month."

I have no doubt that further work was required to enable this, but it's still very cool to be possible to say that.

tekacs··on Show HN: Kedge – Full-stack cloud with forkable VM snapshots and global SQLite
Gotcha, interesting. I would love it to death if there were a way to run some CLI, maybe on the system, that gave you information on the current replication state of the disk!
tekacs··on Show HN: Kedge – Full-stack cloud with forkable VM snapshots and global SQLite
This is phenomenal, great work!

On the HN frontpage is Superlogical (I have no affiliation) right now, and I used the shell provided by the Kedge News demo to `ssh superlogical.jobs` from my browser on a phone on an instance started seconds earlier. :)

I didn't immediately spot the durability contract for volumes -- local NVMe like Fly, or sitting on some underlying system like EBS?

tekacs··on Our position on open-weights models
Very well said, thank you! The triple clause here is exactly what I mean.
tekacs··on Our position on open-weights models
This is completely backwards. Cybersecurity has an attacker-defender asymmetry that heavily, HEAVILY favors defenders.

For starters, a defender gets to pick the surface area, an attacker has to work with what they're given.

tekacs··on Ray tracing massive amounts of animated geometry using tetrahedral cages
For anyone who – like me – was hoping for a video to see what it looks like:

https://youtu.be/h2V8YyfKCVs?t=4421

tekacs··on Lilian Weng (co-founder) leaving Thinking Machines
Although this is interesting, it's worth noting that she's saying that her departure is for health reasons, so it feels like it'd be a hollow place for us – as HN is wont to do – to speculate on the company.
tekacs··on Cloudflare's new AI traffic options for customers
If you read the article, you'll notice that they're explicitly making sure to block the ones that interactively query specific content too.
tekacs··on Cloudflare's new AI traffic options for customers
I mean that they're telling developers that they should use Cloudflare's platform to build agents, the kind of agents that would go across the web and act on behalf of users... but then they're also the ones blocking those requests.

This always engenders a solid amount of distaste from me, because much like Google and Chrome, it creates the incentive for you to treat yourself better than others. Especially coupled with the trust stuff. Of course, Cloudflare is always going to trust their own platform.

tekacs··on Cloudflare's new AI traffic options for customers
> For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.

It's kind of exhausting seeing Cloudflare playing both sides of the arms race.

I just can't imagine bringing myself to use their technology to build agents and build AI products when they're also doing things like this.

> This also lines up the incentive model we want to foster. Losing trusted status across the more than 20% of web domains that sit behind Cloudflare is a deterrent with teeth. Trust becomes something you can carry with you, and something you can lose.

And even more so, LLM language aside, fun and fascinating to see them flagrantly calling out their position here as if it's a positive.

tekacs··on Claude Opus 5
Something fun: on our AWS Bedrock console right now, there's a 'NEW' model called 'anthropic.honey'. Wonder if that's the codename just for this one or in general?
tekacs··on Silent Replacement of Trusted macOS App Executables
Not what I immediately expected: it doesn't give you the cryptographic identity of the target app, but it does read a lot like a high-quality phishing attack where you get to e.g. issue permissions prompts with the totally specific visual identity of the app that you dropped your binary into. A highly technical user might, if we're very lucky, question why they're being asked for permissions again, but an ordinary user never would, for sure. If Apple is going to attest to identity, they need to make sure it's combinatorially right.

I totally understand why the author would love to see this fixed. But I'd honestly continue to go the other direction, (software freedom), and say that I rather wish that we would instead make this whole architecture less rigid. Apple seem to have looked at Windows Vista era UAC... and taken all the worst parts of it.

I tend to think that if we were to give the user a lot more information: people aren't as dumb as Apple treats them. Absolutely many users just do not care, but I think that making the UX as terse as it is is half of what's doing the damage.

---

Imagine:

  A new app downloaded from (signal.org)[1] would like to launch.

  The app is Signal by the (signal.org)[2] team, [which MATCHES where you downloaded it from][3].

  [Signal and signal.org are extremely popular, having been seen N million times this month][4]

  You may wish to check all of this information before launching it, but this is a [COMMON][5] app.
---

[1]: this would be the quarantine attr's source

[2]: this would be from the signing identity – but no need to explain that

[3]: I would literally check this and display the prose

[4]: yes, this involves a bunch of localization strings

[5]: This is where you'd have to decide whether you want to provide an opinion or completely disclaim liability, like Apple has had a tendency to do.

This would leave the door open to incrementally more relaxed versions of this posture, including allowing an unsigned app, or less common apps with more alarming language, up to and including color, to communicate how different that is. Even just a (localized to region) traffic light style appearance would go a long way. I totally acknowledge that, for example, the Chrome folks decided that they have to hide and bury this, because users can't be trusted not to proceed anyway when advised to do so. But this is why I think spelling it out in more detail would go a long way, which Chrome never (competently) did. Even with HTTPS certificates, I'd rather write an EXTREMELY high-quality explainer paragraph to the user about WHAT they're accepting.

tekacs··on Claude Fable produced a counterexample to the Jacobian Conjecture
Given how small the counterexample is... this feels like a great example of where a lot of interesting results are going to be found: not because they were super difficult, but because intelligence didn't scale, and until computers could do this for us, the number of people who seriously poked at many such things was low.

I'm very excited for the impact of this effect in science and medicine and other disciplines too.

tekacs··on OpenAI reduces Codex Model Context Size from 372k to 272k
I have these plan files as well, but it depends on the scope and scale of the things you're executing on, I think. However much detail gets put into the plan, it still doesn't help if part of what the model needs to understand is the fine-grained / perfect detail of a large surface area.
tekacs··on OpenAI reduces Codex Model Context Size from 372k to 272k
That isn't really a way around it for the kinds of things I'm talking about. The whole point is that getting all of the relevant context to even orchestrate it takes up most of the model's 272K.
tekacs··on OpenAI reduces Codex Model Context Size from 372k to 272k
Yep, Max. I would definitely struggle to code with Anthropic's plans on Pro.
tekacs··on OpenAI reduces Codex Model Context Size from 372k to 272k
When I've asked Codex agents about things that were in their context window, they've never – to my experience – been able to actually retrieve something from before compaction when using the proprietary compaction endpoint. Instead, they've had to consult their actual transcript.

So... at least as of a week ago or so, I don't believe so.

tekacs··on OpenAI reduces Codex Model Context Size from 372k to 272k
I've implemented a similar approach – although I'm surprised not to see mention of cache prefix busting in there!
Page 1 of 27Next →