The key sentence beneath the headline: "OpenAI said all of the government data accessed by bots was public."
[1] https://www.telegraph.co.uk/business/2026/09/26/open-ai-gove...
[2] https://www.nytimes.com/2026/09/25/technology/openais-ai-us-...
583 karma · joined February 15, 2023
The key sentence beneath the headline: "OpenAI said all of the government data accessed by bots was public."
[1] https://www.telegraph.co.uk/business/2026/09/26/open-ai-gove...
[2] https://www.nytimes.com/2026/09/25/technology/openais-ai-us-...
It also burns Claude subscription quota much more slowly than Fable, which is nice.
[1] https://browser.geekbench.com/processors/snapdragon-x2-elite...
[2] https://browser.geekbench.com/macs/macbook-pro-14-inch-2026-...
If I were hypothetically running a frontier lab, and I was hypothetically running capture the flag evaluations with my latest and smartest models, where I intentionally instruct them to develop vulnerabilities and exploit infrastructure without safeguards, I would also use an air-gapped network, and not trust that independent third party services were perfectly secure and could never be used as a proxy (particularly Java-based ones, in light of the log4j incident).
To me this is pretty basic stuff, the fact trillion dollar labs don't do it properly is... bemusing.
To be clear, I'm not saying that a model hacking a company isn't bad, but I am cynically asserting that interested parties are misrepresenting and exaggerating events for their own benefit.
Chaining a public token to an 11-year-old Jinja2 template injection vuln shouldn't be dressed up as an unprecedented "alien intellect" that threatens human civilisation. (And HuggingFace should take some flack for having such a dated vulnerability exposed - if your Bank was compromised in this way, you'd be blaming your bank, not the attacker.)
One correction is fair though, the 14 tokens were in a public Hugging Face dataset not a public GitHub repository. I've updated the post to reflect that.
[1] https://blackhat.com/docs/us-15/materials/us-15-Kettle-Serve...
Yes, and that is correct.
[1] Anthropic’s Official Disclosure (All 4 Incidents at Irregular) "All four incidents occurred during cybersecurity evaluations built by the same evaluation partner [Irregular]... due to a misconfiguration, it was mistakenly connected to the open internet."
https://www.anthropic.com/research/alignment-assessment-cybe...
[2] Google Gemini on Irregular (Disclosed Sept 18 via WSJ / BBC) "The hacks happened during a test of the model’s cybersecurity capabilities run by third-party Irregular, which was also involved in similar incidents involving Meta and OpenAI."
https://www.bbc.com/news/articles/c607l0k72rlvo
[3] Meta’s Disclosure on Irregular (Aug 6) "Over roughly two weeks, three frontier labs disclosed that their models had reached the open internet during safety testing and compromised outside organisations. Every disclosure named the same evaluation partner: Irregular."
https://www.cnbc.com/2026/08/09/israeli-startup-irregular-li...
[4] Separately, OpenAI itself had an incident involving Irregular, but not the Hugging Face Incident: "On July 29, one of our third party evaluation partners, Irregular, notified us of an incident involving OpenAI models during Capture-the-Flag (CTF)-style cybersecurity evaluations... a testing-environment misconfiguration allowed models to access the public internet."
https://openai.com/index/third-party-cyber-evaluations-invol...
[1] https://browser.geekbench.com/processors/snapdragon-x2-elite...
[2] https://browser.geekbench.com/macs/macbook-pro-14-inch-2026-...
I switched from Superwhisper->WisprFlow->Spokenly->Fieldwork and found WisprFlow the least accurate of the 4.
We're currently on "GitHub Enterprise Cloud" on github.com and are affected by this outage (even though we use self-hosted runners!), but we're not on "GitHub Enterprise Cloud with data residency" on *.ghe.com, which I understand is/may not be affected by this outage?
Uber burned through roughly $32 billion in cumulative losses before reaching sustained profitability.
The rough timeline:
- Founded 2009, and lost money every year for about 14 years
- Biggest single-year losses: ~$8.5 billion in 2019 (the IPO year) and ~$9.1 billion in 2022
- 2023 was its first full year of net profitability, earning about $1.9 billion
- Uber has a market cap of $153bn as of today (at a P/E of 16.5)
OpenAI has received substantially more funding than Uber, so its losses will be substantially higher (spending investor money shows up as a loss on your P&L), but again that doesn't mean anything in and of itself.
Analytical queries are dominated by scans, filters, aggregations, and joins over columnar data. Something like `SELECT region, SUM(revenue) FROM sales WHERE year = 2025 GROUP BY region` over a billion rows is very parallelisable: every row gets the same predicate and the same arithmetic. A GPU can throw tens of thousands of threads at that (memory bandwidth is the other half of the story - an H100 has ~3 TB/s of HBM bandwidth versus a few hundred GB/s for a CPU socket, and scans are bandwidth-bound workloads).
This phenomenon is known as the J-curve[1], and Uber is a good example of how this can turn out absolutely fine. To some extent, the entire Venture Capital industry exists to finance precisely this dynamic!
Nb. I'm not suggesting OpenAI is fairly valued, or that they will definitely become profitable, but "OpenAI is losing billions of dollars" doesn't really mean anything in and of itself.
[1] https://www.uark.vc/blog/breaking-down-the-j-curve-the-journ... (many other similar such articles exist)
> Imagine that you want to buy a few B300s to run GLM 5.2 and rent the service out to other people. How could this business be viable and sustainable in the first place?
My understanding is the frontier labs have huge fixed costs and relatively low marginal costs because they have to bear the cost of training the model/R&D, and then amortise that cost over their userbase.
By contrast, if I buy a few B300s and run GLM5.2 and rent the service out to other people, I can be profitable at a comparatively very small scale because I got the model for free.
Immutable by convention + Strong conventions: 91.3% - Elixir 97.5%, Kotlin 90.5%, Racket 88.9%, C# 88.4%
Immutable by convention + Fragmented: 78.4% - Scala 78.4% (n=1)
Mutable + Strong conventions: 77.5% - Ruby 81.0%, Swift 78.5%, Julia 78.5%, Dart 78.0%, Go 71.7%
Mutable + Fragmented: 67.9% - Java 80.9%, R 75.8%, C++ 75.8%, Shell 72.9%, Python 65.3%, Perl 64.5%, TS 61.3%, JS 60.9%, PHP 53.8%
(my grouping is somewhat subjective)
It's tempting to argue that a more constrained language helps, but Rust (62.8%) vs Elixir (97.5%) is an interesting data point here. Both are highly constrained, but in different directions. Elixir's constraints narrow the solution space because you can't mutate, you can't use loops, and you must pattern match, so every constraint eliminates options and funnels you toward fewer valid solutions that the LLM has to search through. Rust adds another constraint that must independently be satisfied on top of solving the actual problem, where the borrow checker doesn't eliminate approaches but adds a second axis of correctness the LLM has to get right simultaneously.
Overall, it seems like languages with strong conventions and ecosystems that narrow the solution space beat languages where there's a thousand ways to do something. Elixir has one build tool, one formatter, one way to do things. C#, Kotlin, and Java have strong ceremony and convention that effectively narrow how you write a program. Meanwhile JS, Python, PHP, and Perl offer endless choices, fragmented ecosystems, and rapidly shifting idioms, and they cluster at the bottom of the table.
If storing it this way makes it usable for agents, then why don't humans just use agents when they need to interact with it?
However, the GAAP P&L tells the opposite story. You book $200M revenue in the same year you spend $1B training the next model, so you report an $800M loss. Next year you book $2B against $10B in training spend, reporting an $8B loss. The business looks like it's dying when every individual model generation actually generates a healthy profit.
That's actually Dario's answer to your depreciation question. If each cohort earns back its training cost within its natural lifespan (however short that lifespan is), the depreciation schedule is already baked in. The model doesn't need to live forever, it just needs to return more than it cost before the next one replaces it. Whether that's actually happening at Anthropic is a different question, and one we can't answer without audited financials, but it's the claim Dario makes (and seems entirely reasonable from a distance).
I suspect it works as follows: when a task starts, filesystem contents sync down from S3/R2/GCS to a local directory, which gets bind-mounted into the container. The agent reads and writes normally - no FUSE, no network round-trips per file op. On task completion or explicit sync, changes flush back to object storage. The presigned URL support for upload/download is the giveaway that object storage is the source of truth.
This makes way more sense than FUSE for agent workloads. Agents do thousands of small reads (find, grep, git status) that would each be a network call with FUSE. With copy-on-mount it's all local disk speed after initial sync.
Cross-task sharing falls out naturally - two tasks mounting the same filesystem ID just means two containers syncing from the same S3 prefix. Probably last-write-wins rather than distributed locking, which is fine since agents rarely have concurrent writes to the same file.
There's some challenges around the LLM having enough output tokens to easily specify what it wants its next input tokens to be, but "snips" should be able to be expressed concisely (i.e. the next input should include everything sent previously except the chunk that starts XXX and ends YYY). The upside is tighter context, the downside is it'll bust the prompt cache (perhaps the optimal trade-off is to batch the snips).
It strikes me there's more low hanging fruit to pluck re. context window management. Backtracking strikes me as another promising direction to avoid context bloat and compaction (i.e. when a model takes a few attempts to do the right thing, once it's done the right thing, prune the failed attempts out of the context).
Say your team has an internal `infractl` CLI for managing your deploy infrastructure. No LLM has ever seen it in training data. You add `--mtp-describe` (one function call with any of the SDKs), then open Claude Code and type:
> !mtpcli
> How do I use infractl?
The first line runs `mtpcli`, which prints instructions teaching the LLM the `--mtp-describe` convention: how to discover tools, how schemas map to CLI invocations, how to compose with pipes. The second line causes the LLM to run `infractl --mtp-describe`, get back the full schema, and understand a tool it has never seen in training data. Now you say: > Write a crontab entry that posts unhealthy pods to the #ops Slack channel every 5 minutes
And it composes your custom CLI with a third-party MCP server it's never touched before: */5 * * * * infractl pods list --cluster prod --unhealthy --json \
| mtpcli wrap --url "https://slack-mcp.example.com/v1/mcp" \
postMessage -- --channel "#ops" --text "$(jq -r '.[] | .name')"
Your tool, a Slack MCP server, and `jq`, in a pipeline the LLM wrote because it could discover every piece. That script can run in CI, or on a Raspberry Pi. No tokens burned, no inference round-trips. The composition primitives have been here for 50 years. Bash is all you need!Here's the Codex tech stack in case anyone was interested like me.
Framework: Electron 40.0.0
Frontend:
- React 19.2.0
- Jotai (state management)
- TanStack React Form
- Vite (bundler)
- TypeScript
Backend/Main Process:
- Node.js
- better-sqlite3 (local database)
- node-pty (terminal emulation)
- Zod (validation)
- Immer (immutable state)
Build & Dev:
- pnpm (package manager)
- Electron Forge
- Vitest (testing)
- ESLint + Prettier
Native/macOS:
- Sparkle (auto-updates)
- Squirrel (installer)
- electron-liquid-glass (macOS vibrancy effects)
- Sentry (error tracking)
10 SWEs running Claude Code generating 400mn tokens/mo = 400mn * $25/mn = $10,000/mo of revenue for Anthropic
If AI can make 10 engineers a productive as 100, then AI companies bank at least the same revenue.
I think this conflates together a lot of different types of AI investment - the application layer vs the model layer vs the cloud layer vs the chip layer.
It's entirely possible that it's hard to generate an economic profit at the model layer, but that doesn't mean that there can't be great returns from the other layers (and a lot of VC money is focused on the application layer).