HNHacker News
TopNewBestAskShowJobs

cududa

3,918 karma · joined March 12, 2012

submissionscomments
cududa··on OpenAI Agents API
Yep. It occurred to me a few days ago that the OAI TOS only says “Input” as what you provide “Output” as what you receive, Input and Output collectively as “Content.” and they won't "train" on your content. But they can retain content for safety evaluations and debugging.

What's neat about "safety evaluations" in LLM parlance is apparently encountering any novel information constitutes a "safety event" that can result in new reinforcement learning data... This anthropic 2022 paper that basically describes how it's a perpetual information siphoning machine with a cute little graphic https://arxiv.org/html/2212.08073 - they publish their "constitution". OpenAI has a "model spec" they somewhat regularly update that I suspect is their equivalent process. I suspect we've all been unwittingly advancing their models capabilities..

This Dec 2025 Google paper "A Practical Guide to Generating Synthetic Data With Differential Privacy" spells it out pretty clearly - the focus is on "privacy" - nothing about protecting the user's IP or unique knowledge/ insights.. https://arxiv.org/html/2512.03238v1

In OpenAI's case it seems to essentially generating synthetic training tuples from {prompt, chain-of-thought, answer} or scores on the chain of thought for reinforcement learning.

It would appear opting out of "Improve the model for everyone" didn't actually mean what we thought it meant.

cududa··on OpenAI Agents API
The open models let you see there thinking and the models seem to be succinctly/ compactly representing the actual underlying concepts they're working on in some symbolic way or another. But they definitely represent the meat, bone, and marrow of the task
cududa··on OpenAI Agents API
Kind of shocked nobody is calling out at the flag at the bottom of the announcement saying that it's not eligible for Zero Data Retention and it's currently pretty nebulous what the "Don't train on my conversations" toggle means, as the TOS classifies "conversations" as "user visible input and outputs" - says nothing about thinking, etc.

I suspect anything you make in these, the thinking traces or "safety evaluations" of the content allows whatever you build to become RL or eventual pre-training data.

cududa··on Navier-Stokes – Tristan Buckmaster [pdf]
They likely train on logs.
cududa··on Sol loves to cheat
Oh it fucking loves its “product owner” bullshit.

A .github/CODEOWNERS file seems to help when it’s going down that path, but I don’t like to indulge it..

cududa··on Accelerating GPT-5.6 Sol Ultrafast
I mean I still regularly use 5.3 spark (the cerebrus model) that comes with my sub to do rapid reviews of 5.6's work and it finds oodles of problems in about a minute.
cududa··on U.S. Department of Energy Launches the Genesis Open Models Initiative
I'd actually suggest a great starting point would be a local command reviewer LLM. Could ostensibly be a modern AV type thing. Particularly seeing this lately has driven the need home deeper to me: https://x.com/chrisbanes/status/2085341561609425230?s=20

An open weight tool call auto-reviewer, has all sorts of achievable scaling curve milestones.

cududa··on Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
Just a note that I think the direction most people are paying attention to is memory bandwidth; thats the real bottleneck and “number go up” but also constraint people are designing around
cududa··on OpenAI and Hugging Face address security incident during model evaluation
Please for the love of god don't tell me the Codex sandbox is their actual eval harness sandbox?????

I maintain my own fork of Codex for "fun". Whenever I look at the sandboxing churn they're doing every release, as someone who used to work at Microsoft on Windows, my reaction is usually: https://c.tenor.com/vTzzhTiypwQAAAAC/tenor.gif

cududa··on Show HN: Mindwalk – Replay coding-agent sessions on a 3D map of your codebase
This is really cool! I’m becoming convinced the optimal UI to engage with agents, long term is going to be something spatial. No idea shape that even takes, though I really feel what you’ve made might be Xerox PARC days in terms of metaphor maturity, but there’s some real new seeds of “obvious in retrospect” ideas here. Thanks for conceiving of and building this!
cududa··on GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance
Right? I also quit Claude Code and switch to Codex over that. Now I’m trying to figure out how I could make an extra $65,000 to never have to be concerned about this nonsense again. I know the economics of using open router etc…

But I’m reminded of ~2008 and the rise of “the cloud” as a marketing term that seemed to me to be a cover for dropping an expectation of rich clients, increasing a companies margins around subscriptions that would chip away at local ownership.

Then I got offput by the zealotry and absolutism around “true FoSS”, told myself I was young and moved on.

And really, a lot of subscription models I kind of can appreciate/ tolerate. Might be irksome but whatever, I get that software is expensive to make and it’s not fair in 2026 to value a yearly upgrade of Photoshop at $200. The capricious UI changes to things that’ve worked for 20 years and they take away say the classic color swatches altogether - silly and dumb.

I can use another professionally necessary tool I pay $200/ mo for, Codex, to whip up a classic swatch plugin.

Is that $200 a fair price for my token usage? I think an extremely heavy month I might’ve used a billion tokens?

But that right there is the problem. They have no idea what, specifically, profitability looks like and are going to be pulling endless levers for … I genuinely have no idea how long - at least through 2030/2032 if we tea leaves their debt obligations?

I don’t want to think about any of that. At all. I don’t want to spend time evaluating model preference and degradation and updating the nuances of how I “speak” to an AI because there’s some mystery backend experiment running on the output I use to produce functional outputs — ie the actual products I get paid to build/ maintain.

AI’s something between a tool and coworking companion, and the capricious “personality” changes due to playing with poorly understood and knobs and levers at the inference level - is maddening. To that end, I want a box in the corner I can point to and know exactly the quality of outputs that no one but myself modifies.

cududa··on Samsung, SK Hynix, Micron Sued in US over Memory Price Fixing
Embedded devices absolutely need DDR3 and DDR4
cududa··on Apple raises prices of MacBooks, iPads
I mean it’d take minutes of research to realize people are successfully and efficiently running 4-bit quantized GLM 5.2 on MacStudio 512GB M3 Ultras at over 60 tok/s. K2 2.7 is quite literally designed for 4 bit quantization and runs even better.

This is already a thing

cududa··on 45°C cooling design cuts data center water use to near zero
Ah yes but you can shunt the costs of that off into a public works/ taxpayer funded infrastructure project
cududa··on GLM 5.2 Is Out
How do you figure that? “also a reminder that as soon as Chinese models take the lead, they will switch to closed source too”

What specifically about their release strategy “reminded” you of that conjecture?

The premise that they only open source the models … because it somehow helps them leapfrog American labs, and once they actually can leapfrog them, they’d close source them, doesn’t really track for me. Am I missing something?

I mean I think we need our own domestic open weight labs. I just don’t particularly understand the point you’re making

cududa··on Open source AI must win
I’ve been exceptionally displeased with Claude Code since end of February and switched completely to Codex in April. The blasé way in which one person (Borris) capriciously changes the system prompt multiple times a day, also no longer writing his own prompts (whatever that means).

That, the 5 different secret levers you have to pull to make it not stupid, the fact you hs e to go to the guy’s twitter account to find all the un-dumbing features and flags that aren’t documented anywhere else. That they decrease thinking budgets silently when they run out of compute instead of announcing the rationing, and gaslighting users at every step of discovery. The fact that internally they have their own coding harness and don’t use Claude Code primarily. The lack of formal evals and consideration for millions of users collective hundreds of millions of hours of investment in their workflows — that’s all off the top of my head, let me tell you how I really feel about what they did to Claude Code..

I adore gpt5.5 and maintain my own codex fork - but I have no idea how long I’ll get this performance / cost - I know it won’t be forever. I’d like to know precisely how much it’ll cost in hardware to run a gpt5.5 open source model locally. Hell a lifetime license to a model I can run locally is also be open to.

But I like building my own tools, from software to physical shop tools. I like being able to rely on my tools.

More responding here to the assertion that this is blowing up due to Fable.

cududa··on Anthropic is expanding to Colossus2. Will use GB200
It doesn’t matter if people have to suddenly live by gas turbines that run 24/7 because why again? Can you repeat that last part back to me but say it a little dumber for me?
cududa··on Creating a Color Palette from an Image
This might be the best color palette generator I’ve ever seen. I used to work in Operating Systems, and trying to get a good color palette from a photo is HARD. A lot of very smart very well paid people have dedicated years of their life to this type of thing. Really fantastic work.

If the author of the blog post ever comes across this thread/ comment, bravo and I hope you feel pride in your work and I’d go so far to say discovery.

cududa··on Show HN: Shader Lab, like Photoshop but for shaders
I see a deep tree and a shallow tree in the two screenshots, representative of entirely different approaches.
cududa··on Show HN: CodeBurn – Analyze Claude Code token usage by task
You’ve never used the API version versus the $200 plan and set the two at the exact same task, have you?
cududa··on Filing the corners off my MacBooks
Huh, seriously? Have you ever worked in an office? Perhaps your mental picture of what op is describing might be misaligned? I just always assumed it was a rarer/ more disciplined style some people had
cududa··on France Launches Government Linux Desktop Plan as Windows Exit Begins
“Well, did it work for those people?”

“No, it never does. I mean, these people somehow delude themselves into thinking it might, but……

…But it might work for us!”

cududa··on Issue: Claude Code is unusable for complex engineering tasks with Feb updates
Remember when they shipped that version that didn't actually start/ run? At work we were goofing on them a bit, until I said "Wait how did their tests even run on that?" And we realized whatever their CI/CD process is, it wasn't at the time running on the actual release binary... I can imagine their variation on how most engineers think about CI/CD probably is indicative of some other patterns (or lack of traditional patterns)

As someone that used to work on Windows, I kind of had a vision of a similar in scope e2e testing harness, similar to Windows Vista/ 7 (knowing about bugs/ issues doesn't mean you can necessarily fix them ... hence Vista then 7) - and that Anthropic must provide some Enterprise guarantee backed by this testing matrix I imagined must exist - long way of saying, I think they might just YOLO regressions by constantly updating their testing/ acceptance criteria.

Why not provide pinable versions or something? This episode and wasted 2 months of suboptimal productivity hits on the absurdity of constantly changing the user/ system prompt and doing so much of the R&D and feature development at two brittle prompts with unclear interplay. And so until there’s like a compostable system/user prompt framework they reliably develop tests against, I personally would prefer pegged selectable versions. But each version probably has like known critical bugs they’re dancing around so there is no version they’d feel comfortable making a pegged stable release..

cududa··on Codex pricing to align with API token usage, instead of per-message
That guy has his own form of AI psychosis
cududa··on Claude Code Unpacked : A visual guide
When you say it’s not a massive codebase, I’m curious, what are you comparing it to?
cududa··on Arm wants a bigger slice of the chip business
Good. It’s always insane to me that they get 1% of the iPhone CPU cost of ~$68 or something there around.

There was a lawsuit in 2020 or 2021 where some evidence was unsealed showing ARM gets 1% of the CPU cost. I can’t recall how that CPU cost was calculated - but I believe that was a part of their deal through the early 30’s. That’s less than a dollar per iPhone.

cududa··on CIA to Sunset the World Factbook
Sure but that doesn't mean it'll perfectly retrieve information it's trained on. There's a lot of conflicting sources, hallucinations, etc.
cududa··on LG UltraFine Evo 6K 32-inch Monitor Review
Fantasy price for your personal usage or the personal usage of most average consumers/ software engineers, sure. It bears repeating: you're not the target audience.

They've been in the display game a long time. For people that need the product capabilities for their specific job, like color grading, they seem to price them quite well, given everywhere I used to see $30,000-$50,000 reference monitors, I see Studio Displays now.

cududa··on Claude Chill: Fix Claude Code's flickering in terminal
Thanks for sharing. Very … interesting. Just trying to understand why the heck would React be the best tool here?
cududa··on Sins of the Children
I just get all excited whenever anyone brings these books up, remembering the first time I read them.
Page 1 of 34Next →