> Except costs associated with releasing internal documentation/code.
Jesus Christ, I had a list of comebacks I was expecting, but “doing stuff costs money therefore we shouldn’t ask trillion dollar market cap companies to do stuff because then they have to spend a bit of their trillions of dollars” on it, somehow. Out of baseline respect for the median HN reader, I think.
> The tiny minority of HN readers that wants to run asahi linux on their mac?
Please at least try to make points worth responding to. Asahi Linux community is tiny precisely because Apple does precisely nothing to support this and similar projects. Surely you can comprehend this. I believe in you.
Not all hardware companies, just the one with easily the world's largest (customer device count)x(closedness and ill will)x(cash in bank) volume.
Although, now that I think about it, maybe NVIDIA should also be compelled to support an open-source, community-controlled driver stack. I mean, if AMD can do it, then so can they.
Laws exist to make everyone's lives better. Apple can afford it, and doing this would gift the world a breathtaking amount of value. There are no downsides.
Apple still only start treating their own users with respect once we lawfully require them to support Linux on all Apple devices. Which is what we should be doing.
This says “Anthropic gives you 5× more API-priced model usage”, which is not equivalent to “Anthropic is 5x as much value”. This does not take into account model/harness efficiency. I tried looking up some benchmarks and Claude seemingly uses roughly 5x as many tokens while taking its sweet time to complete a task. No such thing as a coincidence.
It's so that, when judgement day comes, our Lord Jesus Christ of Nazareth has a concrete record of what you've been doing with your semen. It's very important, you see. God's work, even.
Holy shit: the software that works is already there, it’s open source, you just have to clone it, the code writes itself, and Google still manages to fuck it up. I swear, these guys are beyond salvation.
I keep hearing these anecdotes, but never see it in practice, nor in benchmarks. Do you think it's possible that there exists someone who asked Astra to make a game and they got an ugly one on day 1 and a nicer one later? And don't you think blender modeling is a bit of a "svg of a pelican" problem? It's not what the type of task the model is trained to be doing, and I wouldn't really expect the results to be particularly good or reproducible.
I don’t mean to advertise OpenAI - let’s make it clear, fuck OpenAI - but I’ve never seen a degradation like that in Codex. All models have always seemed completely stable over their release lifetime. Meanwhile, I rolled back my attempts at using open weight models because providers start throwing “too many requests” errors after just a few requests and the pricing is roughly 10x worse for same model quality, except the inference is much slower. Getting your weights silently downgraded sounds like fun.
Meanwhile, Codex devs just merge their changes and hit the release button. If a thingy breaks, they fix the thingy and release a hotfix. How irresponsible! It’s a miracle the software works at all!
You can override the system prompt in Codex, but AGENTS.md should probably work as well. Ask the agent to communicate intermediary updates more often using the “commentary” channel.
This is true - but also, models are absolutely terrible at writing articles and I don’t want to read them.
The issue isn’t that a 300 bit idea is padded with 15 KB of content. You can take any human-written article and reduce it by 90% with next to no information loss. What you lose is what makes the article a compelling read instead of a fact table.
I think the reality is that we will see quality long form AI-written content at some point. It doesn’t even feel like labs are particularly interested in chasing that now; code sells way more tokens. Right now the trend is that subsequent models degrade in writing quality as long as that pulls them up on coding benchmarks.
Let’s keep the Codex session JSONL there too, why the hell not. And the debug build logs, since they’re easily greppable text useful for diagnosing recurring problems. And logs/reports from every test run - a ton of useful info there, lets you track regressions over time; would be a shame to throw it away. We could also store screenshots of every app run to have a LLM-compatible historical record of how each component changed visually. And the token provider billing documents, since we’re gonna have a lot of those once we’ll start maintaining all that.
But abilities (and inabilities) are intrinsically tied to phenomenal experiences. And you can even shape your experience by learning. It feels different to look at code, speak a language, play an instrument or do a kickflip when you’re an expert (or exceptionally (un-)talented). Some people are biologically or developmentally predisposed to different abilities, and if you talk to them enough, you’ll learn that the way they experience these abilities is different from yours. I am conflating it because this is conflated in reality. This should be obvious in hindsight - your brain doesn’t have separate ability and qualia compartments, you know.
But there are obvious giveaways. Among your friends and colleagues there are people who can’t sing and people with perfect pitch perception despite no training, ones that can barely coherently speak but can write beautiful prose and vice versa, some can stare at code for 12h straight while others find the sheer idea of it horrifying. There’s people who believe in ghosts or even believe they’ve seen ghosts while others are grounded and rational. People with photographic memory and people who cannot recall what they were doing 5 minutes ago (me). The list of distinctions could keep going forever. Everyone’s good at something and shit at something because everyone’s mind is different. It’s silly how most people completely underestimate the diversity of human internal experience - maybe this type of close-mindedness is part of it too.
Right now agents are good enough for throwing semi-random ideas at the wall. Experiment compute is the bottleneck because it’s not much more than brute force search. A sufficiently intelligent agent with a deep model of its own architecture will more quickly and confidently locate improvements, the same way that high end LLMs can point out a bug and write a correct fix without even needing to observe and probe the program at runtime. If this level of research performance is reachable, experimentation may become much less of a bottleneck. Hopefully it isn’t.
Do you genuinely believe people always act logically? I don't know a single person that does. Don't underestimate the human capability to hold several completely contradictory beliefs with zero self-reflection. There's physicists who believe in god, for fuck's sake.
> How can you truly believe this and be ok with it?
Are you saying they don't truly believe this, or that they aren't OK with it?
Not sure about this. Codex already uses structured diffs, which are much simpler conceptually and shouldn’t require full file rewrites. The GPT model is probably tuned for its apply_patch tool, too. I’d like to see some real world benchmarks/comparisons.
The trick is to switch input to the laptop/external mic. Audio quality stays great. It’s kind of ridiculous that Bluetooth still needs the shitty lo-fi mode when the mic is on - I was recently surprised that this is still the case with modern headphones like it was a decade ago.
> Best of all no advertisements pinging me every few hours.
Skill issue.
The first thing anyone should do when they get a smart watch is disable all notifications. It's so easy to do and it makes the experience so much better. It's completely backwards that notifications are opt out. Sadly, I've never seen anyone change their defaults.
Your dream developer experience is already here; since I've switched to Linux, I feel like I'm experiencing Linux psychosis. It's such a breath of fresh air. I've reconfigured my entire workflow around terminal programs and every action feels instant, even with maxed out power saving settings. There's an upfront time investment, sure - a weekend or two to set things up and figure out how you want the basics configured - but I've found myself immediately productive, and after that it's all compound interest, baby. Over the next few months, I've pushed my muscle memory to a point that's completely unachievable on Windows.
Another big advantage of the platform is that some Linux distributions will come bundled absolutely no software, saving you from having to uninstall and disable annoying bloat you have no use for. Additionally, you don't even need to install WSL, because you've already got the L.
Ok, sounds like a watermark of “this post was written by Claude” is sufficient then. Or don’t add any at all - the average joe is that Canadian politician whose speech included “here is a more natural-flowing version of that section that sounds more like legislative speech rather than a series of short points.”
It is absolutely possible that it would not continue to be picked 10% of the time with a given fixed watermark key. The implementation literally labels tokens using a keyed hash and then modifies their scores. The entire point of the watermarking system is to bias certain tokens against others, and - as you would expect - this reportedly results in a reduced response diversity.
Yes, it is silly. Stripping these is about as trivial as removing "this post was written by Claude" appended in plaintext. You could make a clipboard monitor that does this as soon as you CTRL+C, it's a 1-shot prompt. Not to mention that these wonky Unicode chars will break in every other program.
Stripping Anthropic's watermarking, however, is more difficult - probably about 2 prompts.
What if the next token represents a wrong or low-quality answer, but would have only been picked 10% of the time, but now it's picked 20% of the time? Doesn't that obviously decrease the model quality, even though "it might have picked that token anyway"?