(Is it given that there's a play? Waymo is culturally a little distinct from Google so from personal experience I can't actually rule out that they're just doing this coz they think it would be good for the world).
4,994 karma · joined May 25, 2013
(Is it given that there's a play? Waymo is culturally a little distinct from Google so from personal experience I can't actually rule out that they're just doing this coz they think it would be good for the world).
The goal is "Linux on Mac". In a commercial product, specific details of what gets built usually stem from a more detailed vision that can be articulated as one. In open source projects those details are determined by what individual contributors are motivated to work on.
That doesn't means a cohesive vision can't emerge, there might be one for Asahi actually. But I think the default position is that it's kinda just a mob of people doing what they each feel like doing.
If you have any professional connection at all to Linux then $11/mo is an absolute bargain!
It's not something I expect, but I definitely don't see an issue with people putting a couple of sentences in an "expand to see AI usage notes" tab at the top or bottom of the page.
> Would you expect none?
As in, would I expect no AI usage? Coz no I absolutely expect people to use AI.
If you had got Claude to do this investigation for you then actually I think that would have been notable enough to include details of your AI workflow. But not because "it's cheating not to disclose it", rather because you'd probably have come up with quite a neat workflow that is interesting in its own right.
My main question was "did Claude do the investigation by itself or just write up the article from someone's notes?"
Also though, I'm quite willing to believe this is human written and the human just happens to have a Claudey style. The actual prose isn't that Claudey it's just the structure of how it presents ideas. But, Claude had to get that structure from somewhere. It's not that surprising to see people with that style of communication.
I would reject "Auto Mode is safe" as a message but FWIW I am totally on board with "on aggregate, making Auto Mode the default improves the safety of Claude Code compared to the prior status quo". Coz I would say in the vast majority of cases the access classifier is doing a better job than the thing it replaced.
Anyway yeah. Like I said, conclusion is the same: we should be decoupling this from the harness. We ought to be sandboxing agents the same way we sandbox applications. I wish Claude Code would make this path smoother :(
I don't really understand what Claude Cowork is. Sometimes I use it instead of plain Claude Chat for tasks that "feel big"? And when I've done that I've felt this was the right choice. But until I actually had that session up and running I would not be able to articulate what Cowork is or enumerate any specifics of how I expected it to behave. I just thought "maybe this is a Cowork task?"
(I have only used the cloud one. My understanding is the the local app is essentially Claude Code but for non-coders, which matches the description of the OpenAI equivalent).
Even as someone who's now used it a few times I wouldn't have been able to articulate as many details as Simon W does in this article, it's a weirdly shaped product structure.
To be fair to the authors they don't actually say it is. But then they contrast it with the "0.00% prompt injection attack success rate".
The upshot is kinda the same - this is still evidence that we should be sandboxing our agents. But it doesn't actually challenge Anthropic's "our models are too clever to prompt-inject" vibe.
I think the term "stake your reputation" is quite good for this situation. I.e. "hi new contributor, please acknowledge that you Stake Your Reputation before we move ahead with this code review".
(Doesn't help with legal aspects of course)
This scares me because the stuff that screams "THIS COMMENT WAS WRITTEN BY CLAUDE" has gotta be the tip of the iceberg. There must be much more of these comments that are better-prompted and using better-tuned model pipelines to make them less detectable. (Probably even just the lazy garbage is harder for me personally to detect if you just pass the prompt to GPT instead of Claude, it's just that I've read so much Claude output that I'm attuned to it).
So yeah suddenly dead internet theory is seeming a lot more real today than it did just a couple of months ago.
Basically it lets you treat the commits as part of the thing you are reviewing instead of just a minor detail that the UI doesn't care about very much.
And yeah it's just a bit of a weird black box, and also fundamentally it's exploiting a nontrivial system to do something that had little relationships to its design goals. I don't know anything about electrical systems but I assume there is all kinds of random shit that can disrupt Powerline links while still being fully within the design envelope of the power system.
Now I have a Ubiquiti one, it's antithetical to my DIY instincts but I think I'm OK with my router being one of the proprietary appliances in my life.
Plus now the wired section of my LAN is really fast.
There is about 10m and a brick wall between my AP and my desk so yeah, a fancy new MCS that can theoretically deliver 25Gbps was just never gonna be useful to me.
Also I've worked on (other non-WiFi) 802.11 products where our competitive edge was all about interference motivation. Our competitors at trade shows would see the demos and say "that's fake, you guys are cheating". So I have the feeling that in general there might be alpha in more resilient wireless links.
This seems like a common enough issue that it would be worthwhile publishing a "Markdown library" of "stuff HTML can do" that you can "call" from your AGENTS.md.
Edit: lol, a sibling commenter shows I am not the first to have this idea! https://developer.chrome.com/docs/modern-web-guidance/get-st...
(They made it a skill, which brings me to another thing I've had on my mind lately: that's annoying if you know when you start the harness that you will want it to read the skill. It's a roundtrip through the model to activate it. But maybe there's a way to fix that?)
(Edit2: Claude says if you activate the skill via /skill-name in the initial prompt, you get what I want)
In this case:
"While we would expect children's lungs to grow year on year, our results indicate that the London children's lung growth had accelerated during the four years to 'catch up' with the control group in Luton to reach similar levels of lung capacity," researchers say.
"I was absolutely stunned when I first saw the results," Prof Chris Griffiths, a senior author on the study, at Queen Mary University of London, told the BBC.
"The speed of catch up in lung capacity in the London group was surprising and impressive."
Nothing wrong with AI code either but it suggests the project may have been built in 2 weeks, has no community and low likelihood of ongoing maintenance.
So now overall instead of a flashy website I think the best signal of quality is a boring GitHub README and several months of commit activity, lots of merged PRs.
If it's popping up in some random product I don't pay for I can only assume it's such a shitty model that it's probably not worth talking to.
(Exception: my bank has replaced its shitty phone tree with an AI chatbot. This is faster and the model has a "transfer to a human" tool. It's perfectly capable of following its playbook, determining that I need a human, and putting me in touch with one. So actually I started with "now I've gotta debase myself by talking to a fucking 32B Chinese LLM" but ended with "oh that was great actually").
I guess they're quite happy with that TBF, it doesn't seem like they're that interested in being a git forge anyway.
But anyway I don't think it matters. The hard part here is having a distro image supply chain, not compiling stuff. If they do need to start compiling loads of stuff they didn't before... They just need to add CPUs.
Exactly as you say, x86 is successful coz there are really good x86 CPUs and a SW ecosystem. Not because it's a good arch. Similarly I think, if it wins, RISCV will win because there's enough demand for a nonproprietary ISA, and this trumps basically everything else.
Another way to say this: I've seen internal slide decks about ISA strategy that target C-suites at megacorps. Not once do those slide decks ever go into technical dis/advantages, coz the people steering the 9-figure investments don't give a shit. If you piped up in those meetings and said "RISCV sucks though!!!" you would just be ignored and not invited back.
It's not really any different from startups. Tech stacks matter but there's a huge number of successful companies out there running total garbage and who knows how many perfectly engineered products go out of business every day coz the thing they perfectly engineered was not useful.
This feels right to me. I've worked on embedded systems and OS my whole career and it just doesn't really feel like riscv wins or loses on technical merit to me.
Boots-on-the-ground engineers might grumble but the decisions that matter here just aren't made by them. Otherwise why would we still be using x86 so much?
I've worked on platforms that had barely-functional CPUs, we had to regularly be in touch with the engs that implemented them and they'd regularly go "oh yeah you've found a bug in our memory subsystem". This was kinda annoying for us but those CPUs were cheap as fuck and that opened up business models that weren't otherwise available. This is always gonna be more important than "ah but the CPUs are suboptimal!!"
Sorry I haven't actually read the original article so this is a bit of a driveby.
But, I'm skeptical of this? Surely whichever cloud platform first offers RISC-V compute, the next cloud platforms will ship a CPU that's compatible. Probably they'll be RFTing the same CPU vendors and they won't need to coordinate explicitly for this to happen. Then everyone else building RISC-V servers will essentially be forced to align on the "AWS variant" or whatever.
So yes you'll need a -cloud build of the distro you use but most people are already doing that and the cloud platforms are already providing the infra for distros to ship it (I assume they are also contributing to the -cloud distro builds directly).
RISC-V laptops and phones I could see this being an issue but for server compute it feels like there's gonna be a Schelling point.
(Actually, for phones can't Google just fix this by fiat?)
FWIW neither is xAI, there is no "big 3". xAI has had momentary peaks (I think they are having one right now) but they have never been able to claim to consistently push the frontier in any particular direction. You can also infer they aren't a frontier lab from the fact that they sell their compute.
Google's Opus competitor is 3.1 Pro Preview which is essentially obsolete (competed with Opus 4.6). They do not have a Fable/Sol competitor.