1. Stop AI. Not Pause, Stop.
2. Say goodbye to the concept of contributing labor to society, in the hope that the transition occurs quickly enough to not include 5-10y of working manual labor to survive.
3,038 karma · joined May 6, 2022
robb@doering.ai
1. Stop AI. Not Pause, Stop.
2. Say goodbye to the concept of contributing labor to society, in the hope that the transition occurs quickly enough to not include 5-10y of working manual labor to survive.
I'll be honest that my initial belief it was for OSX exclusively left me feeling indignant, so knowing I was just mistaken (based on the link up top, tbf) turns me around completly. I do in fact run Fedora, so I'll be trying this ASAP!
I feel Fedora+Debian+Ubuntu+Arch covers all but the long tail of devs, based on vibes alone? You might get bullied if you don't support Nix, but you'll probably be bullied by them anyway lol so no advice on navigating those waters.
Doing it all manually better guarantees quality of course, so makes sense for a lot of problems. But if you're looking to move a bit faster with a bit more risk, I'd sacrifice the writing step without sacrificing the reading & validating step.
~~[EDIT: you need to put "only for macOS" in way more prominent places, all over -- that offends my soul greatly and may Linus frown upon you all]~~ [EDIT2: I was mistaken!]
This all looks really solid. That said, two remarks:
1. The integration with OS LSPs is quite fun and commendable. Is it possible that the diagrams might get their own LSP, someday? Or is better just staying as direct TS callsites?
2. The choice of the word "IDE" seems like it might get you in trouble, given the small "cannot edit files" detail. Any comments on the decision there, as opposed to, say... "brainstorming tool"? Or hell, "[architectonic] harness"?
3. The psuedocode "semantic diff" thing is an incredible idea, wow. Props there.
4. This language kinda concerns me: "how the requirements that you set were implemented". In my highly-arbitrary development flow, it ideally goes `idea -> spec/reqs -> plan -> test -> impl -> eval -> land -> review`, and this kind of tool seems explicitly targeted towards just the second and third with some partial coverage of their neighbors on either side. More concretely: by adding implementation, don't you lose a powerful specificity selling point and now have to compete with all full harnesses?
5. Suggesting "GPT-6 Luna and Claude Opus 5.5" is presumably a typo? Cause the equivelant of Opus 5.5 is Astra, and even then not really.
--- Phil & [Pre-]CogSci ---
| rational | intuitive |
| reason | understanding | (Kant & Hegel's goofy wording)
| deductive | inductive |
| animated | automatic | (from ~Aristotle, ultimately)
| ~higher | ~lower | (as in "higher faculties")
--- CS & AI ---
| logical | analogical |
| symbolic | stochastic |
| Neats | Scruffies | (related academic 'camps')
| ~deterministic | ~non-determnistic | (a common heuristic)
| Good Ol' Fashioned AI (GOFAI) | Evil Datacenter AI (EDAI) | (a common honorific)
--- Modern [Cog-]Neuro[-Psych] ---
| slow thinking | fast thinking |
| S2 | S1 |
...ok I assumed I knew more but maybe I don't? Would love to hear from experts!
--- Colloquial ---
| left brain | right brain |
| intelligence | wisdom |
| intention | instinct |
| me | AD[H]D |
At least that's how I see it. Hopefully it goes without saying that there's plenty of nuance specific to each of those pairs, and that many of the relevant experts would balk at this broad characterization. That said... I'm right and they're wrong, I guess!If anyone is in the mood to stomach a paper written by someone who did horrible things with Epstein a few decades later, this is sadly still the best in the biz: https://www.inf.ufsc.br/~mauro.roisenberg/ine6102/leituras/a...
If you prefer geniuses that aren't arguably monsters, this is the core of what Gary Marcus' "neurosymbolic" thing is about.
I greatly resent having to be a biological organism and subject to all the poor engineering and design decisions that entails.
The first word of that sentence implies its absurdity quite beautifully. She resents, therefor she is.And then this:
Brains. After you reach adulthood, neurons don’t divide. If some of your neurons die—which happens every day—then they’re gone. If you get a brain injury, the other neurons will try to “learn around” the injury, but the neurons themselves are never replaced.
We need more neuroscience in schools! This summary is technically correct ofc, but to imply that the brain is missing some obviously-better option is absurd. We're meat that thinks, y'all; it's far from crap.IMHO it's worth keeping in mind that Anthropic employees are some of the least likely to casually pass off artificial prose as authentic, given the company's ethos/brand/cover story (depending on how cynical you are). To them this is all getting pretty high stakes pretty damn quickly; based on my usage of full strength Opus 5.5 today, I can't even imagine what working with their full internal stack must feel like. If they were willing to let the machines speak for them, they'd all be melancholically lounging around home by now instead of coming in to work!
...I am refusing to consider the fact that they probably are still WFH because of Salesforce forcing their shared security contractor to strike. Call that a mental health ignorance on my part :)
First, none of this changes the need for a separate `/docs` dir, also checked into git. Wikis are fun, but docs are essential!
Second, I personally think the new paradigm will be putting all of this into tons of new README.md files, which I'm kinda baffled aren't more common deep into dir hierarchies already. That intuitively tells the human authors and the artificial readers that;
A) it toes a similar "for technical people but not necessarily just our dedicated engineering team" line as the root README --more formal than an ephemeral "prompt" and less formal than a user-facing doc,
B) this isn't the AGENT.md file so should remain human-authored only,
C) this is only an overview with a strong preference for brevity & clarity, and
D) this is focused on this specific directory (along w/ the other benefits of locality, as the author extolls already).
Have I cracked the code? Is there a Turing award for inventing the concept of using a tool we already use but just a bit more extensively -- or at least a YC slot?
*P.S.* OP you dropped this: )
If someone were to share a preprint PDF that I don't believe, my instinct would be to double-check it against the institution associated with its publication. If it was never published then it's basically just a screenshot of a blog post.
Ohhh okay the opening makes more sense now. The blog above is publishing papers in form of blog posts as a protest against scientific publishing. I love the energy, but ruminating on how weird it is that so few people take advantage of the thing that you're weird for doing in the first place feels... obtuse?
I guess, in the end: I think it'll end up being fantastically useful for artificial engineers with their vastly superior ability to keep track of fine details and rapidly context switch, but fairly niche for any of us organic engineers that are left.
All that doesn't apply to low stakes stuff like games, though -- can't wait for the first truly open world game, NGL. A silver lining to the cobalt cloud?
Anyway, the paper I referenced is definitive -- its the founding paper of the field we're discussing, even. I suppose I can debate it if y'all have specific questions besides "but surely it can't really think", but fair warning, I'll be treating the conversation like I treat conversations with people still obssessed with the hockeystick graph.
https://www.hec.edu/sites/default/files/documents/Computing%...
There's really not much else to say, cause it's all just begging the question by assuming that dogma. Like, here:
> The fact that AI can use statistical prediction to answer questions or carry on conversations tells us something important about how regular our real world is.
Sure, it's interesting if you assume that it's "just" statistical prediction. There's a link, but it's just more creative restatements of the dogma, e.g. "But the LLM is just guessing words"
What is the use case for a classifier that works 90% of the time...? I feel like if I'm classifying something, I probably care enough that 90% ain't gonna cut it...
I guess the answer is just agential stuff that effectively gets double checked by the LLM in the driver seat, anyway? That tracks, though it means that jev is mostly just for the people making harnesses. Which is all of us but still!
Some people are still on Usenet after all (?), so I suppose it's not a big deal if a few people want to cling to old communities. I hope you don't mind if we use the word for what it was coined for though in the meantime, back in the real world.
Basic sandboxing is not exactly rocket science after all,[1] and it sure seems like they're missing a whole stack of swiss cheese slices on top of that. Some basic precautions off the top of my head that seem very likely to have caught all of these incidents:
1. Alerts based on telemetry (most importantly, HTTP requests), both explicit (normal) and semilatent (use DL to confirm an intentionally-eager alert before firing it).
2. Latent alerts based on transcripts, e.g. noticing when a thousand agents start mentioning a secret off-premises hangout spot. Even mere embedding comparisons seem likely to catch such a blatantly misaligned sentiment as that one, especially with n>1000.[2]
3. Pausing agents completely until an on-call engineer can rule on ambigious situations or potential issues -- surely security is worth <$1 in lost token cache, especially for a security company?
4. Superheavy orchestrator/baby-sitter models checking in on cybersecurity eval transcripts periodically just in case -- again, would be a neglible cost. Could also be made available to the agent as the first line of defense for clarifing a rule ad-hoc, feeding even confident responses to a queue that is reviewed asynchronously by humans within a workday.
5. Or, hell: just clearer prompts? I'm a cybersecurity noob, but I still feel confident we can write really productive, challenging CTFs without leaving questions open like "maybe I'm supposed to hack my own harness?"
Seeing as they haven't been fired by any of the big 3 yet, they're presumably smart, experienced, dedicated folks. And I'm not normally a "if only I were in charge!" person, I promise. But c'mon.
Perhaps I'm missing something?
[1]: To their credit we have gotten tidbits that indicate some blocklists & such exist, e.g. the German wiki hacks had to work around a blanket ban of POST requests.
[2]: This hints at their insane decision in one or both of the OpenAI incidents to just bandaid up the issue when found, which supersedes all of the above. You can stack swiss cheese slices a mile high and they'll still fail to protect you if the attacker gets to keep retrying & adapting indefinitely.
Cause "guessed passwords" could mean "stole hashes (?) and brute forced them offline" which is basically the quintessential hack. The "found credentials in a public repository" ones could be nothing, but it could be accomplished with a speed & thoroughness that was previously impossible.
The whole thing is made 10x weirder by the partial story -- I don't see any plausible incentive for them to keep the names secret. I guess maybe they're SMBs and thus warrant some privacy, but that would be quite the egregious scope creep indeed. Accidentally attacking the real cloudflare rather than a fake one is goofy but understandable; accidentally attacking Alice's Armoire Emporium or w/e would be baffling.
Y'all, it's Google. They own a money printer and the boring ~half of the AI field. They don't need these weird games to influence the government, and regardless, there is precisely a 0.0000000% chance of regulation happening before Trump's ouster anyway.
So please take it seriously. I'd like us all to survive this, ideally :/
So for clarity I have nothing against Chinese people of any kind, from the PRC, from Taiwan, or otherwise. We’re all on the human side ofc, and I’m a passionate internationalist (antinationalist, even). My country (the US) is in the middle of a fascistic self-coup, so it’s definitely not about superiority.
That said, your comment about conspiracy theories… it’s hard to know how to talk about this productively. But, uh, I’m not exactly picking those examples from nowhere — those are drawn directly from anthropic’s report. The only one that could be arguably a little overstated is the one regarding Uyghur refugees in Syria, where the refugees are often also involved in militaristic activities (supposedly, idk, I haven’t visited).
I don’t want to trip censors, but you can read the report yourself and then type in the zh names for the two departments I mentioned to your local search engine. They’re not hidden or secret or anything, and they’re not exactly bashful about their role in aggressively silencing dissent, either. Again the US sucks, but so far we only have one of those agencies (the monitoring one), and it’s been a tense, lively national controversy since at least Snowden.
I recognize that the PRC sees democracy differently; to you, a world where everyone’s data is always available to the government through its state corporations might not sound so bad. But I beg of you to reconsider. Surely you know that you can’t speak up against the party without being punished, and potentially even sent away indefinitely? Surely that tugs at your heartstrings a little bit, even if you’ve come to ignore it day to day?
I used to work in display ads at Google, which is the economic driver for the vast, vast majority of data collection. I’m not sure what your (firewalled…) internet is like, but over here in the anglosphere the only thing that’s “99.99% crazily cheap” and still quality —that is, the only parts of the “free and open internet” that Google claims to sustain— is shitty mobile games, mostly b/c they can advertise other shitty mobile games in an infinite vicious cycle of whale hunting.
If you’re able to read this message and are interested in replying, I’d be curious to hear about your dreams for the world. Clearly AGI can’t coexist with capitalism, so both western liberal capitalism and your proletarian state capitalism will have to go. I personally think national identities are also a global death sentence in an AGI world, but that’s more controversial. But what else?
Do you dream of a world where you or your kid could say something dumb about politics and not get pulled into a secret court and punished unfairly? Like, regardless of how possible or easy it would be. Is it desirable, at least?
Your English is stellar btw, don’t stress :)
(<3)
I'm sure you're far more experienced than I with basically every aspect of this discussion, but I'd argue that's given you a blindspot, here. I'll hit some specifics below, but the headline is that you're effectively taking a stand against Eternal September II -- a goal that I hope we can all agree would be quixotically antisocial, given what followed the first one!
we arguably do not want FLOSS to "win" over OSS
I think(/hope) that fellow FLOSS proponents would passionately disagree. FLOSS isn't a brand of chatroom, nor even merely a community: it's an ethos regarding labor, property, and liberty. Demanding that all users of your software are also activists for your particular take on intellectual property is clearly a doomed undertaking for anything beyond a toy or library, anyway.Didn't you get into this stuff to change the world? To liberate the oppressed, undereducated, and forgotten with the radical power of the information superhighway? Cause it reads here like you're more motivated by selfishness (not wanting to bother talking to people with less expertise than you) and resentment. On that note...
They want that. They do it themselves all the time.
Here you equate "non-hacker people" with software engineers you don't agree with, it seems. You're ofc welcome to think companies X Y & Z produce "miserable-ness", but as absurd as it sounds, it sure seems like you've forgotten the fact that some users are not developers. Many, in fact! Over 99%, even!Less confrontationally; my mom is in her late 60s, and is pretty computer-literate for her age after decades of knowledge work. Surely you'd agree that she's not, like, evil for using OSX, iOS, GMail, Word, etc.? That she didn't chose those things because of a philosophical commitment to defending IP laws, but rather because of structural reasons? Even if she were pro-IP, wouldn't we want to win good, well-meaning people to our side?
So let them have the "Open" prefix. It's just words, anyway.
I do agree with this still, but as a philosopher I just have to say that everything is just words. It's language games, in fact! Which is why I simply had to reply.I hope none of the above was rude; I'm trying hard to keep my passion for this topic from pushing me past HN guidelines :)
It does have a "mission" feature that's stuck in the strange, distant times of 2025 by way overdoing mandatory verification steps, which means they don't support swarms/workflows/crews/fleets yet -- that is, it's all done in sequence. But they have the boring, corporate engineering attitude that I think we're are all craving rn, and generally seem competent.
I can heartily dis-recommend Vix, even though they gamed themselves to the top of at least one ranking site that shall not be named; exactly like the quasi-bad-faith incompetence described with OpenCode above, but without even the "Open-" branding! Though perhaps that word has been so thoroughly burnt as a prefix by Sam Altman & Microsoft's criminal behavior that we should let it go...
Is this how "FLOSS" wins over "OSS"? Not with an ideological bang, but with a marketing issue?
With that personal failing in mind, I'd ask y'all to permit me to toe the guidelines just once, to proffer a hearty nyah nyah told ya so on a comment thread that spawned ~a dozen disagreeing replies this week! More seriously, I think this[1] is highly-relevant, shockingly-underreported context about the extent to which four PRC companies --Z, Alibaba, DeepSeek, and Moonshot-- are acting in bad faith. Consider it testimony as to their character, just in case anyone is thinking this might just be a simple misunderstanding.
So... nyah nyah, told us so:
> In the PRC, they[1] leaked tons of national secrets on the PRC's latest AI campaigns, the inner workings of their "opinion monitoring" (read: performative panopticon) and "stability" (read: violent oppression) departments, Chengdu's whole CCTV network, direct-energy weapons plans, espionage activities in Syria to hunt down Uyghur refugees, and god knows what else that Anthropic didn't divulge to us common folk.
> In the US, it's very clearly an attempt to rip off a competitor. I'm not sure how else you could possibly see it. Even if you're a distillation fan in general (which A. why and B. plz don't), they did this through a network of Japanese and Signaporean shell accounts, presumably at least some of which were abusing Anthropic's subscription service in a ToS double-whammy, as it would be exorbitantly expensive otherwise. They also had to hack around Anthropic's API to get CoT traces, which seems impossible to explain away as anything innocent.
> I've been beating the "China isn't necessarily an enemy, it's gonna take us all to handle AI" drum for literally years, but this attack was just... gross. Gross in scale and gross in arrogance. Not a good sign for the dawning alignment crisis, to say the least :(
> TL;DR: Use these services if you want, but know that you're supporting aggressive escalations and companies that very clearly don't give a flying fuck about violating the law, much less your ToS. So... buyer beware, I guess.
[1]: https://www.anthropic.com/threat-intelligence-report-septemb... is the report.
I lowkey suspect this PRC-based scandal has been underreported because Anthropic went insane with the sidebar UX on this page for some reason; there were many reports on the reports of Houti and Iranian usage, and very few on these sections. Could a week's mass media cycle be this seriously affected by such a stupid thing as a sidebar experiment?? Strange truth, or just fiction?
People might think I'm a dick, I don't care, my tools, my knowledge, my job security.
If you're feeling ponderous, I'd ask you to consider the outlandish hypothetical that this doesn't win you meaningful[1] job security, and you end up in the same bucket o' crabs as the rest of us. Would you regret your current actions?[1]: anywhere from the fathomable "my harness didn't come up in my layoff email nor in any of my failed interviews following that" to the realistic "the entire concept of employment shifted under our feet so quickly and so profoundly that a dash of solidarity & luck outweighed even the shrewdest preparations".