HNHacker News
TopNewBestAskShowJobs

Willish42

752 karma · joined December 29, 2020

Opinions are my own, except for when they aren't.
submissionscomments
Willish42··on I resigned from Anthropic today
I think many of the skeptics here are focusing too much on the "agent out of control" skynet style scenario and not thinking hard enough about malicious actors using advanced AI to do large amounts of damage faster than it can be mitigated. Even with proper "guardrails" and alignment, the second and third order impacts of e.g. open source models getting better and more usable for cyber attacks could have some pretty scary possibilities.

The comparisons to nuclear and biological weapons seem to me pretty apt, and I don't think we should be so easy to discount the fears of people with a lot of inside knowledge and expertise on how things are progressing and where they might be headed.

Willish42··on Fable 5.1 World Modeling
> I would be especially curious to see the NPC person/car logic and if they're on rails or what, that's a pretty good NPC density for a demo.

Here ya go: - https://github.com/PhiloLabs/fable51-worlds/blob/main/union-... - https://github.com/PhiloLabs/fable51-worlds/blob/main/union-...

Willish42··on Fable 5.1 World Modeling
this is really neat. I wonder if people will start building open world games or "AR" games in the vein of Pokemon Go and Ingress based on similar tech. I've had a similar idea for a long time but don't think it was feasible before now due to AI
Willish42··on METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
More than other AI moments in the last few years, this feels to me like an event in tech that will be seen retroactively as an important watershed moment.

I really appreciated the negative framing and critical tone of this author's overview. Having also read through the METR report, I fee like frontier labs' pattern of getting PR about how impressive their models are "behind the scenes" has poisoned their ability to take security and reliability seriously. Sure these are new failure modes and the agents operate at a scale that's difficult to combat, but the lack of controls and concern for mitigating these sorts of hacks in the future is crazy to me.

The model for postmortems I was taught which has served me well in my career is thoroughly answering the following: - what happened / what was the timeline of events? - how was it mitigated and ultimately resolved? - what went well? - what went wrong? - where did we get "lucky" (meaning it could've gone worse but some arbitrary details about the incident worked out in our favor. usually stuff like "happened during business hours" or "we were already looking at a related thing that brought this to our attention before it was a worse outage") - (action items) how do we detect, mitigate, and prevent this type of failure in the future?

I really hope OpenAI has done an internal postmortem that answers these questions thoroughly. Most SWEs in the industry have to do such postmortems for much smaller outages with way less impact and risk of societal harm. This is probably another area where regulation and governmental oversight would help curb the risks. What's to stop OpenAI and other frontier labs from an intentional "accidental" attack that results in gaining access to competitors' systems?

I'm also curious what, if anything, Anthropic and Google have done differently to prevent a similar event. I suspect they actually monitored the agents as part of their studies and had better guardrails in their infra for how they set up their harnesses etc. for testing models, particularly when the other guardrails are absent as was the case here.

Willish42··on Astronauts describe persistent 'observer' sensation after 6 month missions
> Greg Egan's short story "learning to be me"

I read it and it is indeed quite good!

link for the lazy: https://gwern.net/doc/fiction/science-fiction/1995-egan.pdf

Willish42··on Leaked Document Shows the Surveillance Tech at ICE's Fingertips
> The document is dated May 2024 so is out of date, meaning ICE may not have access to every one of these capabilities at the time of writing. ICE has continued to spend hundreds of millions of dollars on peoples’ personal data and new surveillance technology.

I can't wait to see what kinds of horrific AI-assisted surveillance these organizations end up getting access to. The only saving grace of these sorts of post-patriot-act police state surveillance systems was that they still had to be driven by a human. Imagine this stuff combined with an "agentic" workflow where the "human in the loop" is the dumbest and most aggressive personality imaginable working for the US Gestapo

Willish42··on Making
I think another aspect is the deep understanding one gains from building the thing, and the idea of the thing being entirely the creator's and translated into "reality" through some kind of physical or mental labor. An author of fiction prose would have their vision in their mind or some other ideas of what the prose "should" be, and then created it with their own word choice, and tools. Similar for visual art. This applies even if parts of the creative vision were improvised while the creation was happening, but it started in the creator's mind and if faced with the final product, they could remember the small details and understand the work that went into it.

I think if I'm honest about when I look at LLM output for code, it feels close to a code review of code written by another person, where I spend a decent amount of cognitive effort figuring out those similar small details in reverse, without knowledge or memory of the specific motivations. And furthermore, if I were to read every single line of code, word in a written work, and pixel or brush stroke in visual art of some kind, it would feel like an artificial understanding, rather than something I can fully sign my name to as a thing I understood the ins and outs of while creating it.

I fully get that some people don't care about this, and see themselves more as "designers", but I think those are the same sorts of people who if given a choice, would rather "build" by leading a team of people and claiming credit for "building" as part of their leadership, rather than by doing it themselves by hand and getting a sense of fulfillment from that effort. Also, "agentic" workflows feel a bit like giving up even the leading / designing aspects. With Claude specifically, I'm often impressed (albeit other times often frustrated too) by the small design choices it made while building the code that I hadn't explicitly asked for but end up preferring once I see the output. This to me indicates I didn't really build it

Willish42··on Fable is now included on Max plans (up to 50% of weekly limit)
I've been using Fable on and off since the promotion started and have been fairly impressed with it, so I consider this great news!

At the same time, I sense a lot of negativity (especially on HN) in comparison to Codex from OpenAI (who IIUC is bleeding money to keep this going, though maybe Anthropic is too) which didn't do as many promotional and nerf-ing shenanigans.

As somebody who doesn't want to constantly switch tools, is Codex really that much better, or is this sour grapes from folks who want to stay on the subsidized access plans and are mad at Anthropic's pricing model?

Willish42··on OnePlus halts operations in USA and Europe
I remember the hype around the first OnePlus phone, with the invite system. It was the first time I'd had a device loaded with Cyanogenmod that ran relatively stable without any issues. Eventually the capactive touch home button on my OP2 gave out on me but I was a pretty happy user of the OnePlus One, 2, and 3T. I especially liked the removable backplates that had wood options that were pretty neat and felt nice to hold

I suppose Nothing is carrying that torch forward but it's still disappointing to see. Even though most of it was extensions of Oppo tech and ideas into a US/Europe-friendly market position, it still felt like they were innovating and keeping Android ecosystem healthy and interesting beyond simple slab phones.

I was considering looking into a OnePlus phone as my next device for Lineage or Graphene OS, but I guess I'm glad I waited...

Willish42··on Statement on US government directive to suspend access to Fable 5 and Mythos 5
I'm thankful I tested Fable via my subscription before the cutoff. To me, it seemed like at least part of the improvements were in how it broke up work for a "one-shot" style prompt, but I was very impressed at how quickly and effectively it produced better results than Opus 4.8 at a handful of real-world use cases I threw at it from my own work.

Another interesting finding which I've heard others corroborate is that even though the per-token cost was higher, it seemed to orchestrate the work efficiently and burn through fewer tokens, roughly evening out to approximately the same per-prompt token usage as Opus.

Willish42··on Can we have the day off?
Computers feel like a pretty good analogy for how AI will affect the workforce.

I suspect productivity will massively increase, the complexity and cognitive load of our work will similarly multiply, and yet we'll still being doing the now-more-complex work in some capacity for a similar number of hours.

Willish42··on YouTube to automatically label AI-generated videos
I've been thinking for some time that it wouldn't be too hard to create a third-party browser extension to crowdsource detection of channels that use primarily AI-generated content (for example, the AI slop music channels that put out multiple hour+ long genre or cover "playlists") and hide them from suggestions or home feeds.

My guess is that Google sees some kind of trend in a contingent of users preferring non-AI content and that surfacing AI content misleadingly has a negative effect on retention / watch time, and/or they're trying to get ahead of long-standing creators taking issue with the platform surfacing AI content disproportionately on account of it being excessively easier to upload in large quantities.

Willish42··on Tech CEOs are apparently suffering from AI psychosis
Like a lot of things in tech and pop science, "AI psychosis" had a narrow(-er) definition [1] (psychotic-like symptoms, i.e. believing the AI is in love with you, or being fueled by the AI into some strong delusions or belief in a "mission" so important your faith in it becomes quasi-religious/"destiny"), which aligned loosely with actual psychiatric symptoms.

The much broader version of people getting a bit too full of themselves or trusting an LLM's sycophantic nature as validation that they are right or uniquely smart with their ideas seems to be the version I'm seeing more of in tech news sites and places like HN

[1] https://www.psychologytoday.com/us/blog/urban-survival/20250...

Willish42··on Ti-84 Evo
I loved my TI-84+ SE and wish I still had it (had all sorts of custom programs on it but it got lost or stolen before I finished high school).

That said, I find it really hard to believe that they can't provide better specs and feature set for the cost. User-available memory of 3.5MB is incredibly low, especially with Python support. These could be really cool handheld computers if TI put more effort into their devices that already have a massive install base.

Currently, most of their popularity in my experience is "lock in" effect from teachers who are familiar with TI calculators and lab / curriculum materials that are specifically built around teaching through TI calculators. At this rate they're charging a lot and resting on their near monopoly status in education, which I'm sure is very profitable for TI.

There used to be a great app called WabbitEmu that emulated these devices on Android. I think they got a cease and desist but it was pretty neat to have back in the day

Willish42··on Ti-84 Evo
The cost of these devices isn't the computation, and if anything more connectivity would probably make these more expensive and harder to use (many "smart" devices in classrooms have networking issues and if even one of them can't connect, it hurts the ability to run a lesson). I think standalone computation abilities are pretty important, and connectivity can be a downside for preventing cheating in standardized exams etc.
Willish42··on YouTube now world's largest media company, topping Disney
I wish the audio quality of youtube videos matched other streaming services. Bandwidth-wise it's pretty minimal, but the audio quality isn't quite as good as competitors like Spotify (and the longer they take to upgrade audio bitrate, the longer the problem persists and uploaded content has lower audio fidelity)
Willish42··on I Am Not A Number. In memory of the more than 72,000 Palestinians killed
> By this logic, the Nazis were the good guys in WWII, and Israel would be the good guys if they'd just turn off all their pesky air defenses.

Can you elaborate on this? I thought that the Nazis were pretty obviously the "bad guys" due to committing genocide and mass casualties (combatant and civilian) while trying to expand their borders.

> It doesn't make any sense to try to judge morality based on casualty ratios.

Really, even the ratio of civilian casualties, or ratio of civilian casualties to combatant casualties? Those seem pretty relevant to morality in my book, but I might be misunderstanding.

Willish42··on Do your own writing
The trust component is so critical here. When I get halfway through reading a design doc and hit a part that's obviously slop, it really hurts my confidence in the project and in any faith in the developer having done their due diligence.

Certain communications, especially technical writing, are "expensive" both in terms of the effort of the author(s), and in terms of the person-hours of people reading them to gain understanding. Like mission-critical code, they should be written and reviewed with care, and at the very least heavily edited from an automated LLM output to be unrecognizable as such.

I personally don't use LLMs at all in my designs and I remain skeptical of the value proposition for those who do.

Willish42··on Axios compromised on NPM – Malicious versions drop remote access trojan
> This was not opportunistic. It was precision. The malicious dependency was staged 18 hours in advance.

Another obvious ChatGPT-ism. The fact that people are using AI to write these security posts doesn't surprise me, but the fact they use it to write a verbose article with spicy little snippets that LLMs seem to prefer does make it really hard to appreciate anything other than the simple facts in the article.

Yet another case in point for "do your own writing" (https://news.ycombinator.com/item?id=47573519)

Willish42··on The American Healthcare Conundrum
I think this is a political and economic problem rather than a technological one.

I cannot think of a more important skill than surgery to continue training humans to do and to be wary of AI robotics replacing. Sure, some surgeries could likely be automated, but the entire point of specialist surgeons is to make choices and act in a timely manner in ambiguous situations with extremely high stakes.

What happens when the robot messes up? What happens when the internet is down, or the hospital is operating under abnormal circumstances? How do you teach, train, and collaborate with human medical workers and caregivers in a world where surgeons have been replaced by robots?

Most of the excess costs for healthcare and surgery aren't the humans doing the work. I think there's a lot of other areas we can optimize first, chief among those in healthcare being the cost structure around private businesses and insurers bloating the bill with administrative costs. There's a reason every other developed nation has a single-payer healthcare system and better outcomes, and I don't think an AI breakthrough is the only plausible solution to improving costs in the US. In fact, under the current system, an AI breakthrough in medicine would likely hurt the workforce more than it would improve costs.

Willish42··on Don't post generated/AI-edited comments. HN is for conversation between humans.
This is an angle for people who default to AI-edited written speech that I've tried to be more empathetic to. I think it depends on your audience, but in professional writing that isn't published publicly (i.e. communication with your colleagues, design docs, etc.), or even the "rough draft" form of something that will be published, I think starting with your own words comes across as way more authentic.

I've seen enough GPT-generated slop that I find its style of writing very off-putting, and find it hurts the perceived competence or effort of the author when applied in the wrong context. I'm not sure if direct translation tools serve a better purpose here, but along with the other commenters, I personally find imperfect speech that was actually written "by hand" by the author easier and more straightforward to communicate with despite the imperfections. Also, non-ESL speakers make plenty of mistakes with grammar, spelling, etc. that humans are used to associating with "style" as authentic speech.

It can also become a crutch for language learners of any age / regardless of their primary language, that inhibits learning or finding one's own "style" of speech

Willish42··on Leaving Google has actively improved my life
I've been meaning to get off Gmail, and Proton Mail does seem like my favorite of the alternatives from a quick glance, but I'm also concerned about privacy focused services like Proton getting blocked or compromised in the US... This was a pretty good read

Also,

> I do my best to boycott bad things. And I fail pretty often. I still use Amazon on occasion and I can’t get off Spotify. I use Uber and DoorDash a lot more than I’d like. And I have too many Apple products/services.

OK, I can intuit why most of those are bad, but can somebody give me a good-faith interpretation on what's bad about Apple?

I'd assume it's the working conditions and material extraction processes in China, parts of Africa, and elsewhere, but isn't that true of every piece of consumer technology? The only better companies for consumer hardware that come to mind are Framework and Google for recycling parts and raw materials, but the whole point of the article is about de-googling and Framework's products are relatively niche and at a much lower price and performance / market category.

Willish42··on Claude Sonnet 4.6
I think many used to feel that Google was the standout ethical player in big tech, much like we currently view Anthropic in the AI space. I also hope Anthropic does a better job, but seeing how quickly Google folded on their ethics after having strong commitments to using AI for weapons and surveillance [1], I do not have a lot of hope, particularly with the current geopolitical situation the US is in. Corporations tend to support authoritarian regimes during weak economies, because authoritarianism can be really great for profits in the short term [2].

Edit: the true "test" will really be can Anthropic maintain their AI lead _while_ holding to ethical restrictions on its usage. If Google and OpenAI can surpass them or stay closely behind without the same ethical restrictions, the outcome for humanity will still be very bad. Employees at these places can also vote with their feet and it does seem like a lot of folks want to work at Anthropic over the alternatives.

[1] https://www.wired.com/story/google-responsible-ai-principles... [2] https://classroom.ricksteves.com/videos/fascism-and-the-econ...

Willish42··on Fix the iOS keyboard before the timer hits zero or I'm switching back to Android
> I randomly tried Android again for a few months last spring. Using a functioning keyboard was revelatory. But I came crawling back to iOS because I'm weak and the orange iPhone was pretty and the Pixel 10 was boring and I caved to the blue bubble pressure.

I know this is somewhat a joke site, but I think admitting this really proves Apple's dominance and doesn't really help in making your case. So long as the walled garden / "platform" approach still works, enshittification will continue

Willish42··on Claude Opus 4.6
I feel like this anecdote represents the differing incentives / philosophies of each group rather well.

I've noticed ChatGPT is rather high in its praise regardless of how valuable the input is, Gemini is less placating but still largely influenced by the perspective of the prompter, and Claude feels the most "honest" but humans are rather easy poor at judging this sort of thing.

Does anyone know if "sycophancy" has documented benchmarks the models are compared against? Maybe it's subjective and hard to measure, but given the issues with GPT 4o, this seems like a good thing to measure model to model to compare individual companies' changes as well as compare across companies.

Willish42··on Notepad++ supply chain attack breakdown
> cmd /c "whoami&&tasklist&&systeminfo&&netstat -ano" > a.txt

Naive question, but isn't this relatively safe information to expose for this level of attack? I guess the idea is to find systems vulnerable to 0-day exploits and similar based on this info? Still, that seems like a lot of effort just to get this data.

Willish42··on I'm addicted to being useful
This resonated a lot with me. I am also addicted to being useful and find that off days where my output and usefulness isn't where I'd like it to be really tank my self esteem.

I do think it can be a double-edged sword that often leads to burnout. Respecting your limits and occasional therapy seem to help, as does ensuring you're in as stable and supportive environment as possible so your efforts are sustainable and "heroics" don't get normalized in your org. I wish I had a full solution but have yet to find one in my career that works :)

Willish42··on We pwned X, Vercel, Cursor, and Discord through a supply-chain attack
> Apparently one of the other linked posts shows how you can also gain RCE

Yep, here it is: https://kibty.town/blog/mintlify/

Also linked in his guide (which I missed) and [here in a separate HN post](https://news.ycombinator.com/item?id=46317546). I think this other author's post is a lot more detailed and arguably more useful to folks reading on HN.

Willish42··on Voyager 1 is about to reach one light-day from Earth
> Communicating with Voyager 1 is slow. Commands now take about a day to arrive, with another day for confirmation.

I found this a bit silly given the headline: "well duh, that's the theoretical limit barring fancy quantum entaglement nonsense or similar!"

TIL all electromagnetic waves, including radio which Voyager 1 [uses](https://en.wikipedia.org/wiki/Voyager_1#Communication_system), travel at the speed of light. For some reason I always thought we had satellites doing some slower process or needing to somehow "see" light photons coming back from the probe to achieve near-lightspeed communication.

Willish42··on Cognitive and mental health correlates of short-form video use
There's some thoughtful comments here already, but I wonder the same thing constantly as a fairly addicted user of YouTube who wants to avoid short form video altogether.

I think Premium users tend to be the most affluent desirable group for ad targeting (similar to iOS users on other platforms) and even though YT Premium lets you avoid ads on YouTube, I suspect one's activity feed/"algorithm" on YouTube factors a lot into Google (and others'?) ad targeting. The same eerily effective feedback loop for getting TikTok and YouTube suggestions works better with short-form video, so even if users aren't seeing ads, YouTube still has an incentive to have people use it. So, there's money to be made in dialing in your "algorithm" from using YT Shorts even if you're a premium user.

I'm sure the other stuff about KPIs for increasing usage of shorts to compete with other media sites is accurate too

Page 1 of 6Next →