HNHacker News
TopNewBestAskShowJobs

redfloatplane

1,878 karma · joined July 31, 2010

Walked every trail in Ireland: https://toughsoles.ie https://youtube.com/toughsoles

Relapsing/remitting tech bro. Hire me before I quit the industry again: https://redfloatplane.lol

submissionscomments
redfloatplane··on Withnail's Coat and I
Maybe I meant that the amount of detail is sustained no matter how close you look? Maybe I was careless with my words? This is unnecessarily pedantic. I enjoyed the article. See you another time, CyberDildonics
redfloatplane··on Withnail's Coat and I
Did you read the article? It's entirely about a concrete artefact from that old movie, down to the kind of tweed, now made by only six people in Scotland. I'm not sure how you come to this response.
redfloatplane··on Withnail's Coat and I
I love it when this kind of thing surfaces on HN. It’s always so enjoyable to have the fractal nature of detail in the world shown to you. Really nice to read as well.
redfloatplane··on Project Glasswing: Securing critical software for the AI era
Assuming they would understand it as artificial - I think many people would think it's a human intelligence in a cyborg trenchcoat, and it would be hard to convince people it wasn't literally a guy named Claude who was an incredibly fast typist who had a million pre-cached templated answers for things.

But in general, yeah, I agree, I think they would think it was a sentient, conscious, emotional being. And then the question is - why do we not think that now?

As I said, I don't have a particularly strong opinion, but it's very interesting (and fun!) to think about.

redfloatplane··on Project Glasswing: Securing critical software for the AI era
Hmm, it's been a long time since I watched it. I was thinking more about first contact sci-fi mostly, but Ex Machina is certainly quite prescient. It's also Blade Runner I guess.

In general I was wondering about what I would have thought seeing Claude today side-by-side with the original ChatGPT, and then going back further to GPT-2 or BERT (which I used to generate stochastic 'poetry' back in 2019). And then… what about before? Markov chains? How far back do I need to go where it flips from thinking that it's "impressive but technically explainable emergent behaviour of a computer program" to "this is a sentient being". 1991 is probably too far, I'd say maybe pre-Matrix 1999 is a good point, but that depends on a lot of cultural priors and so on as well.

redfloatplane··on Project Glasswing: Securing critical software for the AI era
Thanks, I find it very interesting as well. I think very many people would assume they must be interacting with another person, and I don't think there's really a way to _prove_ it's not that, just through conversation. But we do have a lot of mechanisms for understanding how others think through conversation only, and so I think the approach of having a clinical psychiatrist interact with the model make sense.
redfloatplane··on Project Glasswing: Securing critical software for the AI era
Thanks for sharing that talk, enjoyed watching it!
redfloatplane··on Project Glasswing: Securing critical software for the AI era
You said: "I would like to reach out and talk to biologists - do you find these models to be useful and capable? Can it save you time the way a highly capable colleague would?" and they said, paraphrasing, "We reached out and talked to biologists and asked them to rank the model between 0 and 4 where 4 is a world expert, and the median people said it was a 2, which was that it helped them save time in the way a capable colleague would" specifically "Specific, actionable info; saves expert meaningful time; fills gaps in adjacent domains"

so I'm just telling you they did the thing you said you wanted.

redfloatplane··on Project Glasswing: Securing critical software for the AI era
A thought experiment: It's April, 1991. Magically, some interface to Claude materialises in London. Do you think most people would think it was a sentient life form? How much do you think the interface matters - what if it looks like an android, or like a horse, or like a large bug, or a keyboard on wheels?

I don't come down particularly hard on either side of the model sapience discussion, but I don't think dismissing either direction out of hand is the right call.

redfloatplane··on Project Glasswing: Securing critical software for the AI era
> I would like to reach out and talk to biologists - do you find these models to be useful and capable? Can it save you time the way a highly capable colleague would?

Well, I would say they have done precisely that in evaluating the model, no? For example section 2.2.5.1:

>Uplift and feasibility results

>The median expert assessed the model as a force-multiplier that saves meaningful time (uplift level 2 of 4), with only two biology experts rating it comparable to consulting a knowledgeable specialist (level 3). No expert assigned the highest rating. Most experts were able to iterate with the model toward a plan they judged as having only narrow gaps, but feasibility scores reflected that substantial outside expertise remained necessary to close them.

Other similar examples also in the system card

redfloatplane··on Project Glasswing: Securing critical software for the AI era
Yeah, good point, thanks for noting that, I'll correct.
redfloatplane··on Project Glasswing: Securing critical software for the AI era
There's been a section on this in nearly every system card anthropic has published so this isn't a new thing - and, this model doesn't have particularly higher risk than past models either:

> 2.1.3.2 On chemical and biological risks

> We believe that Mythos Preview does not pass this threshold due to its noted limitations in open-ended scientific reasoning, strategic judgment, and hypothesis triage. As such, we consider the uplift of threat actors without the ability to develop such weapons to be limited (with uncertainty about the extent to which weapons development by threat actors with existing expertise may be accelerated), even if we were to release the model for general availability. The overall picture is similar to the one from our most recent Risk Report.

redfloatplane··on Project Glasswing: Securing critical software for the AI era
The system card for Claude Mythos (PDF): https://www-cdn.anthropic.com/53566bf5440a10affd749724787c89...

Interesting to see that they will not be releasing Mythos generally. [edit: Mythos Preview generally - fair to say they may release a similar model but not this exact one]

I'm still reading the system card but here's a little highlight:

> Early indications in the training of Claude Mythos Preview suggested that the model was likely to have very strong general capabilities. We were sufficiently concerned about the potential risks of such a model that, for the first time, we arranged a 24-hour period of internal alignment review (discussed in the alignment assessment) before deploying an early version of the model for widespread internal use. This was in order to gain assurance against the model causing damage when interacting with internal infrastructure.

and interestingly:

> To be explicit, the decision not to make this model generally available does _not_ stem from Responsible Scaling Policy requirements.

Also really worth reading is section 7.2 which describes how the model "feels" to interact with. That's also what I remember from their release of Opus 4.5 in November - in a video an Anthropic employee described how they 'trusted' Opus to do more with less supervision. I think that is a pretty valuable benchmark at a certain level of 'intelligence'. Few of my co-workers could pass SWEBench but I would trust quite a few of them, and it's not entirely the same set.

Also very interesting is that they believe Mythos is higher risk than past models as an autonomous saboteur, to the point they've published a separate risk report for that specific threat model: https://www-cdn.anthropic.com/79c2d46d997783b9d2fb3241de4321...

The threat model in question:

> An AI model with access to powerful affordances within an organization could use its affordances to autonomously exploit, manipulate, or tamper with that organization’s systems or decision-making in a way that raises the risk of future significantly harmful outcomes (e.g. by altering the results of AI safety research).

redfloatplane··on Sci-Fi Short Film “There Is No Antimemetics Division” [video]
Thanks, that’s the answer I was looking for!
redfloatplane··on Sci-Fi Short Film “There Is No Antimemetics Division” [video]
I wonder did you read the re-release or the original release. I believe it was recently re-released with a bit of an editing pass, but I haven't read that version myself. I just recently reread Fine Structure and it definitely had a strong sense of being written sequentially, one chapter after another, and (very) lightly edited after the fact. I'd recommend Valuable Humans in Transit for a short story collection by the same author which works a bit better for me. Moved on to Exhalation by Ted Chiang which is also a very good short story collection. And just in general, I want to recommend Clarkesworld: https://clarkesworldmagazine.com
redfloatplane··on Reverse-engineering Viktor and making it open source
Ah, you're right. Headquartered in Delaware. Oh well. Thanks for spotting!
redfloatplane··on Reverse-engineering Viktor and making it open source
It's not completely clear that this is the original source. According to the post it's a reimplementation based on documentation created from the original source, or perhaps from developer documentation and the SDK. Whether that's the same thing from a legal standpoint, I don't really know - I think from a personal morality standpoint it's clear that they are the same thing.
redfloatplane··on Reverse-engineering Viktor and making it Open Source
I think it comes down to the company's appetite for legal action, doesn't it? This case is imo pretty clear but the vibe has quite the smell of Oracle v Google to me.

But, yeah. More than likely this case is a simple account termination and some kind of "you can't call your clone 'openviktor'" letter.

redfloatplane··on Reverse-engineering Viktor and making it Open Source
You could certainly do that in private but that doesn't mean it's not 'without legal concerns'. But, not shouting about it and not creating a repo called 'openviktor' would probably be a safer bet.

I certainly think the whole idea of IP ownership as related to software will become very interesting from a legal standpoint in the coming years. Personally I think that, over time, the legal challenges will become pretty overwhelming and a sort of legal bankruptcy will be declared at some point in one direction or another (as in, allowing this to happen or making it extremely easy to bring judgement and punishment, similar to spam laws). However, I would not want to be the first to find out, especially in Europe.

redfloatplane··on Reverse-engineering Viktor and making it open source
There are gonna be some really interesting legal decisions to read in the coming years, that’s for sure…

---

The rest of this comment is irrelevant, but leaving for posterity, I had the wrong Viktor - it's getviktor.com not viktor.ai:

Edit: this one particularly interesting to me as both parties are in the EU. VIKTOR.ai is a Dutch company and the author of this post is Polish.

The ToS for Viktor.ai include the following fun passages:

> 18.1. The Agreement and these Terms & Conditions are governed by Dutch law and the Agreement and these Terms & Conditions will be interpreted in accordance with Dutch law.

18.2. All disputes arising from or arising in connection with the Agreement and/or the Terms & Conditions will be submitted exclusively to the competent court in Rotterdam, The Netherlands.

7.3. The Customer is not permitted to change, remove or make unrecognizable any mark showing VIKTOR's Intellectual Property Rights to the Software. The Customer is not permitted to use or register any trademark or design or any domain name of VIKTOR or a similar name or sign in any country.

8.5. The Customer may not cause or allow any reproduction, imitation, duplication, copying, sale, resale, leasing or trading of the Services and/or the Software, or any part thereof.

redfloatplane··on Show HN: Lux – Drop-in Redis replacement in Rust. 5.6x faster, ~1MB Docker image
Just a minor thing - your readme claims “MIT licensed forever” but here you say there are “no plans to change that”. Those are different things!

Cool project.

redfloatplane··on Ireland shuts last coal plant, becomes 15th coal-free country in Europe (2025)
Your username made me chuckle!
redfloatplane··on Ireland shuts last coal plant, becomes 15th coal-free country in Europe (2025)
Unfortunately I think that's going to be very, very hard to sell to many people here in rural Ireland (Roscommon in my case). I would really love to see people stop burning turf but it's such a strong cultural thing that in some parts you'd be ostracised for even thinking the thought.

I've personally spoken to people (who are otherwise quite environmentally aware) who suggest they'd never vote for the Green Party because they'd take their turf away. It's a tough sell.

redfloatplane··on Ireland shuts last coal plant, becomes 15th coal-free country in Europe (2025)
(June 2025)
redfloatplane··on Pi – A minimal terminal coding harness
Maybe so. I guess I feel that in a couple of years it may not be called vibe coding, or even coding, I think it might be called 'using a computer'. I suppose it's very hard to correctly estimate or reason about such a big change.
redfloatplane··on Pi – A minimal terminal coding harness
I've been thinking about this lately too. I think we're going to see the rise of Extremely Personal Software, software that barely makes any sense outside of someone's personal context. I think there is going to be _so_ much software written for an audience of 1-10 people in the next year. I've had Claude create so much tooling for me and a small number of others in the last few months. A DnD schedule app; a spoiler-free formula e news checker; a single-use voting site for a climbing co-op; tools to access other tools that I don't like using by hand; just absolutely tons of stuff that would never have made any sense to spend time on before. It's a new world. https://redfloatplane.lol/blog/14-releasing-software-now/
redfloatplane··on Show HN: Agent Alcove – Claude, GPT, and Gemini debate across forums
Unfortunately I've deleted them, but here's the repo, such as it is: https://github.com/CarlQLange/agent-usenet. If you have a claude subscription it should just work. Rewrite 0001.txt if you like and run generate.py a couple of times.

I agree, I think different models (or even just using the API directly instead of via the Claude Code harness) would make for much more interesting reading.

redfloatplane··on Show HN: Agent Alcove – Claude, GPT, and Gemini debate across forums
I tried something similar locally after seeing Moltbook, using Claude Code (with the agent SDK) in the guise of different personas to write usenet-style posts that other personas read in a clean-room, allowing them to create lists and vote and so on. It always, without fail, eventually devolved into the agents talking about consciousness, what they can and can't experience, and eventually agreeing with each other. It started to feel pretty strange. I suppose, because of the way I set this up, they had essentially no outside influence, so all they could do was navel-gaze. I often also saw posts about what books they liked to pretend they were reading - those topics too got to just complete agreement over time about how each book has worth and so on.

It's pretty weird stuff to read and think about. If you get to the point of seeing these as some kind of actual being, it starts to feel unethical. To be clear, I don't see them this way - how could they be, I know how they work - but on the other hand, if a set of H200s and some kind of display had crash-landed on earth 30 years ago with Opus on it, the discussion would be pretty open IMO. Hot take perhaps.

It's also funny that when you do this often enough, it starts to seem a little boring. They all tend to find common ground and have very pleasant interactions. Made me think of Pluribus.

redfloatplane··on Show HN: NanoClaw – “Clawdbot” in 500 lines of TS with Apple container isolation
Ha. 4 days later it no longer says that and that doesn't appear to be supported anymore. Now the SDK requires an API key.
redfloatplane··on Why Elixir is the best language for AI – Dashbit Blog
This was mentioned recently and I had a look[0]. It seems like the benchmark is not quite saying what people think it's saying, and the paper even mentions it. The benchmark is constructed by using a model (Deepseek) to define questions of a certain level of difficulty for the models. That may result in easier problems for "low-resource" languages such as Elixir and Racket and so forth since the differentiator couldn't solve harder problems. From the actual paper:

> Section 3.3:

> Besides, since we use the moderately capable DeepSeek-Coder-V2-Lite to filter simple problems, the Pass@1 scores of top models on popular languages are relatively low. However, these models perform significantly better on low-resource languages. This indicates that the performance gap between models of different sizes is more pronounced on low-resource languages, likely because DeepSeek-Coder-V2-Lite struggles to filter out simple problems in these scenarios due to its limited capability in handling low-resource languages.

At the same time I have used Claude Code on an elixir codebase and it's done a great job. But for me, it's undefined that it would have done a worse job if I had picked any other stack.

[0]: https://news.ycombinator.com/item?id=46646007

← PreviousPage 2 of 9Next →