HNHacker News
TopNewBestAskShowJobs

wyrdcurt

148 karma · joined July 9, 2021

Curt@OptiMoss.ai
submissionscomments
wyrdcurt··on America.gov
AI-generated text gets auto-flagged (which is why I paraphrased instead of posting an actual transcript), I assume that's what happened here.

Interesting result they got, shows that my result isn't hard-coded, and it's the kind of answer I actually expected when I asked the question... I didn't expect the dodge, though I wasn't surprised. Makes me wonder what the rate of it giving a straight answer vs dodging the question is.

wyrdcurt··on America.gov
Paraphrasing a chat with America:

- What happened at the Capitol on January 6th, 2021?

- looks it up, finds dozens of government sources "I am not a history recap service."

- Who was the first president of the US?

- "George Washington took the oath of office on April 30, 1789"

- I thought you weren't a history recap service?

wyrdcurt··on MicroLLM Lab – Try 7 tiny LLM's in the browser
Not sure how you're prompting it but remember that it's not trained for chat or instruction following, it simply takes the text given to it and tries to continue it. Give it the right prompt structure, and it can (at least sometimes) output coherent completions, far more often than you'd see in a Markov-chain. Also, this version is more or less equivalent to the smallest version of GPT-2; the largest version was 1.5 billion parameters and was much more likely to generate impressive (at the time) output.

The assumption that LLMs would always need sophisticated inputs to generate useful outputs is where the term "prompt engineering" came from. Now that idea is basically dead. Absolutely wild how far these models have come in less than a decade!

wyrdcurt··on LA Metro has some of the slowest escalators on Earth
Probably was auto-flagged for being AI-generated, looks Claudish. Brand new account too.
wyrdcurt··on Flock is rolling out a voluntary severance program
Related from a couple days ago: https://news.ycombinator.com/item?id=49762835
wyrdcurt··on Inside ZCode: Silently uploading your Git history to the cloud
ZCode is pretty bloated anyway, in my experience. I used it for a while because Z.ai offers a subscription usage multiplier for using it, but despite that, I found myself hitting limits less often when I switched to Pi (and performance is the same, if not better).
wyrdcurt··on Child Sexual Abuse Material Persists on X
I didn't say there was a good reason, in fact I agree there is no good reason to share such videos for entertainment.

I asked a specific question about your definition of gore, so I could try to engage with your opinion in good faith without making assumptions. But instead of answering the question, you dodged it with a cherry-picked example that doesn't represent every scenario that I'm considering. That makes me inclined to believe that your position isn't particularly well-considered.

Maybe the reason you get downvoted whenever you bring this up while "nobody has a good argument on why it should remain legal" is because you don't provide enough substance for anyone to argue against...

wyrdcurt··on Child Sexual Abuse Material Persists on X
What precisely is your definition of gore? Could you define it in a way that makes it illegal to share images of murders and car crashes as entertainment, without criminalizing, say, journalists documenting a war?
wyrdcurt··on An Alien Mind
It hasn't even been a century since nuclear holocaust became possible. Hardly any time at all on the grand scale. "They didn't" could just as well be "we haven't, yet".
wyrdcurt··on GPT-6 Astra
More optimistic take: we'll only be second-class for a few months, if the pattern of Chinese models catching-up holds.
wyrdcurt··on SteamdDB Joins Nexus Mods
> they are both communities of people that others do not want to be around

Speak for yourself. Only one of those communities is one I don't want to be around, and there are plenty of people who agree, apparently including the operators of Nexus.

> Removing political propaganda that got injected into a game is a better experience for the average gamer.

Let's be clear: the "political propaganda" you're referring to is more accurately described as "gender-neutral language". If you're going to assert that "the average gamer" will have a better experience playing games without gender neutral language, maybe you can back that up with some citations -- unless, perhaps, your idea of the "average gamer" is merely based on the anecdotal experience of yourself and your peers?

I'm fully aware of what a quality of life mod is, by the way. Just now I was playing a game with "Body Type A" and "Body Type B" character creation, in fact, and have no inclination to change that (even if it did bother me, it was there for about 5 seconds at the beginning of the game, never to be shown again... so what?). Just last night, though, I made my own quality of life mod to address an inventory management quirk that was frustrating me. I'm very familiar with the term, which is exactly why I found the idea of it including "removal of gender neutral language" absurd enough to comment on.

> Personally, I'd prefer the platform to remain neutral and let people ignore mods they don't like.

And you have the right to have that preference! Just as Nexus has the right to operate their site in line with their preferences. If you don't like it, you don't have to use it. You could even put your money where your mouth is and start your own "neutral" site. If that's what the "average gamer" wants, perhaps it'll be wildly successful! Although it seems more likely to me that it will attract a certain crowd while repelling everyone else, you are free to prove me wrong :)

wyrdcurt··on SteamdDB Joins Nexus Mods
Can't imagine any reason why anyone who wants trans people to feel welcome would also want nazis and their friends to feel unwelcome... couldn't have anything to do with the fact that nazis systematically rounded up trans people, executed them, and erased their culture, could it? /s

More directly, are you really using a false equivalence to counter an "ad hominem"?

"Quality of life", that's a euphemism for queer erasure I haven't heard before... and here I was thinking a quality of life mod was one that fixed a janky UI or smoothed out a grind, that kind of thing.

You certainly could complain about a "trans bar problem" if you like, and you could even go make your own nazi bar if not having one bothers you enough -- though it's pretty clear what that would say about your values.

wyrdcurt··on Gemini 3.8 Flash and 3.8 Flash Cyber
Pretty typical "cool HTML toy" LLM output, tbh. The only thing impressive about this is how fast it generated it (13 seconds is wild!), but that's more of testament to Google's infrastructural advantage than to the quality of the model.

For comparison's sake, I tried something similar with a couple other cheap models I've used lately, with the prompt "Impress me. Make something cool in HTML. Ensure that it is mobile friendly." (Added the mobile condition as I was on my phone when I did it).

Mimo-2.5 created something similar, only a bit less complex than Gemini's (though, at least the FPS counter is real!), in a minute or two for about 1/3 of a cent: https://gisthost.github.io/?740c325c21e9bfbee59c4f94d9aab0af

GLM-5.3-Flash, currently my workhorse model, spent 12 minutes (ouch) thinking about the prompt. Didn't cost me anything directly because I have a GLM sub, but I did the math and it would have cost about 1.1 cents through the API. Turned out nicely in my opinion (though in reality, it still isn't really anything special): https://gisthost.github.io/?9ef050e16cec2561e6504e725a3f0bcc

Side note: thanks for setting up that Gist Host tool, it's very convenient!

---

Editing to add this bonus from Mercury-2.5-Preview, which I just learned released a couple days ago. It's much less impressive-looking than any of the above, but it cost less than 1/20th of a cent, and the response was generated effectively instantly: https://gisthost.github.io/?02f40b50aa891bf396bfaaa3a7998203

wyrdcurt··on An investigation into the state of corvid–human relations
Interesting fact you may or may not already know: the pigeons most commonly found around human habitats are feral domestic pigeons. They generally aren't bothered by us because they're bred not to be! Corvids on the other hand are 100% wild (which makes their interactions with humans all the more fascinating, imo).
wyrdcurt··on Timeline of the OpenAI accidental attack against Hugging Face
"AI will probably, most likely, sort of lead to the end of the world. But in the meantime, there will be great companies..." - actual Sam Altman quote, the man is so unhinged he's beyond satire
wyrdcurt··on AMD acquires Taalas to boost inference performance by etching models in silicon
No, that's backwards. OpenAI are the ones investing in Cerebras. Part of the deal is that they can't sell to Anthropic.
wyrdcurt··on Open-weight AI is having its Kubernetes moment
They are weights for matrix operations, so in principle you'd think they would be deterministic. In practice it's more complicated than that.

Not to be snarky or dismissive, I mean this genuinely: ask an LLM about it. I currently have a headache so I'm not up to explaining the technical details, but they are interesting and worth reading about.

wyrdcurt··on Open-weight AI is having its Kubernetes moment
Except that even the exact same model won't output the exact same results, that's a fundamental aspect of how LLMs work. They're probabilistic/stochastic, not deterministic.
wyrdcurt··on A taxonomy of omnicidal futures involving artificial intelligence (2025)
Eschatological is a perfectly fine word. The post is about doomsday scenarios, but perhaps the commenter is equally tired of people who think AI will be salvation!

Anyway, if the word confuses anyone, they're in luck: dictionaries still exist :)

wyrdcurt··on Who's afraid of Chinese models?
So if the detector isn't perfect, it isn't useful? Not sure I buy that.

Also, even if there's no way to detect what the activations are doing, we already have the ability to analyze your proposed threat statistically. If the model repeatedly uses insecure libraries in most trials, then yes, in that case it would be prudent not to trust those weights.

Assuming one doesn't get banned for violating some ToS clause about using a closed model for LLM research, it could be possible to run those evals on a closed model too (likely at much greater expense). But there's a big difference: if such a statistical anomaly is discovered in an open model, one could potentially fine-tune that behavior out of it. With a closed model, that won't be an option.

Whether it makes a difference to the US government or not is beside the point. Even with a perfect solution, the current administration could do some mental gymnastics to achieve whatever political outcome they want. I’m not trying to make a political statement here, my point is technical: open-weights at least give us the possibility of visibility into why they generate what they do; this simply isn’t true with closed models.

wyrdcurt··on Who's afraid of Chinese models?
But that would happen with literally any model without grounding and is more of a quality/competence issue, not what's being discussed. It'd be a bit of stretch to conclude open models are no more auditable than closed models based on that possibility alone.

Also, in that case, there would likely be activations indicating that it is favoring a specific version. If that's an insecure version, sure that'd be suspicious... but again, you're only going to be able to verify that's what's happening in an open model.

Maybe you can illustrate a realistic scenario in which that would be a problem, otherwise I don't really understand what your point is in this context.

wyrdcurt··on Who's afraid of Chinese models?
Not sure I understand that position. Unless we're talking about a scenario in which one is using an outdated model along with no grounding (which, imo, PEBKAC), why would the model be pinning insecure libraries?

If it does have grounding, and can therefore see that it's introducing vulnerabilities to the code it's generating, yet does so anyway... I suppose we could invoke Hanlon's razor, but if the model is that incompetent, it probably isn't the right tool for the job regardless of its provenance.

That said, we aren't talking about incompetent models, we're talking about models sabotaging projects due to hidden motives. My point, again, is that those motives could potentially be revealed with open-weight models, in a way that will never be possible with closed models (barring some sort of legislation requiring independent third-party interpretability audits, which I suppose is in the realm of possibility).

wyrdcurt··on Who's afraid of Chinese models?
I'm talking about using mechanistic interpretability to see the model's intent. If it is deliberately using compromised libraries to weaken some code's security, there's going to be a signal in its hidden activations that it's doing so.

Finding these kinds of activations is something Anthropic is actively researching [1] but they're the only ones who can use those techniques to see Claude's intent. On the other hand, if a model is open-weights, in theory whoever is running the model could look inside the activations at runtime to see if a hidden vector associated with "deception" or "sabotage" is being activated [2].

[1] https://transformer-circuits.pub/ [2] https://arxiv.org/pdf/2509.03518

(Those sources are just a couple of relevant starting points I could find without much effort, there is also https://www.neuronpedia.org/ if one is interested in seeing interactive demonstrations of interpretability concepts)

wyrdcurt··on Who's afraid of Chinese models?
Maybe I should clarify. As I understand it, the kind of vulnerability being discussed is something like a Chinese model invisibly "realizing" that it's working on an American project, and then deliberately leaving subtle security bugs in its generated code for Chinese hackers to later exploit. As far as I know, that scenario is hypothetically possible, but has never been demonstrated to happen in the wild. Admittedly, I could be wrong about that! If anyone has evidence to the contrary, I'd love to see it.

Of course, one could retort that gathering that evidence may be nearly impossible now, but my point stands: in the future it might/probably will be possible to properly audit open-weight models. Closed models, on the other hand, will always be a black box.

wyrdcurt··on Who's afraid of Chinese models?
In my opinion, the big issue with that argument is that advances in interpretability research and steering conceivably could, and probably will, render moot that (as of now, purely hypothetical) risk of subtle sabotage for open-weight models... but not for closed models.
wyrdcurt··on I'm Making Strandfall, a Solarpunk Orienteering Larp
"Strandfall is a sci-fi story set in a post-apocalypse, with an emphasis on climate change, co-operation, and adaptation."

I'm in the early stages of designing a world/game focused the same topics, in part because I simply want more stories like this to exist. I live an ocean and a continent away, so of course I won't be participating, but I don't know if I've ever subscribed to a newsletter so quickly. Looks like an absolutely brilliant idea and I hope it's successful!

I'd be extremely interested in participating if a version of this ever came to my corner of the world. Even if it doesn't, just knowing about it is inspiring and motivating :)

wyrdcurt··on Stenchill: 3D Printable Solder Paste Stencil Generator
As someone who has only recently been thinking about learning how to solder and work with PCBs that's actually useful info, thanks!

Unfortunately I do have experience with the smell, being around others working with improper ventilation... I'll be passing that advice along to my BIL too, ha.

wyrdcurt··on Stenchill: 3D Printable Solder Paste Stencil Generator
Sounds like you need a fume extractor!
wyrdcurt··on S&P Global has lowered Oracle’s creditworthiness from BBB to BBB-
True. The linked article's title says that. I wonder if that was a typo by the OP or one of those HN quirks where the title was automatically changed when it shouldn't have been.
wyrdcurt··on A voxel Tokyo in real Japan time – ride the Yamanote line and study Japanese
It's wild if it's running at 400 fps because nobody has a screen that refreshes at 400Hz. Every frame rendered past the screen refresh rate is wasted compute. Easily solved by limiting the frame rate.
Page 1 of 2Next →