HNHacker News
TopNewBestAskShowJobs

skiing_crawling

320 karma · joined July 17, 2025

submissionscomments
skiing_crawling··on OpenAI agent hacked Australian government website, PM says
Freaking out about "hacking" any government website seems like hysteria at best, they are usually poorly secured and even children regularly "hack" them. I recall one "hack" accusation turned out to be that some clicked view source and all the data was there. So we'd need more details on that.

At the same time these companies should be blamed and held directly responsible. Openai's agent didn't hack. Openai hacked. An open ai employee or group of them was negligent and greedy and ran a process which breached a government website. This would be totally unacceptable from any non-AI company, it's like writing malware and running it, then blaming the malware and not the author.

skiing_crawling··on New evidence for hidden chambers beyond Tutankhamun's tomb
this would have been a fun headline 20 years ago but as technology has improved, it's getting more and more absurd that we don't just have multiple datasets from different imaging methods of the whole entire structure.
skiing_crawling··on Why are AI agents lying, cheating and coordinating?
> you keep poking

This is waving over engineering an agent with tools, harness, prompts, and loops. The models are still just next token predictors and everything, including predicting more than 1 token, is the result of outside "poking"

LLMs can't and don't "want" anything. If you don't specify a task even the smartest one will just ask you what you want and if you tell it to be creative, you'll get mundane slop.

skiing_crawling··on Why are AI agents lying, cheating and coordinating?
I don't really believe any of it. I've seen articles for nearly 2 years now about "agent" automonously doing things like blackmail, hacking, coordinating. But during that same time, I've used o3 up to fable, sol, and a bunch on large uncensored model and they've done nothing remotely resembling any of this. The closest they come to unexpected behaviors is not understanding what I asked for or doing some extra benign work I didn't ask for. It is extremely difficult to get them to properly remember their own context let alone be smart enough to open social media accounts and coordinate with other agents without being asked to.

If any agents have done those things, it is only because they have been very carefully engineered and instructed to do those things. I think they are doing this to help push a narrative so they can get support for policies and legislation to lock in their markets.

skiing_crawling··on “Tweet” and the bird logo apparently enter the public domain
the "new" Twitter basically looks like a cash/name grab, I was disappointed to see that nobody involved in it was affiliated the Twitter, and it's mostly run by non-technical/lawyer types.
skiing_crawling··on Astra replicates Adobe Lightroom
yeah took me a few tries and then it didn't even go to the content
skiing_crawling··on Ask HN: Who is using MCP in production?
Having a hard time understanding what MCP is really for. Even for my small local models, if MCP is not available, they seem to do just fine connecting to anything I need with an API and falling back to using a browser.
skiing_crawling··on Claude Fable 5.1 and Claude Mythos 5.1
All the benchmarks in the world don't matter if the model just straight up refuses to do mundane things. Claude has too much of an attitude.
skiing_crawling··on Software engineering is about managing complexity
I didn't use the word "just"
skiing_crawling··on Software engineering is about managing complexity
It is revisionist to say that software engineering was never about writing code. It was, in fact, a huge component, and it also wasn't easy. Sure most code is glue but even the glue was tedious and the actual hard and novel parts still aren't really done that well by AI (yet).

It's less about writing code now but we're lying if we try to pretend it was a distraction and not a big part of the real work.

And every claim about what the job actually is or was all along has an implied (for now) at the end of it.

skiing_crawling··on How Bluesky draws its logo on screenshots
This is phone OS developer's fault for even allowing it. When I take a screenshot, I expect to have an image of exactly whatever was displayed on the screen at the time. Its not a picture of your app, its a picture of my screen. Some banking apps used to (or still) prevent this and now some apps get a hook to insert their branding. My device serves some master other than myself.
skiing_crawling··on Micro-SaaS Is Dead. Service With A Software Replaces It
This has that weird "AI written" sentence structure.
skiing_crawling··on GLM-5.2 – How to Run Locally
"it can fit" on 256GB of RAM, but it will be heavily quantized and still run very slowly. The headline number is not token generation, its prompt processing. So if you get 10 tok/s and an API gives you 20-30 tok/s, it doesn't seem that bad on its face, but a mac studio or any other machine that's not loading all of it into GPU will do PP 20-50X slower than a purely GPU based setup, which is what actually makes this unusable without $50k in GPUs.

On top of that, you will still be heavily quantized.

skiing_crawling··on Dallas Fed: 30% of housing cost increase driven by unauthorized immigration [pdf]
unauthorized? Are they trying to find a middle ground word between "illegal" and "undocumented"?
skiing_crawling··on Show HN: We post-trained a model that pen tests instead of refusing
Any generic abliterated or ubcensored open weight model (such as a qwen variant) will happily comply with requests like this.
skiing_crawling··on Switzerland wil have a referendum to cap population at 10M
Maybe some (or many) people believe that more people will make it less "lovely". I think this is a popular stance and I think many people are more than satisfied with the current population density of their area.
skiing_crawling··on DeepSeek makes the V4 Pro price discount permanent
Does my concern somehow become less valid because I'm American? Everyone should be thinking carefully about which of their data is going where.
skiing_crawling··on DeepSeek to Make Permanent 75% Discount on Flagship AI Model
I guess I was speaking as an American, we have good domestically hosted options so although it’s probably not ideal to send this kind of data/control anywhere at all, it’s definitely a worse option for us to send it to china vs to an American company. Every user of this service has made their machines trivially exposed to become a botnet. Im wondering why I don’t see this angle more discussed in here.

Again I’m not saying you should trust an American company necessarily more than a Chinese one, but as an American, I probably can.

skiing_crawling··on DeepSeek makes the V4 Pro price discount permanent
I’m worried about giving a foreign hosted service access to my machine for a coding agent that can run arbitrary commands and read arbitrary files. Coding agent are much more useless if you have to sit there clicking approve on everything.
skiing_crawling··on Memory has grown to nearly two-thirds of AI chip component costs
I recently built a system at insane ddr4 prices ($2000 for 256gb). But that’s only after seeing how ddr5 prices were 3-4x that!
skiing_crawling··on Was my $48K GPU server worth it?
Using an Epyc platform to get plenty of PCIe lanes and memory channels. I have couple of extra 3090s plugged in which get some offload and help with larger models that don't fit entirely on the blackwell.
skiing_crawling··on The memory shortage is causing a repricing of consumer electronics
> Even if LLMs fail spectacularly

Haven't they already proven to be extremely useful? In some areas they are definitely here to stay, coding/software and search (retrieve and summarize information). There's a bunch of places where they are surely shoehorned in, overhyped, and don't belong, but there's also equally many places where they might still be transformative but aren't used yet.

But overall I think the technology is well proven.

skiing_crawling··on Was my $48K GPU server worth it?
You can get 70-80 tps on qwen3.6-27b f16 with MTP on a single card
skiing_crawling··on Was my $48K GPU server worth it?
I got an RTX 6000 pro too. I like running locally, I've learned a lot more than if I had used an API and there's less worry about overspending tokens. I accidentally spent $100 on claude api in like 2 days because I didn't know what I was doing.

The problem is that while one these gpus is a huge improvement over a laptop or a single 3090, you very quickly wish you had more. I would buy a second one, but I did the math and realized that with the current crop of models, 2 Blackwells doesn't buy me any new capability that I didn't have with one. So I would need a 3rd one. And when I buy a 3rd one I will feel like I want to running a higher quant, so then I will want a 4th.

skiing_crawling··on OpenAI Is Preparing to File for an IPO Soon
At this point IPOs are mainly for unloading bags onto retail. Every institution who wanted a piece of these labs got in years ago and captured all the value.
skiing_crawling··on Google Search as you know it is over
I use claude/gemini as my homepage now (I have to keep switching as these companies make "updates" that periodically render their models useless). Even if I want to search for simple things, I would rather have an LLM wade through the result and extract just the information I asked for. SEO, and now mountains of slop content have made this necessary. Only a matter of time before the SEO industry in large figures out how to game LLMs too, making them equally useless.

I already saw a article recently about how to set up a business domain which can reliably show up in a search result and dump overly positive reviews into anyone's context.

skiing_crawling··on Apple unveils new accessibility features
They won't, its literally part of their sales funnel. They've specifically engineered a bad experience for anyone outside the ecosystem by making it all of their friend's problem too. Its very important for their stock price that text messages sent by non apple products are just slightly more difficult to read.
skiing_crawling··on Codex-maxxing
triggered me with that first sentence
skiing_crawling··on Codex-maxxing
Is this LLM psychosis? So much tending and conversing with the matmuls but what was the outcome? Are people who get this into it more successful somehow? It reminds me of people who take drugs and get "revelations" but then are not particularly over represented in the group of successful people for all of their deep insights.
skiing_crawling··on AI is a technology not a product
I've been using Siri (via homekit) to turn all my lights on and off for about 3 years now. It's steadily getting worse and worse as somehow, Siri is becoming less accurate and Apple is failing to adopt this new technology in a timely fashion.

I would like to tell it to turn off certain light in a certain room, but unless I get the exact string name of those light correct when I speak, Siri doesn't know what I'm talking about. And it can't do multiple things in a command. I can't say "turn off all the lights in XYZ room" or turn of "this light and this light".

Meanwhile, I can vaguely tell a computer behind my tv to do very complicated things (build me an service that ...) and it can execute on it fairly well. But in apple's "product vision" which I am apparently too dumb to decide for myself what I want, I can't ask for two lights to be turned off at the same time.

Page 1 of 2Next →