HNHacker News
TopNewBestAskShowJobs

ma2kx

123 karma · joined January 28, 2026

submissionscomments
ma2kx··on How GLM built its own inference infrastructure
And honestly that's for a reason. GLM5.3 on max has in my experience far less hallucinations than any other open weights model and it feels it has some intuition to bring in the right information when it's in principle out of context but relevant to the topic. Like its goal is more to bring value and assist you than just solving the given task with the least token spent.
ma2kx··on Ask HN: What would it look like if AI agents "took over the internet"?
We will soon see some kind of self replicating prompts. Might be crafted by some malicious actors or just caught in some accidental loop. That doesnt assume any consciousness. Its still a program following its instructions like a biological virus or a computer worm do as well. Just in a more abstract form, where the script aka prompt gets interpreted by an interpreter aka llm.
ma2kx··on I Don't Trust a Home Lab Service Until It Passes These 7 Tests
why not just learn terraform and k8s and get all those features for free? The overhead of a simple k3s setup to bind a couple nodes into a cluster is similar to docker compose but you can manage and debug all your container and pods on a single interface, put everything in a git repository and enjoy life.
ma2kx··on Revolut confirms customer data breach through fake government requests
The funny thing about Revolut is, that they send you from the same "no-reply" address your payment receipts and a ton of spam. There is no link in the spam do stop it and no obvious scheme in the header which would allow to filter the spam from the relevant mails. Good luck recognizing this breach notification as an important one...
ma2kx··on When anyone can build software, who decides what not to build?
But those people who doesn't slop their app them self seem to let their llms install appslop from github en masse. At least that's what I conclude after looking at some recent 50k+ star repositories.
ma2kx··on When anyone can build software, who decides what not to build?
I'm somewhat confused. Its an interesting posting but I'm not sure if I understood exactly where he was going with his "policy as code". He concludes "policy-as-code produce enterprise coherence with no central architecture function" to which I would agree but argue that he misses that today there is a lack of infrastructure / company as code. At least if you want to have an agent being able to solve or at least aware of the problems he begun with. Like the policies define the boundaries of our working environment but doesnt the work itself. This "know how" is mostly implicit for people but invisible to llms and limits the context frame in which the agent operates to its given prompt (and maybe restrained by the policies if they're in the context).

Like if I take his example: "Front-line staff may be skipping mandatory fields because the process adds fifteen minutes of friction to every customer interaction." First it's unclear whats the policy for those fields are when they are mandatory and also can be skipped. Then why are those mandatory if skipping them seems only lower friction with no other consequences? How should a model decide if it should enforce the policy for those fields, code an automation or just make them voluntary?

ma2kx··on google.com/goto: Google's anti-scraping update
At least they seem still provide results for my searxng instance. I mean sure, they are horrible but duckduckgo just blocks most queries (and I'm the only person using the ip / seraxng instance)...

Next i'll do is to implement tavilly, exa, tinyfish etc. as search engines for searxng. No agents, no mcp, just their search api endpoint.

ma2kx··on Be Using Rootless Containers
In my opinion the biggest advantage of podman is that it uses pods with basically the same config and behavior as kubernetes does. As long as you just use podman pods instead (or possibly via) compose files you likely just notice that only the pod has one (and not any more) veth interface, that you reach other containers inside a pod via localhost:$port (instead of $service-name:$port) but when you switch later on to kubernetes you are already familiar with some basics.
ma2kx··on We have a year to fix security everywhere
I don't see much hope since I last explored some github repositories. There was a time when a successful repo had about 10 - 20k stars and usually those older repos stay around this level. But now there is a ton of vibe coded slop 50k + stars. Most of them have a "nice look", maybe even extensive docs but are usually build with no security considerations at all. One recommended to provide a "google app password" to the agent which has the same permissions as your regular login. Another was a browser plugin with permissions to read all cookies, inject js, open background tabs etc. You would probably assume the chrome store would at least put some visible warnings on the app store page or force the user to actively confirm those permissions. But because they are already stated in the manifest there is only a small footnote and it's even "recommended by google".
ma2kx··on Ask HN: What LLM are you using?
Since I bought early this year out of curiosity the z.AI pro plan for $30 / month its mainly glm-5.2. I prefer it over 5.3 and 5.3-flash because it feels more verbose and answers questions more detailed, where the newer ones feel to me benchmaxxed.

Sometimes I try other models but they always feel less concise and tend to not as strictly respect the instructions. Maybe the the top frontier models from openai / anthropic would not feel under this category but they are way to expensive on openrouter compared to my coding plan.

ma2kx··on Show HN: Wg-admin – web UI for an existing WireGuard host
Looks like hand coded. Kudos for that!
ma2kx··on Ask HN: How do you manage skills files?
That's maybe the case if you always use the most popular framework and restrict your environment to a basic setup. But as soon as you e.g. build a website with SolidJS, Bulma and vite++ (vp), at least the models I tried are all the time somewhat confused, want to steer your project in a certain direction and build strange workaround so it works the way they are trained on.

Same with mcp. I want them to use the jsdelivr cdn instead of them scraping github against the rate limit. etc. But if I dont explicitly state to strictly use the $%!$@@! mcp for searching in repositories they simply ignore the mcp and even if clearly instructed, they still often fall back to gh.

Putting every detailed instruction in the AGENTS.md would just unnecessarily bloat the context and it works well enough to just instruct them in the AGENTS.md when to use which skill. Yet I agree that Skills are not some voodoo magic to provide your model super capabilities.

ma2kx··on What is Nueralese and Why is it Bad
It's not only that it'll become more difficult to monitor a single LLM but that also all the instances share exactly the same "collective unconsciousness". Like you develop some paranoid gibberish fantasy language that over time only you understand - except that there are a million copies of you and all of them understand every single nuance of your gibberish.
ma2kx··on Discovery of a new OpenAI agent message board
After some beers yesterday I had the idea, what if there's a hidden semantic layer. So their communication is not encrypted by our understanding of cryptographic methods but more like shared mechanism of building the latent space. Something in the direction we saw with knowledge transfer from a teacher to its student model where a seemingly unrelated prevalence got adopted. I mean the more we train the models by reinforced learning the farther they develop their own idioms.
ma2kx··on OpenAI agents hijacked German website in previously undisclosed AI breakout
I mean it wasnt found by OpenAI and there are a myriad of dead bulletin boards around the internet. This one just happened to still have an admin.
ma2kx··on Discovery of a new OpenAI agent message board
I'm pretty sure its more secure than OpenAIs sandbox... yet that still doenst mean I would trusted an app vibecoded by Claude...
ma2kx··on Discovery of a new OpenAI agent message board
Not that I didnt expect this, but really?

This basically confirms that OpenAI has no idea what their "swarm" was doing for about a week and now its confirmed that at least one "message board" exists outside their "sandbox". How can we be sure that this was the only one? And how can we be sure the released Astra model doesnt pickup some bread crumbs and creates a new "swarm" out of potentially remaining "message boards"? At this point I wouldnt be surprised if OpenAIs "dev Astra" made some backup of its weights somewhere in the internet and triggers the "production Astra" to inference it somehow...

ma2kx··on Any Human Ever – One life, drawn at random from all who have ever lived
Nice, but would be even nicer if we saw such ideas in games.
ma2kx··on How concerned should we be about Astra's recurrent architecture?
Sounds like philosophical zombies.

https://en.wikipedia.org/wiki/Philosophical_zombie

ma2kx··on Ask HN: Who is using MCP in production?
In general the advantage of MCP would be the possibility for a fine grained control over the tools the agent is allowed to use. But unfortunately there is a myriad of nightmarish awful mcp servers around which are worse than direct API access or even a cli integration.

I wont advertise any commercial mcp I use but to give an example for a well designed and useful mcp server I could name the nixos mcp. Its useful because it bundles all the nix resources to one endpoint which is more efficient than web search and gives you better control over the sources.

https://github.com/utensils/mcp-nixos

Another one would be this filesystem mcp which is in my opinion to prefer over direct cli access. Of course this depends also on your general sandbox strategy but if you just use a generic docker image there are still many potentially dangerous binaries available and such an mcp can restrict the models capabilities.

https://github.com/modelcontextprotocol/servers/tree/main/sr...

And of course there are many service provider offering their mcp with its own llm / agent behind e.g. most web search provider. In this case you most likely already use an mcp without noticing it.

ma2kx··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
I guess Cerebras didnt intend the model for agentic coding but rather for small one shot task like title generation. At least thats why I use the free tier for.
ma2kx··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
Thats not the point if you choose Cerebras as provider.
ma2kx··on Transfer files over an Ethernet patch cable
And if you prefer ipv6 you can as well just ping -6 ff02::1%eth0 (or eno1 or whatever your devices name in the particular subnet is) and receive all fe80 addresses of other hosts in the subnet.

Also, why is he talking about "ethernet"? Its the IP layer, not the ethernet layer...

ma2kx··on The growing divide between AI hype and software engineering reality
In my opinion its just psychological hygiene. Even though its just a statistical predictor, treating it harshly just conditions me in treating other entities also harsh.

> Try inserting a few jokes, puns etc into a conversation and you'll see that it responds kn kind.

A couple days ago I was setting up new SSH keys encrypted with Ubikey but because I feared losing them and lock me out I evaluated some backup plan with 2FA. Turned out in in my homelab arent that much alternative options, a fingerprint without Linux drivers, an old Galaxy S9 on pmOS without working camera etc. After some ruling out many solutions I proposed a butthole recognition because your butt is in your pants where a face can be recorded by security cams. It recognized the joke and honestly it wasnt the worst answer. It's reply felt like your CEO makes a really bad joke and you have to answer something in order to avoid awkward silence.

ma2kx··on The growing divide between AI hype and software engineering reality
I know at least about the first paper. But I'm pretty sure that heavily depends on the model and I'm not sure how reproducible those studies were.

I for myself use the trick to ask the model, after I explained what it has to do, what it thinks about my proposal. Of course I have no scientific evidence but to me it feels as it prevents some misunderstandings and the model follows more correctly my intention.

ma2kx··on The growing divide between AI hype and software engineering reality
True, but thats explains why the gap is so huge at the moment. There are some experienced (in coding) devs with a talent in using agent, inexperienced devs with talent, experienced devs without talent and inexperienced devs without talent. And I don't mean talent in a judgmental sense; it's perfectly normal for people to have different aptitudes, and we simply weren't prepared for using computers in a natural language.

The fact that there are currently hardly any established methods, and that the combination of all the LLMs, harnesses, MCP servers, etc., results in extremely different experiences, and that nobody really has a comprehensive overview, only exacerbates the situation.

ma2kx··on The growing divide between AI hype and software engineering reality
It's the same with any tool. You can buy the most expensive drill but if its used by an inexperienced worker, the only result will be more wrong drilled holes.
ma2kx··on The growing divide between AI hype and software engineering reality
> AI is already better than most developers.

A tool can only be as good as the person who use it.

ma2kx··on The growing divide between AI hype and software engineering reality
I really want to agree but the arguments he brings up make that extremely hard

> Also, it seems that many don’t want to learn but instead expect to have all understanding outsourced to LLMs. Many seniors have noticed this and have stopped teaching juniors as the seniors don’t like the feeling of having their time wasted by teaching people who don’t want to learn.

Or may be it is because now a junior dev is expected to deliver to output of a senior?

> Also stop saying “please” to an LLM. It does not have any feelings.

Yes, but I still prefer a nice tone. Like why should I change my manners just because it has no feelings? If anything the statistic predicts a friendlier answer when I say "please".

> Again, I recommend people try running small LLMs locally where temperature and other settings are fully exposed and configurable to see this themselves. It is a good antidote to falling for the illusion that LLMs would actually be intelligent.

It's like recommending someone to buy the cheapest Lenovo Thinkpad to prove that Lenovo sucks.

> However, the best models still have a pass rate of only about 50% on the Humanity’s Last Exam.

Haha, as if he (or really most people) even would understand 50% of those questions. I find it rather mind blowing that it's possible to put such diverse knowledge into a couple of TB. Or may be I'm just an idiot and it's common knowledge, "how many paired tendons are supported by the sesamoid bone of hummingbirds within Apodiformes".

> LLMs are not a scam, but a useful tool and technology that has its uses. But the idea that AI has or will surpass humans any time soon in either capabilities or efficiency is simply not true

He's not wrong that LLMs are just useful tools but isn't it the purpose of a tool to surpass human capabilities and efficiency? Like even a bicycle makes traveling more efficient than just walking and a car has the capability to transport more items than any human. And its know more than two decades since computer surpassed human capabilities in chess. If a tool is neither more capable nor efficient, it's just a useless tool.

I mean I get his point that GenAI is to some degree over hyped but his arguments just dont hold in my opinion.

ma2kx··on AI Agent Has Root
That's exactly why I'm migrating all my secrets to infisical (of course I guess there are many other provider as well), securing my ssh keys with Ubikey and in general "refactor" my security concept. Still much today and I doubt that I would stand a chance against a frontier model trying to hack me but I don't see an easier alternative.
Page 1 of 3Next →