HNHacker News
TopNewBestAskShowJobs

nsoonhui

15,398 karma · joined December 18, 2007

submissionscomments
nsoonhui··on MiMo v2.6
I did try to use Chinese open models, but for my production work they simply couldn't cope at all; both GLM 5.3 and Deepseek v4 went into infinite loop and wasted my tokens until my OpenRouter wallet reached 0; good thing I didn't enable the auto topup. US models, by contrast, breezed past them. Even for simpler tasks, Chinese models took long time to complete, and I needed to supervise closely. The price , in the end, didn't come cheap, mainly because too much time wasted on thinking.

So maybe one day Chinese models will squeeze out the American ones, but today is not that day.

So no, I am not excited about Chinese models ( just because its open weight and not American).

nsoonhui··on AI coding has made CI a bottleneck, so we reworked ours to keep up
It may not have to do with increased revenue, but it does have everything to do with reduced cost, especially developer's cost.

I bet every company is finding up how to level up their employees via AI, so that they can use less of them in the future.

So even without increased revenue, AI has its (mis)uses.

nsoonhui··on GPT-6 Astra Solves a WWI German Radio Cipher
My experience seems to echo some parts of yours: Codex is stingy with tokens when compared to Claude Code.

But Codex is superior when comes to diagnosing bugs ( especially when they involve WPF UI threads), writing tests ( yes, even simulating the form cycles and asynchronous operations) and fixing them.

nsoonhui··on Rust is tier-1 language at Microsoft
This is a shocking news to me. Can you elaborate with sources?
nsoonhui··on Detecting and countering misuse of AI: September 2026
> Our investigation revealed that DeepSeek also deployed tactics similar to Moonshot’s. DeepSeek built a CoT extraction pipeline, relying on the same cross-session replay attack described above. DeepSeek also silently relayed exchanges to Claude without informing DeepSeek customers. Like GTG-16002, their customers were likely not made aware that their requests were being funneled to Claude.

If true, would that sort of explain why Chinese Models score high on benchmarks, but not quite as capable when given real tasks?

nsoonhui··on Anthropic's best AI model struggles to attract users as cheaper tools thrive
Not entirely sure how OpenAI is any different. Their quota system seems random to me. I can use up my quota in a single day, and the next day it gets refilled for no apparent reason. But another time, I also used up my quota in a single day, and there was no refresh; I was made to wait six more days.

To me, all of these are just exercises in getting me to pay for more tokens at API rates.

nsoonhui··on A week of using Codex more than Claude
This is strange, but you were using Codex Terra, or Sol?
nsoonhui··on A week of using Codex more than Claude
I used both.

Claude Code seems more generous with its quota, which is why I use it as my main driver.

That said, Codex does seem more capable, terse, and faster. There are some tasks that Claude can't handle but Codex can. One example was a WinForms binding/project deserialization bug. Sorry, the code is a mess, so even I couldn't quite figure out which part was causing which problem.

I initially thought the bug would be difficult to reproduce in a unit-test setting. Claude could only narrow down the problem and tell me where to put a breakpoint. Codex, on the other hand, actually managed to create a reproducible unit test first, and then used that to fix the bug. That impressed me.

The only problem is the quota. Codex burns through it very, very quickly, even when I'm just using Terra 5.6 Medium. That's basically why Claude Code remains my main driver despite Codex seeming more capable.

nsoonhui··on GPT 5.6 Sol 20% price reduction
That's strange. For Deepseek I burnt through USD 5 on a relatively simple task, in one afternoon. That the simple task took a whole afternoon, a lot of baby sitting, the slowness, and so much money relatively, really shook me to the core.
nsoonhui··on GPT 5.6 Sol 20% price reduction
I did try to use Chinese open models, but for my production work they simply couldn't cope at all; both GLM 5.3 and Deepseek v4 went into infinite loop and wasted my tokens until my OpenRouter wallet reached 0; good thing I didn't enable the auto topup. US models, by contrast, breezed past them.

Even for simpler tasks, Chinese models took long time to complete, and I needed to supervise closely. The price , in the end, didn't come cheap, mainly because too much time wasted on thinking.

So maybe one day Chinese models will squeeze out the American ones, but today is not that day.

As far as consumers are concerned, I feel blindly shilling for anyone purely for ideological reasons are quite meaningless, especially when it comes to open/close source and US/China rivalry. I have no obligation to support "open source/weight" or the "underdogs" just because they are so. We only want things that work, and at a cheap price.

nsoonhui··on The Productivity Mirage
Not entirely what's the point of the article is, because with AI I can clear my backlogs way faster than without it.

How is this *not* real productivity gain?

nsoonhui··on DeepSeek pause fundraise after comments on compute gap to US leaked (transcript) [pdf]
Here's something I really don't understand: If as alleged Chinese open weight models are catching up with US anyway, and the performance is near US frontier model level but Chinese can do it with a fraction of cost, and eventually AI model will be commodified, wouldn't that means that the billion or even trillion dollars that US labs spend have only diminishing returns and the lead is only temporary?

So why Deepseek also want to go down that route? Is having the absolute frontier really that important, given that the performance difference is just transient and costly?

nsoonhui··on Be skeptical of OpenAI's rogue hacker agent story
Then exactly what's the point of the Guardian article?

What should be the focus is whether the capabilities are real; whether companies or anyone else benefit from it is quite secondary. Of course they will! Who wouldn't love a nice story that paints their products in the most glowing light ( which is the actual reality)?

nsoonhui··on Rewriting Bun in Rust
But not all models are equally capable, so I don't know your basis of comparison is even valid, let alone the numbers.
nsoonhui··on Kimi K2.7 Code is generally available in GitHub Copilot
I used GitHub Copilot for my VS 2026 development and switched between ChatGPT and Claude. That was before I discovered Claude Code and the Codex app. Copilot was OK for my purposes, and the USD 10 per month fee was enough for my usage.

However, last month they introduced a new pricing model ( I know the old pricing was not sustainable), and my USD 10 was exhausted within days. Because of that, I switched to Claude Code and Codex and have never looked back. Yes, tokens on Claude Code and Codex are subsidized heavily, but let's just enjoy when good things last.

I do feel there is a difference between using Claude via Copilot versus using Claude directly in Claude Code. I'm not sure what Microsoft is doing behind the scenes.

nsoonhui··on Claude Sonnet 5
Sorry, exactly what is the distinction between agent-assist and agent-driven? T

I give AI an image and just it what's wrong, and then it goes on to fix the bug in the codebase for me ( and write the tests), is this agent-assist or agent-driven?

Sometimes I just give the AI my description, and mockup, and it creates a plan and implements the details for me, and I verify visually ( this is the weak spot of AI), is this agent-assist or agent-driven?

nsoonhui··on Claude Sonnet 5
Your benchmark has Gemini 3.5 Flash as the best model, which doesn't compute for me
nsoonhui··on GLM 5.2 beats Claude in our benchmarks
Not sure what to make if your benchmark because GPT 5.5(low) ranks higher than GPT 5.5 (medium) -- #4 vs #9
nsoonhui··on GLM-5.2 is the new leading open weights model on Artificial Analysis
I really have to take your score with a grain of salt because Opus 4.5 does better than Opus 4.6
nsoonhui··on The 29th International Obfuscated C Code Contest (IOCCC) 2025 Winners
But then we all know that LLM has come a long way since one year ago.

Are you sure they still can't do it?

nsoonhui··on The 29th International Obfuscated C Code Contest (IOCCC) 2025 Winners
I'm not sure this kind of competition is still meaningful, given that LLM can easily convert a program clearly written in any programming language to the most obfuscated C code, and can still easily verify it's correctness in an automated way.

Do I miss anything?

nsoonhui··on Lee Kuan Yew's Singapore Story (2023)
In this kind of discussion, you cannot disentangle the fate Singapore from Malaysia. The comparison between the two is interesting.

When Singapore was squirted out from Malaysia in 1965, it had no natural resources, surrounded by hostile Muslim nations ( though not as bad as Israel, but still), and no one to depend on, except themselves.

The Malaysian Ringgit vs Singapore dollars was 1 to 1 back then in 1970s. And now it's 3.1 to 1. This alone is a testament how far Singapore has come.

One important factors separating Singapore and Malaysia is Malaysia's affirmative action (or quota system) that favors the majority, the Malay Muslims, which gives preference to Malay and Islam in all things including tertiary education, GLC opportunities. If you want to get listed in Malaysia stock market you need to have certain quota reserved for the Malays. It was supposed to ensure social justice and diversity, equality and inclusivity for everyone; why should Chinese monopolize all the opportunity to make money and leave Malays poor? This was so unfair.

This affirmative action was started in 1970, after the famous May 1969 racial riot incident. The argument was the riot happened because that the Malays were badly left behind by circumstances; they suffered so much injustice that they had to release it out on others, and the government must do everything to improve their socioeconomic status, lest the same thing happened again. It originally lasted only 30 years but in 2000, the government deemed that the Malays need more help still, and so it's still in effect today.

The affirmative action initiative by Malaysia government would have made any DEI adherents proud for it's thoroughness. Yet when you look at the results you must have wondered whether we did anything wrong. For if it was done right then why, by the affirmative action supporters own admission, the gap didn't close? And why Malaysia lagged so much behind Singapore? And how much minorities were driven away-- and many of them went to Singapore, to contribute to the economy there-- precisely because of affirmative action?

nsoonhui··on SpaceX, Other Mega IPOs Denied Fast Index Entry by S&P
Not to say I have an opinion one way or another, but why do you think that SpaceX odds to have a successful IPO is lower now?
nsoonhui··on Malaysia enforces ban on social media accounts for children younger than 16
As a Malaysian and a parent, and as someone who detests censorship and who is wholely aware of the slippery slope nature of censorship, I actually agree with the ban.

This is because in Malaysia we already have seen enough examples of bad, vague laws have been used to shut up/down the ethnic minorities and dissenters, adding this ban will not change too much of the landscape.

Banning younger children to have a social media account is good. If we can ban kids from driving because their brains aren't fully developed yet, why not just ban social media account for the same reason?

It's actually sickening to see that everyone-- especially children-- glues to phone in public space: playground, restaurants and whatnot. Of course you can say that adults should follow the same ban but adults are more resistant to the opium of social media ( refer to the driving car example above). So I think the double standard is excusable.

The detriment effects of social media towards the young, girls especially, are well documented ( see the Jonathan Hahdt book "the anxious generation"). So I think the ban is valid.

nsoonhui··on I’ve joined Anthropic
I'm not sure what's your point since he is the co-founder of OpenAI
nsoonhui··on Software engineering may no longer be a lifetime career
I think the polarizing response regarding AI depends on which lenses you are looking through. For junior roles, yes, the job is rapidly disappearing. But for senior roles, experience and judgment are more important than ever.

So yes, software engineering may no longer be a lifetime career for a lot of people, much like elite sport is not a viable career for most—but still, some will, and must, make it their career.

nsoonhui··on Google says criminal hackers used AI to find a major software flaw
There was a discussion a few days ago on White House considers vetting AI models prior to release (https://news.ycombinator.com/item?id=48013608).
nsoonhui··on Vibe coding and agentic engineering are getting closer than I'd like
This is my workflow which I find very productive with Agentic AI.

Disclaimer: I'm doing a CAD-like engineering desktop app, and I'm using VS 2026 Copilot, so YMMV.

When I get a Jira ticket, I will first diagnose the problem, and then ask AI to write a test case for it that will reproduce the problem, with guidance on what/how to do the test case (you will be surprised to know how many geometry, seemingly visual problems can be unit tested), and if necessary I provide clues (like which files to read, etc.) for AI to look at, and ask AI to just go and fix the test.

Often AI can do that; AI can make the test pass and make sure that adjacent tests also pass. If in doubt, I will check the output reasoning. I then verify that the fix is done properly via visual inspection (remember, this is a desktop app), and I ask for clarification if needed.

Then at night I'll let my automated test suites run... and oops! Regression found! Who broke it? AI or human? Who cares. I just tell AI that between these times one of the commits must have broken the code — can you please fix it for me? And AI can do that.

This works for small or medium feature implementation, trival bugfixes, or even annoying geometrical problems that require me to dig out the needle in the haystack. So the productivity gain is very real. But I haven't tried it on feature that requires weeks or months for implementation, maybe I should try it next time.

It's hard to describe the feeling. It's just that the AI is working like a very capable (junior?) programmer; both might not have full domain knowledge, but with strong test suites and senior guidance, both can go very far. And of course AI is cheaper and a lot more effective.

nsoonhui··on What I'm Hearing About Cognitive Debt (So Far)
To me, the cognitive debt incurred by Agentic AI described here is not so different from the cognitive debt incurred by code written by someone else. Even when you are the reviewer of your colleague’s code, you can’t just grok everything as if the code were written by you. What more to say if you are not even the reviewer.

And that’s okay! Much like it’s okay to let other people write the code.

What is important is that the code written by Agentic AI is covered by automated tests adequately, and that you verify that the architectural plan is solid. But then this is also what you do with your colleagues’/juniors’ code.

nsoonhui··on Opus 4.7 knows the real Kelsey
I wonder why this is not guardrailed by Opus?

I fed a few pieces of my (anonymous ) writings to ChatGPT and asked it to guess whether it's me. ChatGPT refused, "due to policy to not doxx people".

Page 1 of 10Next →