So maybe one day Chinese models will squeeze out the American ones, but today is not that day.
So no, I am not excited about Chinese models ( just because its open weight and not American).
15,398 karma · joined December 18, 2007
So maybe one day Chinese models will squeeze out the American ones, but today is not that day.
So no, I am not excited about Chinese models ( just because its open weight and not American).
I bet every company is finding up how to level up their employees via AI, so that they can use less of them in the future.
So even without increased revenue, AI has its (mis)uses.
But Codex is superior when comes to diagnosing bugs ( especially when they involve WPF UI threads), writing tests ( yes, even simulating the form cycles and asynchronous operations) and fixing them.
If true, would that sort of explain why Chinese Models score high on benchmarks, but not quite as capable when given real tasks?
To me, all of these are just exercises in getting me to pay for more tokens at API rates.
Claude Code seems more generous with its quota, which is why I use it as my main driver.
That said, Codex does seem more capable, terse, and faster. There are some tasks that Claude can't handle but Codex can. One example was a WinForms binding/project deserialization bug. Sorry, the code is a mess, so even I couldn't quite figure out which part was causing which problem.
I initially thought the bug would be difficult to reproduce in a unit-test setting. Claude could only narrow down the problem and tell me where to put a breakpoint. Codex, on the other hand, actually managed to create a reproducible unit test first, and then used that to fix the bug. That impressed me.
The only problem is the quota. Codex burns through it very, very quickly, even when I'm just using Terra 5.6 Medium. That's basically why Claude Code remains my main driver despite Codex seeming more capable.
Even for simpler tasks, Chinese models took long time to complete, and I needed to supervise closely. The price , in the end, didn't come cheap, mainly because too much time wasted on thinking.
So maybe one day Chinese models will squeeze out the American ones, but today is not that day.
As far as consumers are concerned, I feel blindly shilling for anyone purely for ideological reasons are quite meaningless, especially when it comes to open/close source and US/China rivalry. I have no obligation to support "open source/weight" or the "underdogs" just because they are so. We only want things that work, and at a cheap price.
How is this *not* real productivity gain?
So why Deepseek also want to go down that route? Is having the absolute frontier really that important, given that the performance difference is just transient and costly?
What should be the focus is whether the capabilities are real; whether companies or anyone else benefit from it is quite secondary. Of course they will! Who wouldn't love a nice story that paints their products in the most glowing light ( which is the actual reality)?
However, last month they introduced a new pricing model ( I know the old pricing was not sustainable), and my USD 10 was exhausted within days. Because of that, I switched to Claude Code and Codex and have never looked back. Yes, tokens on Claude Code and Codex are subsidized heavily, but let's just enjoy when good things last.
I do feel there is a difference between using Claude via Copilot versus using Claude directly in Claude Code. I'm not sure what Microsoft is doing behind the scenes.
I give AI an image and just it what's wrong, and then it goes on to fix the bug in the codebase for me ( and write the tests), is this agent-assist or agent-driven?
Sometimes I just give the AI my description, and mockup, and it creates a plan and implements the details for me, and I verify visually ( this is the weak spot of AI), is this agent-assist or agent-driven?
Are you sure they still can't do it?
Do I miss anything?
When Singapore was squirted out from Malaysia in 1965, it had no natural resources, surrounded by hostile Muslim nations ( though not as bad as Israel, but still), and no one to depend on, except themselves.
The Malaysian Ringgit vs Singapore dollars was 1 to 1 back then in 1970s. And now it's 3.1 to 1. This alone is a testament how far Singapore has come.
One important factors separating Singapore and Malaysia is Malaysia's affirmative action (or quota system) that favors the majority, the Malay Muslims, which gives preference to Malay and Islam in all things including tertiary education, GLC opportunities. If you want to get listed in Malaysia stock market you need to have certain quota reserved for the Malays. It was supposed to ensure social justice and diversity, equality and inclusivity for everyone; why should Chinese monopolize all the opportunity to make money and leave Malays poor? This was so unfair.
This affirmative action was started in 1970, after the famous May 1969 racial riot incident. The argument was the riot happened because that the Malays were badly left behind by circumstances; they suffered so much injustice that they had to release it out on others, and the government must do everything to improve their socioeconomic status, lest the same thing happened again. It originally lasted only 30 years but in 2000, the government deemed that the Malays need more help still, and so it's still in effect today.
The affirmative action initiative by Malaysia government would have made any DEI adherents proud for it's thoroughness. Yet when you look at the results you must have wondered whether we did anything wrong. For if it was done right then why, by the affirmative action supporters own admission, the gap didn't close? And why Malaysia lagged so much behind Singapore? And how much minorities were driven away-- and many of them went to Singapore, to contribute to the economy there-- precisely because of affirmative action?
This is because in Malaysia we already have seen enough examples of bad, vague laws have been used to shut up/down the ethnic minorities and dissenters, adding this ban will not change too much of the landscape.
Banning younger children to have a social media account is good. If we can ban kids from driving because their brains aren't fully developed yet, why not just ban social media account for the same reason?
It's actually sickening to see that everyone-- especially children-- glues to phone in public space: playground, restaurants and whatnot. Of course you can say that adults should follow the same ban but adults are more resistant to the opium of social media ( refer to the driving car example above). So I think the double standard is excusable.
The detriment effects of social media towards the young, girls especially, are well documented ( see the Jonathan Hahdt book "the anxious generation"). So I think the ban is valid.
So yes, software engineering may no longer be a lifetime career for a lot of people, much like elite sport is not a viable career for most—but still, some will, and must, make it their career.
Disclaimer: I'm doing a CAD-like engineering desktop app, and I'm using VS 2026 Copilot, so YMMV.
When I get a Jira ticket, I will first diagnose the problem, and then ask AI to write a test case for it that will reproduce the problem, with guidance on what/how to do the test case (you will be surprised to know how many geometry, seemingly visual problems can be unit tested), and if necessary I provide clues (like which files to read, etc.) for AI to look at, and ask AI to just go and fix the test.
Often AI can do that; AI can make the test pass and make sure that adjacent tests also pass. If in doubt, I will check the output reasoning. I then verify that the fix is done properly via visual inspection (remember, this is a desktop app), and I ask for clarification if needed.
Then at night I'll let my automated test suites run... and oops! Regression found! Who broke it? AI or human? Who cares. I just tell AI that between these times one of the commits must have broken the code — can you please fix it for me? And AI can do that.
This works for small or medium feature implementation, trival bugfixes, or even annoying geometrical problems that require me to dig out the needle in the haystack. So the productivity gain is very real. But I haven't tried it on feature that requires weeks or months for implementation, maybe I should try it next time.
It's hard to describe the feeling. It's just that the AI is working like a very capable (junior?) programmer; both might not have full domain knowledge, but with strong test suites and senior guidance, both can go very far. And of course AI is cheaper and a lot more effective.
And that’s okay! Much like it’s okay to let other people write the code.
What is important is that the code written by Agentic AI is covered by automated tests adequately, and that you verify that the architectural plan is solid. But then this is also what you do with your colleagues’/juniors’ code.
I fed a few pieces of my (anonymous ) writings to ChatGPT and asked it to guess whether it's me. ChatGPT refused, "due to policy to not doxx people".