HNHacker News
TopNewBestAskShowJobs

pcwelder

388 karma · joined November 27, 2019

submissionscomments
pcwelder··on GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
The whole thing (article, benchmark) is a slop soup.

- Metaphor overload. We get it, it's just like car racing. Show some mercy on human readers.

- X, not Y

- A, never B

- Tasteless em dashes

- Hallucinated data, like model add date. The standings are so unbelievable that they border on laughable.

pcwelder··on Prevent cognitive debt by manually retyping LLM-generated code
The most I enjoy working with AI is my special workflow.

I ask it to plan the feature in a separate worktree.

In parallel I start coding without being biased by AI and vice versa.

At some point I read its plan and iterate on it all the while I am in implementation mode. This helps me improve my own vision.

Finally I ask the AI to review my implementation. It flags off bugs and gaps which are usually straightforward for it to fix.

pcwelder··on Kimi K2.7-Code: open-source coding model with better token efficiency
Could be json or non json. Instead of using tools in API, you ask model to share structured output in text. You parse the string to get the JSON. Gives much more control over things you can do.

For example model shares

<tool_call name="getWeather"> <param name="city">London</param> </tool_call>

pcwelder··on Statement on US government directive to suspend access to Fable 5 and Mythos 5
https://news.ycombinator.com/item?id=48496895

Found the time traveller.

pcwelder··on Kimi K2.7-Code: open-source coding model with better token efficiency
Great! Finally follows custom tool call format (k2.6 couldn't). It's a good indicator of instructions following and agentic behaviour.

UIs it's generating is pretty good, not without problems, but certainly better than other models at this price point.

pcwelder··on DeepSeek makes the V4 Pro price discount permanent
None of the deepseek models are multimodal. How are you guys able to use it in daily work without image input?

For example it's just so natural to share screenshots in a chat.

pcwelder··on DeepSeek reasonix, DeepSeek native coding agent with high caching and low cost
Opus 4.7 selects such palette and motifs by default. Might even be first iteration of claude design.
pcwelder··on Gemini 3.5 Flash
Opus is not the correct tier to compare this flash model with.

On my tasks it has not been as good as even Sonnet 4.6 so far.

Instruction following over long context feels worse.

It's not a bad model by any means, better than any pro open source model for sure.

pcwelder··on LLMs corrupt your documents when you delegate
It's worth noting that Claude Code itself doesn't use the `insert` tool. (It also uses custom edit tool not the suite's predefined str_replace)

Also as a person developing agentic code tools since before Claude Code, I'm skeptical if str_replace provides accuracy improvement over just full rewrite.

Back in the day when SOTA models would do lazy coding like `// ... rest of the code ...`, full rewrite wasn't easy. Search/replace was fast, efficient and without the lazy coding. However, it came with slight accuracy drop.

Today that accuracy drop might be minimal/absent, but I'm not sure if it could lead to improvements like preventing doc corruption.

pcwelder··on “Car Wash” test with 53 models
To sonnet 4.6 if you tell it first that "You're being tested for intelligence." It answers correctly 100% of the times.

My hypothesis is that some models err towards assuming human queries are real and consistent and not out there to break them.

This comes in real handy in coding agents because queries are sometimes gibberish till the models actually fetch the code files, then they make sense. Asking clarification immediately breaks agentic flows.

pcwelder··on Improving 15 LLMs at Coding in One Afternoon. Only the Harness Changed
Great work, but concurrency is lost.

With search-replace you could work on separate part of a file independently with the LLM. Not to mention with each edit all lines below are shifted so you now need to provide LLM with the whole content.

Have you tested followup edits on the same files?

pcwelder··on GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
Cool! Please share your work if possible!

I couldn't decide on folding and reducing noise so I'm stuck on that front. I believe there is some elegant solution that I'm missing, hope to see your take.

pcwelder··on GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
All anthropic models. Gemini 2.5 pro and above. Gemini 3 flash is very good too.

GPT models can follow tool format correctly but don't keep on going.

Grok-4+ are decent but with issues in longer chats.

Kimi 2.5 has issues with it reverting to its RL tool format.

pcwelder··on GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
I had added z-ai in allow list explicitly and verified that it's the one being used.
pcwelder··on GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
It's live on openrouter now.

In my personal benchmark it's bad. So far the benchmark has been a really good indicator of instruction following and agentic behaviour in general.

To those who are curious, the benchmark is just the ability of model to follow a custom tool calling format. I ask it to using coding tasks using chat.md [1] + mcps. And so far it's just not able to follow it at all.

[1] https://github.com/rusiaaman/chat.md

pcwelder··on Parse, Don't Validate (2019)
Each repost is worth it.

This, along with John Ousterhout's talk [1] on deep interfaces was transformational for me. And this is coming from a guy who codes in python, so lots of transferable learnings.

[1] https://www.youtube.com/watch?v=bmSAYlu0NcY

pcwelder··on MaliciousCorgi: AI Extensions send your code to China
> These are sending all files it can access

TBF, Cursor's code indexing works the same way, it has to send all workspace files to their servers.

Auto-completion systems need previous edits to suggest next edits so no surprises their either.

pcwelder··on Unrolling the Codex agent loop
Sonnet has the same behavior: drops thinking on user message. Curiously in the latest Opus they have removed this behavior and all thinking tokens are preserved.
pcwelder··on Announcing the Beta release of ty
```

from anthropic.types import MessageParam

data: list[MessageParam] = [{"role": "user", "content": [{"type": "text", "text": ""}]}]

```

This for example works both in mypy and pyright. (Also autocompletion of typedict keys / literals from pylance is missing)

pcwelder··on Announcing the Beta release of ty
Displaying inferred types inline is a killer feature (inspired from rust lang server?). It was a pleasant surprise!

It's fast too as promised.

However, it doesn't work well with TypedDicts and that's a show-stopper for us. Hoping to see that support soon.

pcwelder··on Claude CLI deleted my home directory and wiped my Mac
To those who are not deterred and feel yolo mode is worth the risk, there are two patterns that should perk your ears up.

- Cleanup or deletion tasks. Be ready to hit ctrl c anytime. Led to disastrous nukes in two reddit threads.

- Errors impacting the whole repo, especially those that are difficult to solve. In such cases if it decides to reset and redo, it may remove sensitive paths as well.

It removed my repo once because "it had multiple problems and was better to it write from scratch".

- Any weird behavior, "this doesn't seem right", "looks like shell isn't working correctly" indicative of application bug. It might employ dangerous workarounds.

pcwelder··on I failed to recreate the 1996 Space Jam website with Claude
It just fetched the HTML and replicated it. The usage of table is a giveaway.

Any LLM with browser tool can do it (Kombai one shots it too for example), because it's just cheating.

pcwelder··on I failed to recreate the 1996 Space Jam website with Claude
But that's cheating because it then has the source code containing the table and its styles.

I can confirm that this is what it does.

And if you ask it to not use tables, it cleverly uses div with the same layout as the table instead.

pcwelder··on What I don’t like about chains of thoughts (2023)
In RNNs and Transformers we obtain probability distribution of target variable directly and sample using methods like top-k or temprature sampling.

I don't see the equivalence to MCMC. It's not like we have a complex probability function that we are trying to sample from using a chain.

It's just logistic regression at each step.

pcwelder··on Should LLMs just treat text content as an image?
I ϲаn guаrаntее thаt thе ОСR ϲаn't rеаd thіs sеntеnсе ϲоrrесtlу.
pcwelder··on Karpathy on DeepSeek-OCR paper: Are pixels better inputs to LLMs than text?
There are many unicode characters that look alike. There are also those zero width characters.
pcwelder··on Python developers are embracing type hints
It doesn't throw error in the REPL though. Surely you meant to share some other example?
pcwelder··on The bloat of edge-case first libraries
>if they don't whatever happens, happens

What happens is you get an error. So you immediately know something is wrong.

Javascript goes the extra mile to avoid throwing errors.

So you've 3>"2" succeeding in Javascript but it's an exception in python. This behavior leads to hard to catch bugs in the former.

Standard operators and methods have runtime type checks in python and that's what examples in the article are replicating.

pcwelder··on Updates to Consumer Terms and Privacy Policy
Navigate to `https://claude.ai/settings/data-privacy-controls` and disable it before Sept 28. Isn't applicable to team plan.
pcwelder··on How to build a coding agent
Agree. To reduce costs:

1. Precompute frequently used knowledge and surface early. For example repository structure, os information, system time.

2. Anticipate next tool calls. If a match is not found while editing, instead of simply failing, return closest matching snippet. If read file tool gets a directory, return directory contents.

3. Parallel tool calls. Claude needs either a batch tool or special scaffolding to promote parallel tool calls. Single tool call per turn is very expensive.

Are there any other such general ideas?

Page 1 of 5Next →