HNHacker News
TopNewBestAskShowJobs

ezyang

1,148 karma · joined February 23, 2010

submissionscomments
ezyang··on PyTorch: A Reference Language
One of the reasons why I have been thinking and writing this topic is because there is indeed a malaise among the team, which is along the lines of "AI agents are blowing up coding everywhere, what place does a compiler have in this brave new world?" If compilers truly are pointless, then I would like to update on this and save people time and grief!

But I also don't think compilers are pointless. In this post, which looks at the maximalist AI situation, the compiler resurrects itself as the verifier: the two types of systems do very similar things! You can also defend compilers on more banal grounds: not everyone has a ton of kernel engineers to hand optimize your kernels; token efficient ways of getting good baselines is useful; there's more to ML than transformer models and compiler leverage is really useful there. There's also a historical point which is that torch.compile... by all objective metrics, has really been quite successful!

I do agree it's hard to tell the future. These days, I ask myself, "Do I think this is likely to happen in six months." I don't think compilers die in six months. Ask me again in six months :)

ezyang··on Every GitHub object has two IDs
I just want to point out that Opus 4.5 actually knows this trick and will write the code to decode the IDs if it is working with GitHub's API lol
ezyang··on Tracking Copilot vs. Codex vs. Cursor vs. Devin PR Performance
The happy path way of getting code out of Codex is a PR. This is emphatically not true for Cursor.
ezyang··on GPT-4.1 in the API
Lmarena isn't that useful anymore lol
ezyang··on Everything wrong with MCP
I quite liked this article, actually, and I'm quite an MCP stan. These all seem like legitimate problems that the burgeoning ecosystem needs to figure out. Some of them will be logically solved inside the MCP spec, but I also think some won't.
ezyang··on Local CI. Sign off on your own work
I mean, obviously you want to run the local CI in some isolated way. But it's not such a bad idea for many projects.

At PyTorch, we have the reverse problem, it's basically infeasible to run the CI locally, you really do need the cloud setup to cover all of the operating system / hardware configurations, and the parallelization in cloud also saves you loads of time.

ezyang··on Show HN: Codemcp – Claude Code for Claude Pro subscribers – ditch API bills
No, in fact, codemcp can be thought of as a fancier version of the official filesystem MCP that Anthropic released. It's 100% MCP.
ezyang··on Show HN: Codemcp – Claude Code for Claude Pro subscribers – ditch API bills
One thing I'll say, is that if I was going to make people pay API costs (like cline/claude code) I probably wouldn't actually make an MCP. The MCP box is pretty limiting, and I'm only willing to pay the cost because that's how I get onto flat pricing structure.
ezyang··on AI Blindspots – Blindspots in LLMs I've noticed while AI coding
I definitely agree that for current models, the problem is finding where the LLM has comparative advantage. Usually it's something like (1) something boring, (2) something where you don't have any of the low level syntax or domain knowledge, or (3) you are on manager schedule and you need to delegate actual coding.
ezyang··on AI Blindspots – Blindspots in LLMs I've noticed while AI coding
One problem with keeping the changes separate is the LLM usually wants to test the code with the incremental new changes. So you need a working tree that has all the new changes. But then... why not use the real one?
ezyang··on Show HN: Codemcp – Claude Code for Claude Pro subscribers – ditch API bills
So I have liked the Claude Pro style rate limiting over credits that refresh monthly but mostly because I only get snatches of 1-2 hours while I am on baby leave so I never actually get rate limited. As for learnings, I put a lot of them in my AI Blindspots blog, cuz I did most of codemcp's dev with LLMs
ezyang··on AI Blindspots – Blindspots in LLMs I've noticed while AI coding
The thing is that, as many junior engineers can attest, randomly blundering around can still give you something useful! So you definitely can get value out of AI coding with the current generation of models.
ezyang··on AI Blindspots – Blindspots in LLMs I've noticed while AI coding
I have used Cursor and my own MCP codemcp. Cursor has a lot of nice QoL that you can't get from an MCP package; the TAB is really good for traditional coding. Haven't used copilot so I don't have a comparison there. Definitely use agent mode.
ezyang··on Show HN: Codemcp – Claude Code for Claude Pro subscribers – ditch API bills
When I posted this to Reddit there was a pretty lively discussion there: https://www.reddit.com/r/ChatGPTCoding/s/wRmnREUWzn
ezyang··on Show HN: Codemcp – Claude Code for Claude Pro subscribers – ditch API bills
Cline has famously huge API token costs as it is profligate with context. Because codemcp plugs into Claude Desktop you only pay for your Claude Pro sub, similar to Cursor's pricing model
ezyang··on Show HN: Codemcp – Claude Code for Claude Pro subscribers – ditch API bills
In my experience, the model is king, and codemcp will operate fairly similarly to Cursor with the same failure modes as Cursor due to Sonnet 3.7. One thing that I like about codemcp is I can customize aspects of the interaction as I discover new things I want to do :)
ezyang··on AI Blindspots – Blindspots in LLMs I've noticed while AI coding
Yeah, Sonnet 3.5/3.7 are doing heavy lifting. Maybe the SOTA Gemini models would do better, I haven't tried them. Generating correct patches is a funny minigame that isn't really solved, despite how easy it is to RL on.
ezyang··on AI Blindspots – Blindspots in LLMs I've noticed while AI coding
In fact, the test framework I was using at the time (jest) did in fact support this. But the person who had originally written the tests hadn't had the foresight to use snapshot tests for this failing test!
ezyang··on AI Blindspots – Blindspots in LLMs I've noticed while AI coding
They are listed on one page right now! Haha
ezyang··on Show HN: Codemcp – Claude Code for Claude Pro subscribers – ditch API bills
I'm surprised you managed to get Desktop running on Linux lol. You don't need the filesystem/git MCPs alongside codemcp; in fact, it's better not to have them so that Claude consistently uses' codemcp's equivalents to do edits. I'm not sure why codemcp's built-in git support did not work; you can probably find out more by looking in .codemcp/codemcp.log.

If you need to start a new chat, it works just fine. Tell Claude what's happened so far and what you want it to do. You can also ask Claude to summarize the old conversation, that's how /compact in claude code works too.

ezyang··on AI Blindspots – Blindspots in LLMs I've noticed while AI coding
Hi Hacker News! One of the things about this blog that has gotten a bit unwieldy as I've added more entries is that it's a sort of undifferentiated pile of posts. I want some sort of organization system but I haven't found one that's good. Very open to suggestions!
ezyang··on Hacking Your Own AI Coding Assistant with Claude Pro and MCP
And if you don't like clicking, even once, check out my other project, Refined Claude :)
ezyang··on Hacking Your Own AI Coding Assistant with Claude Pro and MCP
Most of the problem is installing the MCP servers, which is more annoying on Windows. https://github.com/ezyang/codemcp#getting-started has instructions that I've personally tested for installation on Windows, which might help you out some.
ezyang··on Hacking Your Own AI Coding Assistant with Claude Pro and MCP
I don't really recommend using filesystem MCP directly, it won't checkpoint changes so it's easy to end up in a state where you can't recover an older working version of the code. Use an actual coding oriented MCP.
ezyang··on Hacking Your Own AI Coding Assistant with Claude Pro and MCP
But only Claude Desktop gets flat $20 pricing from Claude Pro lol
ezyang··on Hacking Your Own AI Coding Assistant with Claude Pro and MCP
Not built in, you'll have to use something like https://github.com/sparfenyuk/mcp-proxy
ezyang··on Hacking Your Own AI Coding Assistant with Claude Pro and MCP
And with Claude Code, it's typically not $0.08... it's more like $0.50, $5.00 just for a roll LOL. Variable rewards gambling addiction? Definitely...
ezyang··on Hacking Your Own AI Coding Assistant with Claude Pro and MCP
If you'd done it with MCP it would only have cost you $20 and you would still have had the rest of the month to use your Claude Pro sub :P
ezyang··on Hacking Your Own AI Coding Assistant with Claude Pro and MCP
You can try (self promo) https://github.com/ezyang/codemcp . https://github.com/rusiaaman/wcgw is also quite popular, although they allow unrestricted shell access (that's why it's named wcgw lol).
ezyang··on Hacking Your Own AI Coding Assistant with Claude Pro and MCP
There's also some fundamental limitations to the Desktop MCP experience that are probably never getting fixed; Claude Code can spin off subagents and play around with the context, I assume that Claude Desktop's form factor is basically going to stay the way it is until the end of time lol.
Page 1 of 8Next →