HNHacker News
TopNewBestAskShowJobs

joshmlewis

1,402 karma · joined March 7, 2011

I build software and run long distances in the wilderness.

PS Would love to discuss cool projects with AI. Email is hn [at] josh.ml.

submissionscomments
joshmlewis··on Gemini 3.5 Flash
It's interesting they use output tokens as an eval because all tokens are not made equal. Even from model to model (like Opus 4.6 to Opus 4.7) the tokenizer can be different and it's no longer an apples to apples comparison. No one really talks about this but it directly affects stats like usage limits. Certainly comparing models between providers on an apples to apples comparison token wise is not a good test.
joshmlewis··on Audio is the one area small labs are winning
Speechify has been good for me although there might be better / cheaper alternatives I'm not aware of.
joshmlewis··on Claude Opus 4.6
I think the OP was implying that it's probably already baked into its training data. No need to search the web for that.
joshmlewis··on OpenClaw – Moltbot Renamed Again
"They" being the guy (Peter Steinberger) who created it as a personal project that he open sourced.
joshmlewis··on How to code Claude Code in 200 lines of code
This is cool but as someone that's built an enterprise grade agentic loop in-house that's processing a billion plus tokens a month, there are so many little things you have to account for that greatly magnify complexity in real world agentic use cases. For loops are an easy way to get your foot in the door and is indeed at the heart of it all, but there are a multitude of a little things that compound complexity rather quickly. What happens when a user sends a message after the first one and the agent has already started the tool loop? Seems simple, right? If you are receiving inputs via webhooks (like from a Slack bot), then what do you do? It's not rocket science but it's also not trivial to do right. What about hooks (guardrails) and approvals? Should you halt execution mid-loop and wait or implement it as an async Task feature like Claude Code and the MCP spec? If you do it async then how do you wake the agent back up? Where is the original tool call stored and how is the output stored for retrieval/insertion? This and many other little things add up and compound on each other.

I should start a blog with my experience from all of this.

joshmlewis··on Claude Memory
This just feels like the whole complicated TODO workflows and MCP servers that were the hot thing for awhile. I really don't believe this level of abstraction and detailed workflows are where things are headed.
joshmlewis··on Claude Memory
This should not really be necessary and is more of a workaround for bad patterns / prompting in my opinion.
joshmlewis··on Claude Memory
How big is your claude.md file? I see people complain about this but I have only seen it happen in projects with very long/complex or insufficient claude.md files. I put a lot of time into crafting that file by hand for each project because it's not something it will generate well on its own with /init.
joshmlewis··on Slack has raised our charges by $195k per year
It's also not a coincidence that Slack is neutering the ability to access channel history via the API very soon. With a very generous rate limit of 2 requests per minute I believe it was and a max of ~10 messages. This is already enforced for new marketplace apps and will apply to all apps starting in March according to their docs.
joshmlewis··on We put a coding agent in a while loop
One of the biggest nuggets people need to take away from this:

> At one point we tried “improving” the prompt with Claude’s help. It ballooned to 1,500 words. The agent immediately got slower and dumber. We went back to 103 words and it was back on track.

Keep your prompts / agent instructions short. Focus on the wide view, not specifics.

joshmlewis··on Tidewave Web: in-browser coding agent for Rails and Phoenix
As someone who builds AI products and having used agentic coding tools since they came out (often with Rails projects), I don't get this. There was a similar project called Rails MCP Server which said:

> "This Rails MCP Server implements the MCP specification to give AI models access to Rails projects for code analysis, exploration, and assistance."

And again I don't get the value. I can see some slight benefits to having a tight browser integration but I don't think that's worth leaving the IDE / CLI tools and the workflows they bring. You can also use Playwright MCP or just screenshot easily for more context. Claude Code can now run your server in the background and view logs as well. In a perfect world where LLM's can one shot whole features, maybe. But I can't let Claude Code go 10 minutes without it introducing a bad pattern and having to stop it. Reducing that visibility even further with this does not seem like a good combo.

I'm not wanting to tear down others projects either, just giving my perspective. I should try it to see how it does in the wild but the Copilot license or Anthropic API key requirement also deters me as well as having to have a project specific dependency.

joshmlewis··on GPT-5
It is funny how it can be like this sometimes. I think a lot depends on coding styles, languages, prompting, etc.
joshmlewis··on GPT-5 for Developers
Cursor
joshmlewis··on GPT-5
When it came out on Tuesday I wanted to throw my laptop out of the window. I don't know what happened but results were total garbage earlier this week. It got better the past couple days but so far with gpt-5 being able to solve problems without as much correction I'm going to use it more.
joshmlewis··on GPT-5
Whoosh, it went right over my head.
joshmlewis··on GPT-5
The data is made up, the point is to see how models respond to the same input / scenario. You're able to create whatever tools you want and import real data or it'll generate fake tool responses for you based on the prompt and tool definition.

Disclaimer: I made PromptSlice for creating and comparing prompts, tools, and models.

joshmlewis··on Cursor CLI
I would highly doubt it. Even when you BYOK inside of Cursor they still say it's routed through their servers.
joshmlewis··on Cursor CLI
I noticed it was taking awhile on the first large-ish task I gave it. I'm assuming it was just a bit overloaded at the moment.
joshmlewis··on GPT-5
Where'd you get 720 from?
joshmlewis··on GPT-5
Did I say GPT-5? I said o3. :) That was a rebuttal to you saying you have never needed to add your key to use an OpenAI model before.
joshmlewis··on GPT-5: Key characteristics, pricing and system card
It seems to be trained to use tools effectively to gather context. In this example against 4.1 and o3 it used 6 in the first turn in a pretty cool way (fetching different categories that could be relevant). Token use increases with that kind of tool calling but the aggressive pricing should make that moot. You could probably get it to not be so tool happy with prompting as well.

https://promptslice.com/share/b-2ap_rfjeJgIQsG

joshmlewis··on GPT-5 for Developers
It's free in Cursor for the next few days, you should go try it out if you haven't. I've been an agentic coding power user since the day it came out across several IDE's/CLI tools and Cursor + GPT-5 seems to be a great combo.
joshmlewis··on GPT-5 for Developers
It does seem to be doing well compared to Opus 4.1 in my testing the last few hours. I've been on the Claude Code 200 plan for a few months and I've been really frustrated with it's output as of late. GPT-5 seems to be a step forward so far.
joshmlewis··on GPT-5 for Developers
It does really well at using tool calls to gain as much context as it can to provide thoughtful answers. In this example it did 6! tool calls in the first response while 4.1 did 3 and o3 did one at a time.

https://promptslice.com/share/b-2ap_rfjeJgIQsG

joshmlewis··on GPT-5 for Developers
I've been testing it against Opus 4.1 the last few hours and it has done better and solved problems Claude kept failing at. I would say it's definitely better, at least so far.
joshmlewis··on GPT-5
You had to use your own key for o3 at least.

> Note that BYOK is required for this model. Set up here: https://openrouter.ai/settings/integrations

https://openrouter.ai/api/v1/models

joshmlewis··on GPT-5
It's more efficient with tools for one and the input cost is cheaper (which is where a lot of the cost is).

See comparison between GPT-5, 4.1, and o3 tool calling here: https://promptslice.com/share/b-2ap_rfjeJgIQsG.

joshmlewis··on GPT-5
I am convinced. I've been giving it tasks the past couple hours that Opus 4.1 was failing on and it not only did them but cleaned up the mess Opus made. It's the real deal.
joshmlewis··on GPT-5
It's a really good model from my testing so far. You can see the difference in how it tries to use tools to the greatest extent when answering a question, especially compared to 4.1 and o3. In this example it used 6! tool calls in the first response to try and collect as much info as possible.

https://promptslice.com/share/b-2ap_rfjeJgIQsG

joshmlewis··on Denver rent is back to 2022 prices after 20k new units hit the market
As a side note, if anyone here is in Denver and wants to grab a coffee or cowork sometime, hit me up.
Page 1 of 21Next →