HNHacker News
TopNewBestAskShowJobs

wcallahan

64 karma · joined April 8, 2022

Founder and CEO of aVenture (aventure.vc), a venture capital research platform. San Francisco-based - let's grab a coffee if you're local!

Personal website: williamcallahan.com

submissionscomments
wcallahan··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
From my experience, I suspect you'll need to be able to use both thinking modes (medium and xhigh).

I did set the default at medium, but the xhigh still seems to be a big part of its benefits.

I found it really easy to chat with it and instantly see flaws in its thinking when using the medium level of reasoning... in a way that had me consider that the frontier models reasoning abstractions (both at the harness level, e.g., claude code, codex) AND in the server-side obfuscation) that I found refreshing, because it made it easy for me to step in and precisely identify the failures of reasoning that the frontier models were getting stuck on, and because of the 'black box' hiddenness of their reasoning, it made it harder for me to diagnose.

So, I suppose I'm saying 'there is a time and place for each'.

Here is my setup: https://williamcallahan.com/blog/qwen-3-8-27b-is-a-great-ope...

wcallahan··on Show HN: SeaTicket – AI agent that resolve GitHub and Discord issues
Hey @Daniel-Pan, I’m a big fan of your work with Seafile. It’s by far the best open source alternative to Google Drive and Google Docs!

I love the idea of this as well.

Do you plan to make it available as open source software also? I think it goes a long way towards trust in the platform/software.

wcallahan··on Building a SaaS in 2026 Using Only EU Infrastructure
I just looked at Scaleway’s pricing for two popular open source models (gpt-oss-120b and qwen3.5-397b) and it’s meaningfully more expensive than alternatives (e.g., many you’d find on OpenRouter).
wcallahan··on Qwen3.6-Plus: Towards real world agents
Openclaw
wcallahan··on Java is fast, code might not be
Isn’t that what Convex is doing?
wcallahan··on Java is fast, code might not be
Jooq with Kotlin for a back-end has been the best of both worlds for me.

Much cleaner, shorter code and type safety with Postgres (my schema tends to be highly normalized too). And these days I’ve got it well integrated with Zod for type safe JS/TS front-ends as well.

wcallahan··on GrapheneOS – Break Free from Google and Apple
American here who values individual liberties greatly. I know things are politically tense at the moment, but I’m not sure I understand this popular contemporary sentiment.

I’ve always believed governments and companies should be regarded with fairly low trust, and the behavior of big tech companies and some recent government actions are great examples why.

But what disappoints me a bit about this moment is (the perhaps inevitable?) response to nationalism with more nationalism.

Just as I didn’t seek to punish the EU over authoritarianism in Hungary and Poland, I feel the current moment has many responding to the symptoms instead of the sources of the problems. This is not a defense of policies I believe concern you, it’s a question of priorities.

I think the author of the article got it right. Because in addition to privacy, I believe one should be able to navigate the internet freely without a mandate to do business with monopolistic dominant companies, which includes rights like ownership of your data.

wcallahan··on Warcraft III Peon Voice Notifications for Claude Code
42 here, played a ton of Warcraft II, but my favorite to return to now is definitely Warcraft III (or AoE II).
wcallahan··on What has Docker become?
I suspect the timing of this and comments is not coincidental.

I pay for Docker licenses, even though not meeting the criteria for business size requiring it, as I wanted reliable image fetching for my self hosted container CI/CD pipelines failing docker hub image fetches.

But as of now, my oAuth logins to Docker expire within hours now, and I’ve been left with no choice but to scatter in search of diffuse container image alternative sources for my Dockerfiles to stop this madness.

My one way permanent migration from Docker Hub sourced images has finally left me with no reason to keep paying for Docker licenses due to whatever this misguided or blundered rate limit implementation is.

wcallahan··on Why NUKEMAP isn't on Google Maps anymore (2019)
I do a lot of maps API calls, and found I get better results (and can save money) by using multiple providers.

So I use Apple Maps, Mapbox, OpenStreetMap, and Google Maps… sometimes to check the results with multiple providers, and sometimes to divvy up the free allotment.

For anyone using Java/Kotlin/JVM, I made an SDK for Apple Maps: https://github.com/WilliamAGH/apple-maps-java which is one with a generous included tier.

wcallahan··on 2026: The Year of Java in the Terminal?
Hope you enjoy tui4j and brief!
wcallahan··on 2026: The Year of Java in the Terminal?
I made some tweaks to the Github releases config, you should be able to do this now as well:

curl -L -o brief.zip https://github.com/WilliamAGH/brief/releases/latest/download... unzip brief.zip cd brief-*/ ./bin/brief

wcallahan··on 2026: The Year of Java in the Terminal?
Great ideas for both :)

I wasn’t expecting the main topic of what I’ve been building to appear on the cover of hacker news today, so I was caught a bit unprepared, but they were definitely on the todo list next!

wcallahan··on 2026: The Year of Java in the Terminal?
It just so happens that I’ve built one already: TUI4J (Terminal User Interface for Java).

https://github.com/WilliamAGH/tui4j

It combines a port of BubbleTea from Go, and Textual and other inspired rewrites of other functionality.

It’s a fork of someone’s earlier work that I sought to expand/stabilize.

I built a beautifully simple LLM chat interface with full dialog windows, animations, and full support for keyboard and mouse interactivity parity, showing what this Java library is capable of.

Example chat app: https://github.com/WilliamAGH/brief

Would love to see others build similar things with it!

wcallahan··on Nvidia Nemotron 3 Family of Models
Great advice. Have you observed any other differences? I’ve been wondering if there are any specialized variants yet of GPT-OSS models yet that outperform on specific tasks (similar to the countless Llama 3 variants we’ve seen).
wcallahan··on Nvidia Nemotron 3 Family of Models
Yes to both comments. I said that to:

1. disclose my method was not quantifiably measurable as the not model, because that is not important to me, speed of action/development outcomes is more important to me, and because

2. I’ve observed a large gap between benchmark toppers and my own results

But make no mistake, I like have the terminals scrolling live across multiple monitors so I can glance at them periodically and watch their response quality, so I care and notice which give better/worse results.

My biggest goal right now after accuracy is achieving more natural human-like English for technical writing.

wcallahan··on Nvidia Nemotron 3 Family of Models
Yes, I run it locally on 3 different AMD Strix Halo machines (Framework Desktop and 2 GMKTec machines, 128gb x 2, 96gb x 1) and a Mac Studio M2 Ultra 128gb of unified memory.

I’ve used several runtimes, including vLLM. Works great! Speedy. Best results with Ubuntu after trying a few different distributions and Vulkan and ROCm drivers.

wcallahan··on Nvidia Nemotron 3 Family of Models
I don’t do ‘evals’, but I do process billions of tokens every month, and I’ve found these small Nvidia models to be the best by far for their size currently.

As someone else mentioned, the GPT-OSS models are also quite good (though I haven’t found how to make them great yet, though I think they might age well like the Llama 3 models did and get better with time!).

But for a defined task, I’ve found task compliance, understanding, and tool call success rates to be some of the highest on these Nvidia models.

For example, I have a continuous job that evaluates if the data for a startup company on aVenture.vc could have overlapping/conflated two similar but unrelated companies for news articles, research details, investment rounds, etc… which is a token hungry ETL task! And I recently retested this workflow on the top 15 or so models today with <125b parameters, and the Nvidia models were among the best performing for this type of work, particularly around non-hallucination if given adequate grounding.

Also, re: cost - I run local inference on several machines that run continuously, in addition to routing through OpenRouter and the frontier providers, and was pleasantly surprised to find that if I’m a paying customer of OpenRouter otherwise, the free variant there from Nvidia is quite generous for limits, too.

wcallahan··on Show HN: Real-time system that tracks how news spreads across 200k websites
Instead of filtering them out, I’d imagine you’d want to establish their equivalency instead? Then they can be made available as equal/similar alternatives to the same article (i.e., from your outlet of choice).
wcallahan··on Show HN: Real-time system that tracks how news spreads across 200k websites
Bad bot.

‘masterphai’ is evidence of how effective a good LLM and better prompt can be now at evading detection of AI authorship… but there’s no way this authors comments are written by a sane human.

From the comment history it appears it has tricked quite a few humans to-date. Interesting!

wcallahan··on You can see a working Quantum Computer in IBM's London office
I suspect I’m not alone in pausing around the statement:

> "It’s not likely to be something you’ll ever have at home"

I’m curious… what would need to be true to make this statement wrong?

wcallahan··on DBCrust – A modern database CLI
It would be great to have Convex Database support
wcallahan··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
I just used GPT-OSS-120B on a cross Atlantic flight on my MacBook Pro (M4, 128GB RAM).

A few things I noticed: - it’s only fast with with small context windows and small total token context; once more than ~10k tokens you’re basically queueing everything for a long time - MCPs/web search/url fetch have already become a very important part of interacting with LLMs; when they’re not available the LLM utility is greatly diminished - a lot of CLI/TUI coding tools (e.g., opencode) were not working reliably offline at this time with the model, despite being setup prior to being offline

That’s in addition to the other quirks others have noted with the OSS models.

wcallahan··on Gemini 2.5 Flash
Tried Gemini Codes yesterday, as well as anon-kode and anon-codex. Gemini Codes is already broken and appears to be rather brittle (she disclosures as much), and the other two appear to still need some prompt improvements or someone adding vector embedding for them to be useful?

Perhaps someone can merge the best of Aider and codex/claude code now. Looking forward to it.

wcallahan··on Show HN: Windsurf – Agentic IDE
I wanted to watch the video, but the keyboard typing being the loudest part of the video made it rather hard to listen.

I wonder if a tool exists to strip keyboard noise from YouTube videos?

wcallahan··on New Mac Mini with M4
Just ordered one the UM690S after this comment - I had been looking for another machine like it and saw the sale. Thanks for sharing!

I just deployed another System76 machine that I’m using for the same purpose — using both for redundant containerized web servers I’m self hosting. I got it with 64GB RAM.

Has the machine been a solid performer?

wcallahan··on Show HN: Dead man's switch without reliance on your infra
Was thinking the same!
wcallahan··on A web scraping CLI made for AI that is idempotent
Nice work. Would love a similar repository for Google cloud’s equivalent services!

Or a PR on this that accomplishes the same, as @clemlesne mentioned.

wcallahan··on OpenDevin: An Open Platform for AI Software Developers as Generalist Agents
Aider is written in Python (they have a great Discord community, btw). My experience matches yours: for Python, aider/Sonnet seems to do much better than for Javascript so far. I strongly recommend aider despite LLM limitations at the moment for anyone interested in this space.

It's also very sensitive, unsurprisingly, to development documentation that is moving quickly, e.g., most AI APIs right now. A lot of manual intervention is still required here because of out-of-date references to imports, etc.

wcallahan··on OpenDevin: An Open Platform for AI Software Developers as Generalist Agents
I'm an active aider user, I spent ~$120 last month on a combo of Sonnet and Opus. It was much more expensive, as you probably know, with Opus. Now it's rather reasonably priced and more sustainable, IMO.
Page 1 of 2Next →