HNHacker News
TopNewBestAskShowJobs

asabla

257 karma · joined February 18, 2021

submissionscomments
asabla··on Prompting Claude Opus 5.5
> All that keeps jumping out at me is how they've set it to refuse giving users thinking tokens and prompts for full reasoning in output

I keep seeing comments added to code, which reads like reasoning output instead of meaningful words. I see this behavior for both OpenAI and Anthropic models (for several harnesses as well).

But this is a sample of one. And I may be in a situation where I'm more negative to the output from LLMs in general.

asabla··on Reverse-Engineering Claude Web's MicroVM: Uncovering Anthropic's Hidden Antspace
The way I interpret this is:

After some inactivity it goes into some dormant mode. Where it will only consume some storage. The next time someone tries to access this artifact it just restores a snapshot from last working state.

I guess the public url part is working similar to how OpenAI and their Sites are. E.g: their personally only available to you, until you change the sharing options.

asabla··on Anthropic boss Dario Amodei calls for AI development to slow down
> My new suspicion is now that they didn’t got drastically smarter, but they got trained on the user input on the previous generations

I think this happened around Opus 4.5 or 4.5 and the same for GPT 5.4.

The more I use those models, harnesses, techniques for guidance etc etc. The more I land in going back to writing software by hand again. Maybe not all of it, but at least the crucial parts + foundations.

asabla··on Fine, I'll build my own text editor
I used sublime text over so many years. Still find it to be a marvel when it comes to large text files, and how well it still handles them
asabla··on 44% on ARC-AGI-1 in 67 cents
Thank you for answering these questions. Looking forward for the next write up about this.
asabla··on Keychron announces first open-source firmware for gaming mice
I had a similar issue with mine. But was able to reset it with Fn + z + j for a couple of seconds.
asabla··on Command Line Interface Guidelines
It took me way too many years before I set this as a default in my profile files.

The funny thing is. I kept missing this detail every time I read through git. Feels like I was blind to the whole concept of it

asabla··on Project Hail Mary – Stellar Navigation Chart
Super cool! how long did it take to generate all those custom images?
asabla··on Agent Safehouse – macOS-native sandboxing for local agents
Yee I gotcha.

Did a migration myself last week from using playwright mcp towards playwright-cli instead. Which has been playing much nicer so far. I guess you would run into the same issues you've already mentioned about running chrome headless in one of these sandboxes.

I'll for sure keep an eye out for updates.

Kudos to the project!

asabla··on Agent Safehouse – macOS-native sandboxing for local agents
Oh woah!

I've been trying to get microsandbox to play nicely. But this is much closer to what I actually need.

I glimpsed through the site and the script. But couldn't really see any obvious gotchas.

Any you've found so far which hasn't been documented yet?

asabla··on GPT-5.4
I really don't have any numbers to back this up. But it feels like the sweet spot is around ~500k context size. Anything larger then that, you usually have scoping issues, trying to do too much at the same time, or having having issues with the quality of what's in the context at all.

For me, I would say speed (not just time to first token, but a complete generation) is more important then going for a larger context size.

asabla··on Running Claude Code dangerously (safely)
Very nice!

I've been experimenting with a similar setup. And I'll probably implement some of the things you've been doing.

For the proxy part I've been running https://www.mitmproxy.org/ It's not fully working for all workflows yet. But it's getting close

asabla··on A guide to local coding models
From my point of view, you're either choosing between instruction following or more creative solutions.

Codex models tend to be extremely good at following instructions, to the point that it won't do any additional work unless you ask it to. GPT-5.1 and GPT-5.2 on the other hand is a little bit more creative.

Models from Anthropics on the other hand is a lot more loosy goosy on the instructions, and you need to keep an eye on it much more often.

I'm using models interchangeably from both providers all the time depending on the task at hand. No real preference if one is better then the other, they're just specialized on different things

asabla··on Fifty Shades of OOP
This is such a good video. I really like the way he presents it as well.

His rant about CS historians is also a fun subject

asabla··on AI can code, but it can't build software
> The last thing I'll mention is that Claude Code (Sonnet 4.5) is still very token-happy, in that it eagerly goes above and beyond when not always necessary. Codex (gpt-5-codex) on the other hand, does exactly what you ask, almost to a fault.

I very much share your experience. As for the time being I like the experience with codex over claude, just because I find my self in a position where I know much sooner when to step in and just doing it manually.

With claude I find my self in a typing exercise much more often, I could probably get better of knowing when to stop ofc.

asabla··on A proposal to add GC-less, unmanaged memory spaces to C#
I can't tell if this is satire or not. And some parts read like it was written by AI.

Either way, a more fine grained control over the GC is probably preferred over something like this.

asabla··on GPT-OSS Reinforcement Learning
I'm always so confused by those statements as well. Because just like you, I feel that the 20B version is really good at following instructions.

Some of the qwen models are too, but they seem to need a bit more handholding.

This is of course just anecdotal from my end. And I've been slacking on keeping up with evals while testing at home

asabla··on Are OpenAI and Anthropic losing money on inference?
And by GPT-5 you mean through their API? Directly through Azure OpenAI services? or are you talking about ChatGPT set to using GPT-5.

All of these alternatives means different things when you say it takes +20 seconds for a full response.

asabla··on The issue of anti-cheat on Linux (2024)
I fundamentally agree with you.

But anti-cheat hasn't been about blocking every possible way of cheating for some time now. It's been about making it as in convenient as possible, thus reducing the amount of cheaters.

Is the current fad of using kernel level anti-cheats what we want? hell nah.

The responsibility of keeping a multi-player session clean of cheaters, was previously shared between the developers and server owners. While today this responsibility has fallen mostly on developers (or rather game studios) since they want to own the whole experience.

asabla··on From M1 MacBook to Arch Linux: A month-long experiment that became permanenent
I think so far some of the surface devices and some of the Razer (yes, the one making computer mices, keyboards and such) had been the closest.
asabla··on I run a full Linux desktop in Docker just because I can
I still remember how much I liked the idea. Really tried to use it, but the experience with both browsers and vscode was....not that great.

Kinda hope they revisit this idea in a near future again

asabla··on AGENTS.md – Open format for guiding coding agents
Been using a similar setup, with so far pretty decent results. With the addition of having a short explanation for each file within index.md

I've been experimenting with having a rules.md file within each directory where I want a certain behavior. Example, let us say I have a directory with different kind of services like realtime-service.ts and queue-service.ts, I then have a rules.md file on the same level as they are.

This lets me scaffold things pretty fast when prompting by just referencing that file. The name is probably not the best tho.

asabla··on Counter-Strike: A billion-dollar game built in a dorm room
Squad, Arma (and especially Arma reforger), Dayz, Battlebit, Heretic + Hexen and thr list goes on.

Arma usually gets the more complex and janky stuff (in a fun way). While the others are more modified experiences.

Like Squad, we're they've re-created star wars battlefront

asabla··on Counter-Strike: A billion-dollar game built in a dorm room
Mostly AA and indie game titles. The simulator scene is still going strong with dedicated servers (like squad, arma, farming simulator, the hunter etc etc).

Larger titles swapped over to more control in order to extract more money from the players, but also control the experience.

There is however some AAA titles every now and then which support hosting your own servers. But they're quite few these days

asabla··on GPT-OSS vs. Qwen3 and a detailed look how things evolved since GPT-2
> that rely on tool use for facts, and “knowledge bases” tuned for retrieval-heavy work

I would say this isn't exclusive to the smaller OSS models. But rather a trait of Openai's models all together now.

This becomes especially apparent with the introduction of GPT-5 in ChatGPT. Their focus on routing your request to different modes and searching the web automatically (relying on an Agentic workflows in the background) is probably key to the overall quality of the output.

So far, it's quite easy to get their OSS models to follow instructions reliably. Qwen models has been pretty decent at this too for some time now.

I think if we give it another generation or two, we're at the point of having compotent enough models to start running more advanced agentic workflows. On modest hardware. We're almost there now, but not quite yet

asabla··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
Initial testing has only been done with ollama. Plan on testing out llama.cpp and vllm when there is enough time
asabla··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
I'm on a 5090 so it's not apples to apples comparison. But I'm getting ~150t/s for the 20B version using ~16000 context size.
asabla··on Building MCP servers for ChatGPT and API integrations
For ChatGPT and DeepResearch yes, not when using the API. I guess you could just return empty results if you want to offer other tools as well (can't test it now, since custom connectors only supports Workspace or PRO accounts for this moment).

Quote we're talking about: > To work with ChatGPT Connectors or deep research (in ChatGPT or via API), your MCP server must implement two tools - search and fetch.

Reference links:

- Using remote MCP servers with the API: https://platform.openai.com/docs/guides/tools-remote-mcp

- Which account types can setup custom connectors in ChatGPT: https://help.openai.com/en/articles/11487775-connectors-in-c...

asabla··on Understanding Tool Calling in LLMs – Step-by-Step with REST and Spring AI
I know I'm a bit late. But for MCP servers running over HTTP/custom it should use OAuth 2.0. If it's served via stdout, it should use configuration files and/or environment variables.

ref: https://modelcontextprotocol.io/specification/2025-03-26/bas...

asabla··on AI note takers are flooding Zoom calls as workers opt to skip meetings
It's not about being a prima Donna. It's about business value. Too many meetings over the years should either be better planned, not taken place at all or could have been an email/chat message.

Business value first

Page 1 of 5Next →