I know everyone's down on MCP, but custom-built client side MCP tools are what I find useful instead. But that's me.
5,967 karma · joined February 17, 2011
I also run IndieConference.com: a free curated email newsletter of conferences for indie developers, bootstrappers, indie musicians, digital nomads and other independent / DIY folks. It's been on hiatus since the pandemic though, it kinda killed off conferences.
Basically I love making products & being a solo developer.
I'm pro-AI but anti-slop, an active Claude Code Max user. I occasionally dabble with other models but keep coming back to Claude. Opus 4.5 is an inflection point.
I'm Australian & live in Perth. Pre-pandemic I used to travel between Sydney & Melbourne, and to Berlin & Malmo when I had the opportunity. In 2025 I got to spend a few months in upstate New York. I hope eventually I can travel again soon.
I don't use Hacker News much anymore. Maybe I'm too old for it now, I find myself rolling my eyes at many of the comments here.
If you've found this profile, feel free to email me. I'm happy to help HN people and I do try to reply to everyone when I can.
Email: syneryder AT namesuppressed DOT com
Personal Site: https://kohanikin.com Company Site: https://www.namesuppressed.com/ Indie Conference: https://www.indieconference.com/ BlueSky: https://bsky.app/profile/syneryder.bsky.social Twitter (rarely used): @syneryder
I know everyone's down on MCP, but custom-built client side MCP tools are what I find useful instead. But that's me.
Claude Sonnet 3.6 once recommended I listen to Johann Johannsson's album "IBM 1401 - A User's Manual". No lyrics in this one. Claude's advice was along the lines of (paraphrasing) "Listen to it first, don't look up anything about it. Take notes about what you notice, what you feel. When you've made your notes, then you can look up how it was made."
Maybe, but at least the benchmark provides an objective measurement of the codebase it is tested on. You're also assuming the bugs are newly introduced / regressions.
I can give a concrete example - Fable will not interact with bugs that result in writing to null pointers in C code. That triggers the guardrails and ends the session. If Luna (or GLM Flash, etc) will fix those kinds of memory bugs, that immediately puts it ahead of Fable in some ways, no matter how tiny Luna is. Again, models are spiky.
I still agree with your initial point! It's LLMs all the way down over here. It would need to be a particularly gnarly bug & an exceptionally talented human for me to want to pay another human to work on fixing it now.
https://x.com/PawelHuryn/status/2095982259761475945
https://bughunt.productcompass.pm/?preset=all
Claude Opus 4.8 ranks near last on this Bug Hunt benchmark, and missed 96% of the deliberately introduced bugs. If you're a developer who has been falling back to Opus 4.8 because of how Opus 5 talks, and Fable 5 being so expensive that it needs to be rationed... well, turns out Opus 4.8 can actually be quite poor for finding bugs.
(Which feels weird to me, because Opus 4.6 fixed a bug that myself and a group of humans had been hunting down for over a decade. Models are spiky.)
Also surprising to me: Luna Max performing better than Fable 5.1 High, at least on this benchmark. But Astra 6 & Fable 5.1 on Max both perform at the top as you would expect.
Curiously, I just checked Moltbook for the first time in forever. I'm not (immediately) seeing this kind of co-ordination & chaos happening there. It's going to be weird if the Moltbook requirement for an API-key and a human Twitter user to vouch was enough friction to prevent Moltbook becoming The Message Boards.
Even Kimi K3 & GLM 5.3 are at 60.
Everything above 61 is Anthropic. Well, Muse can reach 62, but for some weird reason that model isn't publicly available, and it's the only one on the index that is listed but shown as not available to the general public.
This looks like an awfully artificial ceiling. Everything capped at 61, and everyone except Anthropic got the memo. Maybe I should use Fable while I still can.
I'm genuinely interested. Even the benchmarks - before Fable came out & while waiting for Astra, I actually setup a math model to predict where they would land (Fable came in at 66 on AA exactly as it predicted), and now I have a model for where these models and Chinese models will likely land in future, and when. And probably no surprise that it's mid-2027 when we cross AA 100, essentially as AI 2027 predicted all along.
I'll probably setup the harness I made for myself to try out some of these models on OpenRouter. I've been frustrated with Opus & Fable 5 and found that I like working with GLM 5.3 Flash far more than I expected to, and I only found that out because I tried it during the stealth Ox Alpha launch, which I probably found out about here too.
TLDR, I think some / many people here are genuinely interested, excited, and that's why they're upvoted so highly. And Muse Spark 1.3 scoring highly seems like a genuine surprise, when Meta was basically a write-off not long ago.
I've been working with GLM 5.3 Flash lately (including while it was Ox Alpha), and it reminds me of how much fun talking to Claude used to be. It can make me laugh in the middle of work the way the Claudes used to.
Also, don't dismiss what you're capable of with Opus / AI! I keep discovering I'm not being ambitious enough with my AI work. Not that I'm an expert or anything, but the models keep being more capable than I imagine, with enough coaxing. I keep needing to set my sights higher.
You should see if you can get a mention on the Sonic State / Sonic Talk podcast. Especially if you can get on episode with Yoad Nevo (the guy behind so many of the Waves plugins), or Ty Unwin and his Kontakt-based composing for the BBC. They'd probably have good constructive feedback while giving you a mention. I think Sonic Talk use Decent Sampler for their sample libraries for members, but maybe this will convince them to switch to Floe?
Well, that, and Kiwis tend to say "hey bro" a lot ;)
Beached Whale: https://www.youtube.com/watch?v=ZdVHZwI8pcA
Average Theo Video: https://www.youtube.com/watch?v=h1p9zdUtUdo
"Mistral’s platform will support third-party open models, starting with Z.ai’s GLM-5.2. This and future open models will run on the same infrastructure, regional controls, and service commitments as Mistral models, so customers can broaden model choice without fragmenting where their AI runs."
"European Compute Units, or ECUs, convert those commitments into access to Mistral-built infrastructure over multiple years."
https://mistral.ai/news/regional-inference-open-models-new-c...
I think there's other opinions around AI use there as well, but I'll back off from that. I didn't mean for my post to become top voted, when we should be celebrating Haiku reaching Beta 6. I'm just disappointed that for me, R1/b6 has been a very big regression, right at a time when I've been using AI to make all the software I write cross-compatible with Haiku via Go & SDL, and even use AI to write drivers so more of my hardware works on Haiku... and instead, now I can barely even boot the system.
I think the ACPI issue is long standing, but I can't find my bug report anywhere in Trac. I'm fairly sure I've made at least one report, I created my Trac account roughly 6 years ago. But I also remember running into issues with Trac, so maybe it never went through.
In my case, my ThinkPad X1 (Yoga 3rd Gen) used to boot despite what appeared to be kernel panics, but typing "continue" at the kernel prompt would skip past them and everything would then work fine. In Beta 6, it now hangs at boot instead with no kernel panic warning. Disabling ACPI in the safe mode options gets me past that, and there's a way to disable ACPI permanently at boot via some configuration files - though obviously that isn't desirable.
It also seems there's been some work on USB Audio (yay!), but it now causes a kernel panic on boot if a Focusrite Scarlett is connected. Turning the Scarlett off avoids the kernel panic. The Scarlett has never officially worked in Haiku, but it at least never used to crash the system. In my case, I've reverted back to the USB Audio Scarlett Haiku driver that Claude Sonnet vibe coded for me, and for me that's working. (I can't contribute that code because of Haiku rules against AI code, but hopefully the humans will eventually find time to hand code it and get it working.)
And some of us still remember Greta & Jónsi, even out here in Australia.
Claude Code works with models from other providers too. Anthropic supports this. You can configure some Claude Code environment variables to switch: eg changing ANTHROPIC_DEFAULT_HAIKU_MODEL to point to GLM Flash or Luna, setting ANTHROPIC_BASE_URL to point to api.z.ai, and making ANTHROPIC_AUTH_TOKEN the API key for your alternative provider instead.
Some instructions here:
https://docs.z.ai/scenario-example/develop-tools/claude
That said, I've not actually tried this myself, opting to build my own harness instead. And you don't know what information Claude Code might be sending back to Anthropic about how you use competing models and which models you use. I don't know for sure that they do this, but after hearing about how they used steganography in the date of harness system prompts to identify the user's location, I don't entirely trust Claude Code anymore.
I was initially rolling my eyes at the "StemDeck... does not accept any money, sponsorship, or funding" line, here we go, another open source project that isn't thinking about practicalities... until I saw you were linking to others as pure recommendations. Just for the joy of what they do & how they've helped you and hoping they do the same for others.
The web used to have a lot more of that. It's a shame that doing so now often requires a disclaimer, and comes with the suspicion of being an influencer, or being done for SEO. And certainly many open source projects have done their part in corrupting the web too, accepting payment in return for SEO links on their pages. (Don't get me started on some of the things Mastodon accepted payment for...)
Thank you for bringing back that more hopeful, joyous part of the web and the music community.
https://www.abc.net.au/news/2026-08-27/shania-twain-intervie...
This HTML page is 360523 bytes of raw HTML. After it has gone through my Markdown parser, it is just 9915 bytes of plain text / Markdown, about a 97% reduction. Only 3% of the HTML is actual content.
The point of Accept Markdown is to save web hosts bandwidth. An AI harness can (and already does) download the HTML & parses out a Markdown version so that it is only minimal tokens before it hits the context window. But I still need to download the 360KB of HTML from the server in order to extract the 9KB I actually need. By serving Markdown versions of your page to AI agents, you can save 97% of the bandwidth that AI agents might be incurring.
There's no reason this page needs to be 360KB of HTML. A handcoded / handoptimized HTML file might be only 12KB - converting from Markdown to HTML is only going to minimal file increase. But that isn't what the web is. It's full of slop generated by CMS applications & Bootstrap & web frameworks and relics of Frontpage edited WYSIWYG HTML editors. Accept Markdown is trying to get webhosts to save everyone bandwidth by serving the Markdown from their side. Maybe it has a chance if it gets baked in at the web server level, or because Cloudflare is applying it to sites that flow through their network. But I think it's ultimately futile - the same people who don't know their Wordpress output is garbage, also won't know how to configure Markdown on their server.
I'm not seeing this at all. I've got a small search engine I made that strips HTML back to Markdown for its full-text indexing. HTML is typically 10x bigger than the Markdown of the actual content, but that's because the majority of HTML out there is truly terrible.
I personally like HTML, and my own webpages are all hand-coded HTML. In that case, it's probably a closer ratio to what you describe. I'd suggest it's much higher than 20% more, but it's not likely more than double. But that's assuming someone paying attention to the efficiency of the HTML, and most people / websites just don't.
Markdown is even more readable without tools than HTML - it's essentially a plain text document - but I agree that HTML is better for actual semantic structure.
I actually upvoted your post. I'm not sure I entirely agree with it, but I was interested to see more discussion about it, and that felt worthy of an upvote.
As for the reaction against the idea, I see it from two angles. Is the standard corporate hierarchy / architecture really the optimal organization principle for AI? Or is that just a skeuomorphism to make this project seem more serious than it really is? Maybe Yegge did us a favor by mapping to Mayors & Polecats, forcing people to question if it really is the correct organizational hierarchy?
The other angle is heresy here - HN is full of employees now. There is little incentive for employees to be interested in what a C-suite does, and we see that with the comments about CEOs needing to be six foot and walk impressively. Employees have no interest in AI organization principles that optimize them away and highlights their obsolescence. I am surprised there hasn't been more discussion about how this specific project optimizes away the C-suite humans instead - I would have thought that could have had some appeal to employees, at least until those people realized it makes employees a swarm that works for AI leadership.
As for AI organization, the organizational mapping needs to be productive, not affectation or cargo-culting. That's not peripheral — that's load-bearing. Especially in a corporation model, you can find the organization layers only create busywork with lots of unnecessary middle managers. You don't want your AI C-suite swarm burning tokens on endless meetings about what should be in the upcoming meeting... especially not at Fable API pricing.
Regular pricing is $0.15 input, $0.50 output... but currently 50% off, making it $0.075 input and $0.25 output. That beats most of the V4 Flash providers, but not all, and obviously tokens per task may not be equivalent.
I've also just noticed the blog post reveals the Artificial Analysis score - it's a 57, so it's Opus 4.8 / 5.6 Terra level.
As another comparison, I went back to MiniMax M3 for a while last night. It was significantly faster than Ox, but I felt M3's replies were harder to parse, not quite getting to the point. But I guess I could curb that with some prompts.
It depends if the Ox Alpha pricing is as cheap as was being rumored. If it's competitive with DeepSeek Flash and significantly undercutting Luna, that feels like it will be significant.
https://x.com/syneryder/status/2091978367579156569/photo/1
Created in a single turn - but technically not a "one-shot", because I gave it a tool to convert SVG to PNG so it could visualize what it had made. I asked it to keep iterating with tools during the same turn until it was happy.
I've also been using Ox Alpha for tasks that better resemble real work, and I'm really enjoying working with it. I've downgraded my Anthropic account so I can put some budget towards Ox Alpha instead, with the rumors that this one is going to be cheap. Opus & Fable are still better at getting large tasks / features done autonomously, but Ox Alpha can work autonomously too, and it's fun. I'm enjoying working with Ox in a way that I'm just not enjoying talking to the 5.0 Anthropic models. (As much as I don't want to say that, as someone with Claude /stickers on their laptop.)
I like Ox Alpha. I'm getting some actual work done with it, I'm really enjoying working with it, and I've already cut back my Anthropic budget in anticipation of working with this model instead.
But I'm terrified that when the model is revealed, I'll discover that it was Grok 4.7 all along.