HNHacker News
TopNewBestAskShowJobs

jfaat

931 karma · joined January 7, 2014

joe@bonniebuilds.com

building cool ai stuff, hit me up!

submissionscomments
jfaat··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
You absolutely can. On a large codebase. Written in multiple languages. I also use gpt models and for most tasks I prefer ds. It can orchestrate really well, and it stays on task without going off to reinvent the w̵h̵e̵e̵l̵ state machine.
jfaat··on Show HN: Raven – The harness of harnesses, built for RSI
I'm using paseo heavily too. The main differences here seem to be around how opinionated Raven's orchestration is. They're providing agents, workflows, memory, skills, etc. Paseo gives you some orchestration tools but it's mostly letting the 'native' harnesses do the work. For my workflow, I'm interested in some of what they're doing here but I'm not in buying into their whole system (markdown for memory and calling it RSI, as someone else called out, is not doing it for me). The follow up actions and meta-harness tuning look pretty cool.

They also don't ship an app which is one of the best parts of paseo. Then again Paseo's perf leaves a lot to be desired.

jfaat··on Plan mode is dead
I'm using a set of skills for planning now which does something similar to plan mode but it's creating documentation that later skills reference it while building. It works so much better than anything else I've tried, because it keeps things on track throughout the feature buildout and across different agents/models/tasks. From what I read here (& in general) my take is that a lot of people haven't adopted something like this and it's just totally the wild west right now as everyone is cooking up their own flows.

I agree that as I move away from holding agents' hands through actual coding I need a different way to monitor what's going on. What step of the plan are we on, what are the tests actually doing, what agent owns what, etc. I haven't found a product that does a great job of that yet and it seems like the next frontier of the 'IDE' to me. It'd be more like an IM(management)E really. The closest thing I've seen was whiteboard [0] but I didn't have a great experience trying it out.

[0] https://dev.fast

jfaat··on Plan mode is dead
Do you mean when you're making changes to a production DB?
jfaat··on Waymo pulls over, calls cops on juvenile riders who had 'ghost gun"
Probably
jfaat··on TSMC Uses Old Fabs to Make New Chips [video]
> there are ordinary looking trucks puttering about on Taiwan's highways with likely tens of millions of dollars of chip inventory sitting inside. Like what happens if one of them gets into a crash? Will Nvidia miss earnings because of this?

Wow that's amazing. When he was talking about needing to move wafers around I was thinking surely they can't just like truck them around but nope, that's the whole thing

jfaat··on The AI Credit Resale Economy
By lighting VC (public soon) money on fire...
jfaat··on DeepSeek V4 Flash 0731
That's fascinating, it's WAY better than luna ime. What sort of things are you testing it for?
jfaat··on I'll be stepping back from leading product for X
> truly impressive growth hacking on children

I haven't been more viscerally disgusted by a phrase I've read about tech in a while. Astonishing word choice.

jfaat··on Gemini Robotics 2 brings whole body intelligence to robots
Your napkin math only shows that the market is identical. There's about 35m businesses operating in the US. So by selling to 10% of those you sell the exact same number of robots...
jfaat··on FAA lets Boeing sign off on 737 MAX, 787 airworthiness certificates again
In the US (parent mentioned US specifically) I think that's just Frontier now that Spirit is gone. I mean technically that's doable sure but idk if I would say trivial it's really limited on routes and the experience is terrible from what I understand.
jfaat··on LM Studio Bionic: the AI agent for open models
I disagree. The point of the frontier models is to do everything as well as possible ("AGI" race or whatever) but smaller models with some RL are going to be the clear winner for a ton of use cases. Think about all the use cases for LLMs that would never be economical at frontier inference costs, and in no way need it. You don't need or even want a phd polymath helping you with small productivity tasks that most people use computers for every day. It's often overwrought and annoying. I don't even really like the frontier models for coding for this reason. They're constantly blowing up scope and you have to fight it constantly.
jfaat··on M 3.9 Experimental Explosion – 147 Km ENE of Ponce Inlet, Florida
Is your point that that makes this better?
jfaat··on Kimi K3: Open Frontier Intelligence
Ha, Konami code and all! That's impressive. I think Kimi has had the best design sense of any frontier-ish model for about 6 months.
jfaat··on Kimi K3: Open Frontier Intelligence
Not to mention the bill you would be paying anthropic could be way higher at that scale
jfaat··on Female US rower completes historic solo journey from California to Hawaii
The extent of my expertise here is what I saw in a 30s video. I guess the point is it drifts while pointed in the right direction. That could be closer or further from the goal depending on wind and current but keeps you on track at least.
jfaat··on Female US rower completes historic solo journey from California to Hawaii
I saw some of her videos on instagram during the journey. There's a rudder that keeps the boat pointed in the right direction without propulsion while she sleeps.
jfaat··on Anthropic's Method to Losing Goodwill in a Few Easy Steps
That's why I'm switching to open weight models. I'm running models locally, self-hosting and using third party for various things. Like it says in the article each model has its strengths and none of the proprietary harnesses allow you to orchestrate different models depending on their strengths. It's cheaper, more flexible, private, more resistant to rug pulls, and brings back some fun to building for me vs just auto-accept, yell BAD CLAUDE when it breaks, repeat ad nauseam.
jfaat··on GLM 5.2 beats Claude in our benchmarks
Yeah this is sounds close to my workflow and its good to hear you've find a nice flow too! It frees me up to spend that effort on doing more things in parallel and focusing way more on the specs which is usually a good idea anyway.
jfaat··on GLM 5.2 beats Claude in our benchmarks
My whole point is that I don't want it to build an entire feature from one prompt. At most, I want to work with an agent to nail down the spec and then work with an agent that orchestrates the implementation via other agents, same for testing, etc. None of that requires frontier capabilities, it requires a little bit of work on a harness, a little bit more of my input, a little more of my brainpower. I _want_ to build tools that make it work better and don't change when the CC team gins up some default for their harness and foists it on me. I don't see that as a tradeoff at all and I think engaging in my work process more than fire and forget (and literally always in my experience fix stuff later) is more fun and rewarding once the 'holy shit this is now possible' high wears off. Doubly so once the frontier model gets nerfed mid-cycle and now I have to undo the mess because they released v*.x++ and I fell for it again by trusting it to do these agentic tasks without my involvement.
jfaat··on GLM 5.2 beats Claude in our benchmarks
We built our own and aren't done open sourcing it but before that I got to a really good place with opencode plus some custom agents, pi family is good too although I haven't used it as much. We made an agent to design a spec, one to implement by dispatching subagents, one to validate against the plan, things like that. All of this helps claude/gpt too IME. For open models it has helped them stay out of loops (e.g. Kimi's but WAIT) and for frontier it helps them stay on task and not invent bloated patterns
jfaat··on GLM 5.2 beats Claude in our benchmarks
> but if you only want to use the best model available, it isn't there yet

I'm trying to wrap my head around exactly why so may people seem to want the best model available when it has recently become clear that most halfway decent models can write damn good code for a fraction of the price. And the frontier models get nerfed constantly so you with open weight you can get something slightly less performant but way more stable. Almost like buying a Ferrari for your daily commute instead of a Toyota or even a Mercedes.

I think there are several factors. Certainly marketing making us think we need the shiny thing which is rampant online and very smart people think they aren't susceptible to. There's a lot of really odd 'I trust Anthropic/OpenAI more than Deepseek' which tends to ignore, for starters, that you can run choose your provider and still save a ton. I also think there's some amount of addiction and brand loyalty where a Ferrari is one hell of a drive so that you turn your nose up at that sensible Toyota. Oh the other one I see used is like oh only fable can oneshot updating my embedded systems thing from 1975 to rust which is great but let's recognize how niche that is.

And it ends up just coming across as people are getting SO reliant on the tools so fast. Maybe it's ok to think and like read a few lines of code and work with these agents to convert your thing to rust or center your div. Even if coding is over which in some sense it certainly is, don't turn your mind into the wall-e people yet. I found myself guilty of this so often. It takes way more time and effort to do things via prompt and I wouldn't just open the editor and fix it because that dopamine hit of the magic the abstraction provided was so strong.

So I'm pretty much done using the 'best' (on benchmarks, if money isn't an object, etc etc) models available. After a year on Sonnet/Opus/GPT5x I'm having way better results with open weights models that don't get lobotomized weekly. I'm finding ways to do the crafting part of building software by focusing on honing my harness and workflow. I'm enjoying changing the oil on my Toyota after a year of almost flying off cliffs in my Ferrari and if I can check my ego it's a purely positive thing.

jfaat··on What Ozempic does to the gut-brain axis
Known to who? I just did a quick search didn't see anything that like a scientist or physician said about this. Someone who describes himself as a 'biohacking educator' said it without citing sources though.
jfaat··on The gap between open weights LLMs and closed source LLMs
I'd call 'more likely' an extremely safe take given that it's exactly what's happening right now
jfaat··on A recent experience with ChatGPT 5.5 Pro
I'm going to start using malapropisms so people know I didn't use an llm to write things
jfaat··on Where things stand with the Department of War
Which of the countries that the US has recently attacked are you comparing to Nazi Germany?
jfaat··on Over 80% of 16 to 24-year-olds would vote to rejoin the EU
With what's happening in the US post covid, I'm gonna have to disagree
jfaat··on GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
Yeah that's a good idea. I played around with kimi2.5/gemini in a similar way and it's solid for the price. It would be pretty easy to build some skills out and delegate heavy lifting to better models without managing it yourself I think. This has all been driven by anthropic's shenanigans (I cancelled my max sub after almost a year both because of the opencode thing and them consistently nerfing everything for weeks to keep up the arms race.)
jfaat··on GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
Lol wat? I mean you certainly have enough control self hosting the model to not let it join some moltbot network... or what exactly are you saying would happen?
jfaat··on Show HN: LemonSlice – Upgrade your voice agents to real-time video
I'm using an openAI realtime voice with livekit, and they said they have a livekit integration so it would probably be doable that way. I haven't used video in livekit though and I don't know how the plugins are setup for it
Page 1 of 6Next →