HNHacker News
TopNewBestAskShowJobs

mchusma

4,637 karma · joined February 9, 2011

Founder & CEO SignNow Founder & CEO tidy.com CTO HotMic
submissionscomments
mchusma··on Gemini 4 Argon (High): Intelligence, Performance and Price Analysis
I’m guessing this is considered something like a C grade from Google if they are being honest with themselves.

After being nowhere near the frontier for a long time, they are pre announcing a model that ranks 3rd, roughly on par with models today that are cheaper.

Good for them to think about releasing to stay in the frontier game.

(I do think 3.7 flash was a solid release, so they are around the conversation. And their image and audio and live models are good)

mchusma··on Gemini 4 Argon
Cool. But until I can play with it it’s vaporware. I will say I think 3.1 pro is still really good for legal type things. They cooked with that one.
mchusma··on Responsible Release of AI-Generated Mathematics
Either mathematical progress helps advance society, in which case progress is a good thing.

Or mathematics is more like a hobby, and while ai may spoil their fun, they need to move on like chess and go players.

mchusma··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
47 is my crossover for serious things (e.g. Grok 4.7 is below the line and GPT 6 Sol is above the line). I mean, Opus 5.5 is way better, but GPT 6 Sol still gets the job done for anything that doesn't require design thinking.

Although I do think Luna 6 max is ok for some basic things, would never use it for coding myself.

mchusma··on Dots: Always-on agents
Jev is orders of magnitudes cheaper and faster, its a very interesting new paradigm. (Jev pioneered this about 2-3 weeks ago).
mchusma··on Sonnet 5.5
My guess is that haiku will be a mid release. They don’t seem to care to compete for the low end. Something akin to gpt6 sol level intelligence at $1 / $5 pricing. Then not release an update for 6+ months.
mchusma··on Sonnet 5.5
I feel like sonnet is priced too close to opus right now. If Sonnet 5.5 were half its current price it would make sense to use. At its current prices, I won't use it in applications (I would use cheaper models) and I won't use it in my subscriptions ( just use Opus instead). At least that is my initial reaction.
mchusma··on Nissan's third generation e-POWER powertrain
I agree I don’t like having both engines, more to fail. BUT I could see a world where you basically have basically an EV with a smaller battery pack (150 mile range) and basically a swappable $300 generator. I saw swappable so that if anything goes wrong they just swap it for $300, and credit you back async any recovery value. Just electricity generation from it, for rare needs of longer range. I think in theory, it could be a good way to get a cheaper EV supporting rare bursts of range and low maintenance cost.

But I’m not even sure regulators would allow such a design. Even then I’m not sure it’s worth it.

But I am totally done with ICE car maintenance. I want nothing to do with it anymore.

mchusma··on SpaceX's Starship launching to orbit for first time ever today
I think for a sense of scale if this works as the SpaceX team wants, they will be able to launch more mass to orbit in a single day than all payloads ever before 2022 (rough math). And at a cost roughly 1/100th the prior average.
mchusma··on Tokens too cheap to meter
This I agree with.

Right now, I do actually use OpenAI's gpt-oss-safeguard-20b for somethings, was released 11 months ago, and is $0.075/M input / $0.30/M output now. I could see this model being in fairly widespread use at 10x speed and 1/10th cost if it was introduced today. Meaning, that for some usecases (moderation) i think dedicated chips can pan out today.

But for more general models, its tougher. Gemini 3 pro was launched in November, if ASICs brought it down 1/10th in cost, it would be $0.20/$1.2. GPT 6 Luna is $0.1/$0.50. Luna is better at a lot of things, but not everything. So 1/10th doesn't really make the ASICS investment worth it in my opinion, but if it brought it down to 1% ($0.02 / $0.12) it would be a really compelling model with a lot of use.

BUT, do i think something like Luna is probably generally capable of doing a huge amount of knowledge work. So if Luna came out at 1/10th the cost a year from now, it would probably be compelling for a while.

It all depends on the rate of improvement in cost/capability.

mchusma··on The Price of Intelligence Is Falling Rapidly
What? The report is about ai performance and cost.
mchusma··on GPT-6 Sol and Luna
My initial takeaway is that GPT-6 is mostly a lower cost win, for Luna. GPT-6 Max is an upgrade on intelligence too, but its mostly a cost play (which is great, not complaining).

Overall, I expect for most people think the winner of today was Anthropic. I personally am preferring Opus 5.5 at medium over GPT-6 Sol Max, in very very early tests. Similar price range, more capability.

But competiton is great, these are solid releases by OpenAI today.

mchusma··on GPT-6 Sol and Luna
Opus 5.5 is incredible so far, its going to get used. Fable is much better than Astra for me in practice, and Sol is not marketed as better.

Its a great release, I will use both heavily.

mchusma··on Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
"High" to me looks like the one to use. https://artificialanalysis.ai/models/claude-opus-5-5-high

Many benchmarks start to plateau after high, this benchmarks better than Fable, and my initial tests show it working really well.

mchusma··on GPT-6 Sol and Luna
What a day! I couldn't really use the last Luna for much (wasn't smart enough) or Astra (too expensive). So this release is really exciting. I can probably use Sol 6 as much as I want in the week, which as great.
mchusma··on Claude Opus 5.5
I hope Haiku is Pareto better than Luna/Deepseek, so slashing its price by about 90%.
mchusma··on Grok 4.7
Coding, agentic flows like logging into my accounts and gathering data, grok bot.

4.6 made more mistakes than SOL or Opus overall. Gave up a lot. And in my opinion, the rate of mistakes is kind of more important than how brilliant it is.

I think 4.7 may still be better, but I was hoping for clearly Sol/Opus level and so far it just isn't there for me.

mchusma··on Grok 4.7
Initial impressions, Grok 4.6 for me just didn't really hack it for any usecase I tried. I seem to have a floor for my usecaseses (coding and a bunch of agentic workflows) and Sol/Opus are above some kind of intelligence floor.

4.7 is definitely slower & more expensive. It feels kind of like they really had it burn tokens to claw up the benchmarks. But it's not super clear to me whether it's above the line or not. A part of that is that it is so slow that i haven't been making fast progress today with benchmarking it.

Overall, it it gets above my intelligence line its a good release...but you can read the tea leaves and tell the Grok team thinks this was a miss.

mchusma··on Uber arbitration award over Emily Normandin-Parker’s death
If you have been through both processes, you would more likely say the traditional civil process should be illegal.
mchusma··on Uber arbitration award over Emily Normandin-Parker’s death
I have been on both sides of arbitration, winning and losing. It’s much better. Basically legislation done right (for civil matters).

The only people who really win from traditional legislation are lawyers (and plaintiffs counsel who use the long expensive process to blackmail people - which is 90% of civil cases)

mchusma··on Gemini 3.8 Live and 3.8 Live Extended Thinking
Their live models have been and continue to be at the frontier. I like them a lot!
mchusma··on Gemini 3.8 Live and 3.8 Live Extended Thinking
Yeah its good. Reasonably priced too (at current prices, if they do raise them in January I would stop recommending it). 3.7/3.8 were good releases.
mchusma··on Steam Frame starts at $1059
With every VR release I check the resolution per eye to see if I can use it as a work desktop replacement. IMO the current Meta Quest 3 is not there, and this roughly matches it. I am not sure what the resolution needs to be to render good text...but I'm hopeful we get maybe 50% more pixels at some point in some version. Unfortunately, i think the screen stack at that level is well outside mobile usecases so there is kind of a chicken and egg. Manufacturers don't want to work on advanced VR display tech without demand, and it can't be used for work purposes until display tech improves.
mchusma··on Pion, an agent designed to run any company autonomously
Our control panel is on a server but individual agents actually are run on anyone’s machine. This allows us to use the native harnesses including subscriptions. Yes it uses a lot more tokens, but i have have Claude $200, OpenAI $200, and SuperGrok Heavy $300 (or whatever it is called) that includes Cursor Ultra. My machine runs most of them, but some other team members have agents running on their machines using 1-2 $200/mo subs.

I would say this setup probably costs us about $1,000/month total. (I’m excluding traditional engineering use of LLMs from this number. This is the cost of all the “AI employees”.

One reason we did it this way was to use subs.

mchusma··on Pion, an agent designed to run any company autonomously
I hope to post something within a few weeks. But our philosophy is generally if it can be done deterministically with software (eg run payroll on autopilot via an api to gusto) then do that. If it can be done by an ai agent, do it with that but add as much software as possible to make it reliable at that thing. And the ai agents are all tuned to escalate to humans as needed. Then there is also just things that are completely human.

A typical example of something that is AI vs human is the AI most commonly operates like “managers”. For example, reviewing transcripts of every demo call, compiling results, figuring out insights, learnings that need to update our company docs, feedback to humans (who run the demos).

We are big fans of having AI agents “own” koi’s because now anytime we say “we really should be doing this” we try to set it up on the spot.

The “downside” here is that I do occasionally get busy, and if I’m the only one who can approve or unstick one of these bots, it just keeps harassing me until it gets done. This is a sign generally that I need to hire someone to own a set of bots.

mchusma··on Pion, an agent designed to run any company autonomously
We do have a control panel, but it’s in effect a server that has things in a database. All chats, tasks, assignments are all there. This felt easier to debug and manage versus putting it all in slack, although we considered it. I don’t think anything we are doing is particularly “fancy”, but basically one agent sends a message to another one. It saves the message in the db, then adds the message to that other bot same as any other user chat. We have an internal website where anyone in the company can see the bits they are authorized to see and can see all chats. These AI workers are single threaded, but we see that as more of a feature than a bug (minimize complexity). They all work their way through a shared task list, which is just another table in our remote server sqllite.
mchusma··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
We had a two human PR requirement until recently we dropped it. It was slowing us down too much now the human developer creating the future is obviously writing it all with AI so they need to check it then depending on the feature and it’s use it requires a PR but it’s not universal and we’ve stepped up our automated test Tan X what it used to be it’s been so far fewer bugs better delivery
mchusma··on iOS 27, iPadOS 27, and macOS 27
Apple's search is so comically bad. For example, I can click on my applications, and I have an app called Chess. I can often type letters C-H-E-S-S and have it not show the Chess app. This is inside the applications folder!

Many people have complained before. Honestly I ONLY search for app names in spotlight search or application window, it should be trivial to make it actually work. Maybe i just need to vibecode some alternative search bar :(

mchusma··on Pion, an agent designed to run any company autonomously
There was surprisingly little information on how they actually do this, but we run our business with a large number of, what we call "AI employees" in addition to regular employees, and they act in interesting ways. We've been building out orchestration tools to handle this, and we'll probably do a write-up or blog post on it soon. May even open source some of it.

The preview is that the problem with most agents (and this includes frameworks like Grokbot and Openclaw and Hermes) is that for many of them, they're black boxes. They say they learn or improve, but it's a black box in what they do. Getting agents to reliably do things is hard, and getting agents to build out software tools to help themselves improve and do better over time is also hard.

Our approach at a high level is pretty simple: every single AI employee is a standalone GitHub repo that shares some characteristics, but we direct them to build as much software as possible to make their goal as easy and reliable to manage as possible. Then we have a shared communication layer for bots across the company to interact with humans and AI. We have decided to organize these like departments similar to the way you might hire out humans. I'm not 100% sure if that's the best approach, but I will say it's been easier for people to understand because they're more mentally easily able to traverse the bot org chart if it somewhat reflects a traditional business org chart.

Each of these AI employees has specific sets of goals and KPIs, instructions that they manage the business with manager bots. We have layers of management, which we actually have found helpful. We also run different bots with different models and harnesses, and some using different models and harnesses to check the work before anything can get done, along with lots and lots of testing.

Every single time, actions have a massive amount of tests based off of previous failures to prevent failures in the future. Sorry for rambling. I do think this is a very interesting space. I didn't see anything interesting in Pion that was public on this website, but I do anticipate that more companies will be "AI and software first," as in the substrate of the company is basically a software application powered by autonomous agents, with humans as a fallback.

mchusma··on The case against JPEG XL
TIL that AVIF supports progressive decoding, cool!
Page 1 of 34Next →