HNHacker News
TopNewBestAskShowJobs

mchusma

4,645 karma · joined February 9, 2011

Founder & CEO SignNow Founder & CEO tidy.com CTO HotMic
submissionscomments
mchusma··on Inkling: Our Open-Weights Model
I don’t know if it’s a great business model but it makes perfect sense to me. Open models when fine tuned are capable at better than frontier performance at a fraction of the price for many (probably most) domain specific tasks. If companies help make that easy to implement, there is value to capture. But I kind of like Unsloths model here which is to be really good at just layer, and not bothering with building their own models.
mchusma··on Neko Health Raises $700M
Non paywalled alternative: https://www.theverge.com/science/965849/spotify-founder-ek-s...

This looks similar in concept to the Midjourney health, but much more about multiple modalities (different surface scanners). Its interesting, I do think it could be better then dermatologists at picking up skin issues if widely deployed.

mchusma··on Bonsai 27B: A 27B-Class model that runs on a phone
I do think this would be interesting if they made these easy to finetune, as I do think this level of intelligence is likely sufficient for many applications and could be extremely cheap to run.
mchusma··on Australian energy retailers must offer three hours of free daytime electricity
Incentivizing usage during peak times makes total sense, but if price swings are this wild, how are grid scale batteries not highly economical? My rough ballpark math was that you need roughly 20 kilowatts of battery storage to make this issue basically nonexistent, and that would cost about 10 billion dollars, which doesn't seem that much for this.
mchusma··on Precursor
I would say that overall there are pros and cons to this, I really want to be allowed to use agents on my behalf, and don't want to see sites prevent me from doing this. On the other hand, I do recognize there are cases when its good/ok to have only humans allowed to take some action. In my opinion, the line is likely when you are representing you are a human, its ok to prevent bots. otherwise, you can't.
mchusma··on LAPD lets contract with surveillance giant Flock expire
I stayed in downtown LA recently and looked like the set from the walking dead. Literally blocks of people wandering in traffic. I guess you could argue you definitely don't need flock cameras to see the problem, but also I don't know how anyone would not do everything possible to stop it.
mchusma··on Former NOAA employees built Climate.us to preserve climate data and resources
I agree with parent, the full quote is: "The whole thing relies on donations to keep it afloat, which is really what tax dollars are for."

I think this is a great site, love what they are doing, and support them (including a literal donation). But a government maintained website for this data is low on my list of things of what tax dollars are for. In fact, I think this is better done privately. To be clear, many of the things every US administration does including this one I also think is better done privately.

mchusma··on A full body MRI earns you a year of smoking
I like the linked Scott Alexander post, but I also genuinely wonder what is the rate of change on these tests? The linked test Prenuvo has competition from Ezra + Function and others. It this drops from $2k to $500 over time, it makes it look considerably better. The more we can use different testing modalities, we should be able to reduce false positives in each modality.

I will say, that for cancer specifically, tests like Galleri seem better, but as that cost comes down I could see in 5-10 years an annual $500 scan that offers a full body scan of some kind, plus comprehensive bloodwork including blood cancer screening, and the type of thing that could be done annually by many in the US.

mchusma··on A full body MRI earns you a year of smoking
I was similarly confused. Saying a MRI is the equivalent of stopping smoking for 1 year earlier, or driving a motorcycle 10,000km less seems actually really good! Go MRIs!

As another point, most of the negative costs of getting full body scans are actually poor reactions to the full body scans. The phrasing is "Hey, if you get more information, we are going to act badly on this information." I think the solution here should be just acting better on the information, not getting less information.

mchusma··on Show HN: Super Dario
This was my favorite as well.
mchusma··on Apple's new SpeechAnalyzer API, benchmarked against Whisper and its predecessor
I will plug Willow for mac recording. IMO it's basically to me a "better than perfect transcription" as it cleans things up and is almost instant. I liked Superwhisper but switched to Willow as it was a big difference.

Its so good that I'm not sure that it's possible to get any better. Speech to text seems like basically a solved problem, if not now then definitely in 5 years. I don't know if any of these speech to text businesses will work in the long run, but for consumers they are great. My guess is the 2030 version of Apple's SpeechAnalyzer will be so good that nobody will need to use 3rd party software.

mchusma··on We Must Act Now – A Statement on AI's Transformation of the Economy
I think the content here is not controversial EXCEPT that it sounds way too doomerish. The only mention of positive effects of AI is "It could bring...major gains in living standards." It really should say "In the US, social security runs out in 2032. 65 million people die annually. AI progress is critical to solving these and other issues."

Full text of statement: "1. AI may become radically more powerful over the next 10 years.

2. This could drive an unprecedented transformation of our economy, larger than the Industrial Revolution, but unfolding over a vastly shorter time frame. It could bring risks, including large-scale job displacement, as well as opportunities such as major gains in living standards.

3. Economists, policymakers and technology leaders must act now to understand the economics of transformative AI and to build the incentives, guardrails, and institutions needed to steer AI in a direction that complements humans and benefits society."

mchusma··on Autopsy Study Finds Replicating SARS-CoV-2 in the Hearts of Long Covid
After seeing studies like this, and how the shingles vaccine reduces dimensia, I have become increasingly convinced that it’s bad to get almost any disease, even transiently. I used to think that it was kind of good to train your immune system (kind of whatever doesn’t kill you makes you stronger). I no longer believe that. I believe that diseases often cause unknown effects and it’s better to avoid disease entirely, that vaccines are actually more beneficial than current studies show in this regard, and new universal vaccines to prevent the common cold and flu will likely have significant health span improvements over time beyond the acute prevention of symptoms.
mchusma··on Under federal rule, colleges must leave grads better off or lose financial aid
This is great. Federally subsidized loans is directly (not solely) responsible for rapid inflation of college costs in the US. Anything to limit its use is a good thing. I’d argue that this test would be better expanded to actually having an ROI, not just do no harm, to encourage schools to not only provide value but also constrain costs (eg your school may make sense at $10k debt not $100k debt).

This and/or making loans dischargeable in bankruptcy.

mchusma··on Muse Spark 1.1
I agree with parent, Meta has been at this a long time and its only because they have recently fallen off that they pushed this "oh give us credit its really a new org" thing. Basically, if you can't actually "win" then try to fake a restart and say we are the fastest.

Even given that, this is their second try (they had Spark 1.0). Spark 1.0 was uninteresting, this is potentially interesting, but we can't really try it yet it seems (at least not in Openrouter).

Ultimately, competition is now fierce in this broad level of intelligence/cost: Spark 1.1, Grok 4.5, GPT 5.6 Luna, GLM 5.2

Sonnet not in the same ballpark of pricing (more expensive than Opus in many cases). Haiku has been basically abandoned.

mchusma··on Muse Spark 1.1
This not being available on Openrouter really makes it hard to test. I was going to compare vs Grok 4.5 and GPT-5.6 Luna, but I don't want to deal with signing up for Meta for it unless it checks out. Please Meta make this available.
mchusma··on GPT-5.6
Looks like a great set of models, but there are about 20 different thinking/model levels here in this family and they are very complex to pick the right one for the task

E.g. for GeneBench Pro, it looks like you would always use GPT-5.6 Sol over Terra/Luna, its pareto optimal.

For Agents Last Exam, you would maybe want Luna, then Terra, then Luna, then Sol as you increasingly budget for tasks.

I feel that there may need to be a new auto mode in many of these cases. It selects the best model and thinking given a particular problem.

Feels like it's going to have to go that way eventually, because here we have about 20 different model and thinking levels you could use, and they're not obvious which ones are right for the given use case.

mchusma··on Muse Spark 1.1
Yeah, this is most directly comparable to xAI Grok 4.5. In both cases, directionally "opus level intelligence for haiku prices" which is a really big deal for application developers who want to include models like this in their applications. I have been testing switching out haiku and sonnet for Grok 4.5, and may give this a try too (it is quite a bit cheaper, particularly for cached).
mchusma··on Grok 4.5
Great model, very nice. Opus class performance at Haiku level pricing (or cheaper with the token efficiency). This seems like a GLM-5.2 killer and this is what Sonnet 5 should have been.

This is a model I could really see used inside applications, where Opus or Sonnet or GPT-5.5 are too expensive.

I would really like to see a strong Deepseek v4-Flash competitor, which ideally is something like Sonnet 4.6 performance at <$0.30 per token. This is missing from main US labs.

mchusma··on GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday
I agree. Gemini actually is pretty good for isolated components too. But fable is much better at design than opus or gpt5.5. I have not seen as much difference elsewhere, but definitely design fable is great.
mchusma··on Automating AI Away
I think scaffolds and the app layer are really the two big things needed for the deployment of AI in most use cases. In general, my company says for a given problem, we prefer deterministic software as the solution first, followed by LLMs, followed by humans. That's how we approach pretty much every problem. Yes, there are many things that we do with LLMs that we can eventually get to be done with software, and many things that are done by humans that we can get to be done by LLMs.
mchusma··on In Praise of Observational Evidence
Good article. I will only add that I think the problems outlined in the article are exacerbated by the regulatory apparatus. If this were a debate related to truth seeking alone, it would be good. But trials drive banning/unbanning of treatments and medicines, as well as their mandated coverage by insurance.

We could be much more flexible in our approach to things if we would un-ban things. For example after phase 1 trial or based on sufficient observational evidence, things can no longer be banned but have a higher standard for "insurance is now required to cover it".

mchusma··on Performance per dollar is getting faster and cheaper
I was hoping they would be discussing some path to improving things faster and cheaper. But in this post it looks like they offer quantized version for the same price as full version, and a fast version at much higher cost.
mchusma··on Physical disc production ending in Jan 2028 for new games on PlayStation
I do think if you are claiming something is not permanent you should basically have to make some specific claims about the timeliness of it (eg you can use it for at least 10 years or your money back).
mchusma··on Global review confirms mRNA vaccines are safe, effective and full of promise 
I really feel that many of the issues with mRNA vaccines and health studies in general are generalizations like “safe and effective”. Everything has statistical risks and benefits, and we should just share those front and center with people. Eg test results for X mean you have a Y% chance of having X, given your history and symptoms and other results. Here are low cost low risk marginal things you can do to improve statistical significance.

Similar for vaccines, just give us the numbers clearly and upfront.

This bypasses regulators from having to make claims beyond “we reviewed the data and agree with these numbers and feel that this should not be banned.” I do think it would also help to separate something “not banned” and being “required to be covered by insurance” or “required for professions like the military”. I think trying to simplify things makes things worse, because this abstraction is not real.

mchusma··on Claude Sonnet 5
Really if they wanted a standout model that would really take the wind out of GLM's sails, they should have made this the new Haiku, priced at Haiku levels with this performance.
mchusma··on Claude Sonnet 5
I agree with this assessment, IMO my takeaway from this is "Generally run Sonnet on low, otherwise use Opus". It's kind of like an "extra low" setting of Opus. (depends on the application for sure).
mchusma··on Claude Sonnet 5
This is much more interesting of a model at $2/$10 (their launch pricing) than at full price. There are many competing models at around this level of performance.

I also like that the difference between low, medium, high, xhigh seems more spread, which is actually a good thing for people trying to tune applications. Running Sonnet 5 on low with the launch pricing makes this potentially a better fit than Haiku or open source models for some tasks. I don't think it will make sense at full price.

mchusma··on Previewing GPT‑5.6 Sol: a next-generation model
I don't see this as that different. Anthropic was the first one to get involved in the "AI models must be approved" regime. OpenAI just has the advantage of being second.

(To be clear: I do not like this new paradigm)

mchusma··on Previewing GPT‑5.6 Sol: a next-generation model
I've struggled with this. You definitely can have great cheap models. There are many of them open source and served profitably by neo-clouds. The big labs have basically given up on cheap models, and it is frustrating. It means applications are not likely to build as much on them anymore (we are shifting workloads from Haiku/Sonnet to Deepseek v4, for example).

I suspect the problem is that they need to charge a lot to keep revenue numbers up, and they are more worried about cannibalizing themselves than others cannibalizing them.

← PreviousPage 4 of 34Next →