HNHacker News
TopNewBestAskShowJobs

spuz

4,868 karma · joined June 18, 2009

submissionscomments
spuz··on Gemini 3.8 text-to-speech
Arguably, no AI voice should sound like a human voice:

https://youtu.be/M-IVVJkZnuo?t=236

spuz··on XCancel service is suspended until further notice
The decision to force you to sign in to view replies is clearly not by accident or ineptitude. It's probably not malice either - just a business decision.
spuz··on Hackers Had a Live Feed of Every ID Verification Company Scanned for over a Year
Haha, damn not enough caffeine this morning to parse correctly
spuz··on Hackers had a live feed of every ID verification company scanned for over a year
Am I missing something? What do you mean the DMV makes tens of millions of dollars a year selling data to itself?
spuz··on Go grandmaster Shin defeats AI KataGo with a two-stone handicap
What does "on five handicaps" mean? The human received a 5-stone handicap?
spuz··on Unusual Suspects
I can't wait to see how actual artists do with this game.
spuz··on Claude Fable 5.1 and Claude Mythos 5.1
No, training tends to override system prompts at lot of the time.
spuz··on Nancy Grace Roman Space Telescope
I guess for one thing, these launches don't tend to fail and second, they actually do make duplicate missions sometimes. See Voyagers 1 and 2, Spirit and Opportunity and Curiosity and Perseverance.
spuz··on A 3D fruit fly on macOS desktop powered by the real FlyWire connectome
I didn't downvote it but I imagine it's because your comment comes across as manic rambling. You managed to say "it will run in the browser" three times - once is enough. I'd recommend writing out your entire comment then deleting it an rewriting it to be more concise. No we don't expect literature in HN comments but good writing is a courtesy to your readers.
spuz··on Climbing Guide as a Shared Infrastructure
This article seems to be an AI warping of this original article which contains a lot more useful details and screenshots of what the application looks like: https://community.openclimbing.org/d/24-v200-a-big-step-forw...
spuz··on DeepSeek peak/off-peak pricing update
What do you mean? Most providers on OpenRouter offer the same or lower prices than DeepSeek themselves:

https://openrouter.ai/deepseek/deepseek-v4-flash#providers

spuz··on DeepSeek peak/off-peak pricing update
I wonder whether all the DeepSeek providers will follow suit or are they going to try to stay competitive with the old prices?
spuz··on Human vs. AI – Diff-based line-level provenance for text under agentic editing
Does Claude run git commit on every change it makes? Doesn't that pollute your git history?
spuz··on Human vs. AI – Diff-based line-level provenance for text under agentic editing
> A git repository is already a history of versions each carrying a provenance marker — every revision of the file, in order, with the author of the change that made it.

Maybe I'm missing how people use AI these days but when I have an agent working locally, all git commits have my authorship attached.

spuz··on Understanding the AI Economy
It's not a report about the economics of AI. It's a report on how people are using Google's AI. They don't mention anything about the supply side.
spuz··on You only need the frontier model for one single edit
I cannot imagine local models running on an igpu could get anything close to either a useful plan or execution of a plan. I've tested Qwen 3.5 27B locally and its solutions to coding problems are usually flawed and running in thinking mode is too slow and that's on a discrete GPU. How do you get anything useful done on a model that runs on an igpu?
spuz··on 1-Bit LLM in the Browser
Unfortunately, this doesn't work for me. After loading for the first time, my first prompt had it generate an infinite series of exclamation marks. My subsequent queries just had it return nothing. I have plenty of RAM and VRAM. This is with Chrome on Linux Mint with 32GB RAM and a 12GB 3060.
spuz··on Claude Fable produced a counterexample to the Jacobian Conjecture
Thanks that's a very nice summary.
spuz··on $100 AI Music Video: Claude Fable 5 vs. GPT-5.6 Sol
The problem with this discussion is we are conflating artistic value with economic value. A music video can be valuable as art, but it can also be valuable as a tool for promotion or for generating revenue. If the goal is pure economic efficiency then an AI has the potential to create a music video more efficiently than a human. If the goal is to produce something that makes people feel some kind of emotion then AI will work against you. The same goes for code. I use AI to write code that I'm not precious about but if I want to feel proud of my work, you can be sure I'll write it by hand.
spuz··on Ice water drowning survival of young patient (2025)
The article describes their decision making process:

> As rescue divers searched for the boy's body, we deliberated whether to attempt resuscitation and likelihood of meaningful neurologic recovery of a child submerged for at least 90 minutes. We reviewed literature for guidance2-4,6 and drew from institutional experience with a 2-year-old submerged in ice water for 40 minutes who received 101 minutes of CPR.3 The toddler recovered with no sequelae. For our current patient, the decision was made to resuscitate and rewarm the boy because of his young age and protective effects of ice water submersion. We reasoned that if meaningful neurologic function were not observed after rewarming, end-organ preservation on ECMO may allow family goodbyes and organ harvest for transplantation to give other sick children the gift of life.9 This important point should be considered by providers faced with the difficult decision to attempt resuscitation of a patient with asystolic hypothermia >90 minutes.

spuz··on Benchmarks in Leipzig
Ah you are right. I think I started reading the results of Stage 2 thinking it was Stage 1.
spuz··on Benchmarks in Leipzig
As well as measuring how many questions each model was able to answer correctly, I think it's equally important to measure how many questions each model answered incorrectly. After all, if you consider using them as a tool, you will need to have confidence that any answer they give is correct.

If you look at Table 3 you can see the difference in performance between for example GPT 5.5 and Opus 4.7 for each of the 20x 100 runs:

- GPT 5.5: 1389/2000 questions answered, of which 1043 were correct (75%)

- Opus: 1306/2000 questions answered, of which 294 were correct (22%)

So while you can claim that Opus solved 40% of the problems it still had a failure rate of 78%. That means if you chose this model to answer your homework question, there is a good chance you would fail.

Perhaps a more useful benchmark for future models is measuring how many of these types of questions they can answer in one shot. I.e. how confident can you be when using them for real world tasks.

spuz··on Sagrada Família Lego set
I normally love Lego's interpretation of various real architectural works but I don't believe there is enough detail here to really capture the unique style of Gaudi's design.
spuz··on Volkswagen blocks Home Assistant by requiring client assertion
What does client assertion mean here? I don't see any mention in the GitHub issue.
spuz··on OpenAI Adopts Google's SynthID Watermark for AI Images with Verification Tool
What are some of the millions of legitimate use cases that are harmed by having metadata added to generated images? It's funny you mention Photoshop because that software also adds metadata to jpeg images that it creates. Is the difference here that the SynthID is hidden and can't be removed?
spuz··on OpenAI Adopts Google's SynthID Watermark for AI Images with Verification Tool
There's a world where all the big platforms automatically flag AI images thanks to SnythID and other techniques. The more ubiquitous it is the easier it will be for them to make that a reality.
spuz··on Princeton mandates proctoring for in-person exams, upending 133 year precedent
When I had exams in the 90s we'd have to hand phones in at the start. If the phone was seen during the exam, your test would be forfeited. If the rules were that strict then, I can't imagine how they could be less strict now given how much more powerful phones are today.
spuz··on Princeton mandates proctoring for in-person exams, upending 133 year precedent
What is this honour council I've heard in a few comments? I thought Princeton was unique in having and honour system as opposed to strict academic integrity rules.
spuz··on Colorado grandma keeps getting pulled over due to database error
Yeah this is a problem even without technology. I believe the UK does not use the letter O in standard registration numbers so it cannot be confused with 0.
spuz··on Self-updating screenshots
The only problem with this idea I can forsee is that the application and therefore the screenshots can change but the documentation does not. For example, if the documentation says press "Options > Customize" but the application is updated so this becomes "Preferences > Advanced" then the screenshot will show the new text but the documentation will still show the old labels. This would be very confusing as it would be hard to correlate what is being shown on the screenshot with the text. If the user saw the old screenshot they could more easily identify that they were looking at an out of date documentation.

Having said that, have a process to automatically grab screenshots is going to make it significantly easier for a developer to update the docs so the motivation to keep the text up to date is going to be much higher.

Page 1 of 34Next →