HNHacker News
TopNewBestAskShowJobs

regularfry

9,430 karma · joined January 6, 2009

submissionscomments
regularfry··on Mistral OCR 4.1
Lots of tech money is demonstrably wrong, at least in the general case. Microsoft's whole business model is to be the second mover. Patents exist because first mover advantage alone isn't enough to hold onto a win.

If AI is different, the onus is on AI folk to explain why, not the reverse.

The frontier labs are racing as hard as they can right now because they know as well as anyone else that they're at best 6 months, probably closer to 3 months, ahead of the generic commodity AI market. The majority of the tech money currently boosting them is a bet that either they can stay that far ahead without tripping, or that a moat can be engineered. It's not at all clear that it can - or, if it can, that it will before the money runs out.

regularfry··on Pixel Watch 5
I got a pixel 4 after some surgery basically for the fall detection. Otherwise it's mainly doing what the Fitbit I had 8(?) years ago did for me, with added GPS. Pulse, steps.

As an ADHD-haver it's good to get notifications to the watch so that I can avoid unlocking the phone and risk falling down a rabbit-hole. I think that's under-sold.

What I want is a decent (by which I mean SIMPLE) pomodoro app for it, and I'll probably have to write that myself. Does that make me a power user? I was surprised that finding such a thing would be difficult.

I tend not to use it for payment because the ergonomics are terrible. Someone needs to tell that team that having the face on the inside of your wrist is entirely valid (especially if you don't want to be whacking your shiny watch face into every passing door handle), but that makes you do contortions if you want to get it close to a card reader. Not only that but the way screen locking works is ungainly enough that the fact you need to use it more with the face towards your body makes it a constant annoyance. I've unknowingly reorganised my widget order far too often because of this.

regularfry··on Why Target Common Lisp for Code Generation?
I found that explicit instructions to keep nesting depth below a limit helps limit the damage. That limit was 5 in my case, which can be inconveniently tight in some situations.
regularfry··on Why Target Common Lisp for Code Generation?
It's more opportunities for a screw-up to lead to a vulnerability in ways the tooling won't catch by default, in a system where P(screw-up) > 0.
regularfry··on Why Target Common Lisp for Code Generation?
I have a hunch (based on using the Kimi models to write some clojure) that the article's AST point is exactly wrong. I had to spend a lot of time cleaning up when it miscounted closing parens, which implies that while the LLM may be operating on the AST, mapping to and from the token stream is harder, not easier, when the individual tokens carry less information.

If the author is finding that it works well, I suspect there's something else (code or comment style, maybe) that's compensating for it which didn't seem worth mentioning.

regularfry··on DARPA heavy lift challenge ends with winner at a 3.84:1 payload to weight ratio
The article makes reference to this being expected from the physics, but I don't think it's particularly intuitive why it should be the case. Unless it's as simple as smaller swept areas needing higher tip speeds for the same lift? That would make the energy loss to drag worse for multirotors.
regularfry··on Illinois just passed a law that puts Linux on the hook for age verification
Making a stupid idea more convenient does not make it less stupid.
regularfry··on SQLite Critical CVEs or LLM Slop?
Any org large enough to have separated the people responsible for the security exposure of the organisation from the developers with familiarity of what's deployed is likely to have done exactly this.

The thing you have to remember is that CVEs can be a) scanned for without exerting mental effort, and b) counted.

regularfry··on Prevent cognitive debt by manually retyping LLM-generated code
The bottleneck is very rarely the typing.
regularfry··on Qwen3.8-Max: A New Bar for Coding and Cowork
Oh, I'm sure. If it wasn't a bit of a tricky needle to thread we'd have seen them start to do it already.
regularfry··on Qwen3.8-Max: A New Bar for Coding and Cowork
Cowork is to work as coding is to... ding?
regularfry··on Qwen3.8-Max: A New Bar for Coding and Cowork
If you're literally using it as code autocomplete that's not functionally true. But it's still neither well specified nor a good idea.
regularfry··on Qwen3.8-Max: A New Bar for Coding and Cowork
For certain specific uses. As a consumer I can still buy Huawei phones. As a business I can still buy Huawei routers, if I feel so inclined.
regularfry··on Qwen3.8-Max: A New Bar for Coding and Cowork
With their current API approach they're essentially a commodity. They need to start moving parts of the harness behind the API, otherwise they'll remain a commodity.

Recursive self-improvement changes the parameters a bit, especially for the market-leaders, and it's the one thing that makes me wonder if they'll be able to extend their lead faster than the smaller labs can keep up, but it's an option available to everyone.

regularfry··on DeepSeek-V4-Flash Update
It's good but (at least on openrouter) it's got an annoyingly tight output token limit. So if it does get stuck in a reasoning pit, it won't work its way out of it in time.

It's replaced the Kimi models for me though.

regularfry··on Some thoughts about Anthropic's new cryptanalysis results
For me it's more an expression of how much mileage you can get out of "glorified", how many tasks devolve down to "if you model language accurately enough, look what drops out" because it turns out that to model language you need to model how the world works. I don't use it to minimise the capabilities at all.
regularfry··on Some thoughts about Anthropic's new cryptanalysis results
The wikipedia definition is:

> ...a hypothetical type of artificial intelligence that matches or surpasses human capabilities across virtually all cognitive tasks.

You can argue that paperclip maximising is an inevitable consequence of that (and the huggingface breach is interesting from that point of view) but it's not fundamental to the definition.

The question then is what "surpasses human capabilities" means and we're there in some niches but not all, and not across many models.

regularfry··on Keychron announces first open-source firmware for gaming mice
I've just grabbed myself a Nape Pro trackball. It's really quite neat. ZMK firmware (which was actually the main thing I wanted) and a configurable positioning angle. I'm still getting used to it but so far it looks like a decent addition to my wireless toolkit.

I would prefer a bigger trackball as a daily driver, but it's an interesting device otherwise. It feels like they could reasonably do a mini version, too, just by chopping the non-main buttons in half. Definitely going to crack it open and see what the innards look like.

regularfry··on Handbook.md shows that long policy documents do not reliably govern agents
Control vectors for the win? It feels like the way to fix this is to pull out a control vector immediately after processing the handbook and use that to steer later inference. That should stop the drift over long distances, but it doesn't guard against the handbook itself being too big.
regularfry··on Show HN: Formally verified 3D CSG: Trust 93 lines spec, not 1000 lines AI code
CSG on meshes isn't really where FreeCAD struggles or focuses. That's more BRep, which is a harder problem than CSG.

Blender used to struggle with boolean operations, though. Seems to have got much better some time in the last couple of years, but without knowing any of the backstory I don't know if they've bodged it into shape or would benefit from something proven.

regularfry··on Vehicle Motion Cues
When did you try it? I got a bit confused by your comment because there's no obvious reason why the accelerometers in an android phone would have a problem reporting forward/backward acceleration any less accurately than lateral. So I installed it and the 'Dots' mode definitely responds to acceleration normal to the screen, but that mode is marked as 'New' so might not have been present when you last looked.
regularfry··on What Even Are Microservices?
If you're needing to share code for domain objects between microservices something has gone badly wrong.
regularfry··on What Even Are Microservices?
The very first time I came across the term, it was used to refer to something so small that you'd comfortably rewrite it rather than fix it if it was wrong.
regularfry··on What Even Are Microservices?
It's not just the modular organisation that you've got to get right though, although that's definitely a big thing. Deployment independence matters too: if I've got a perfectly modular monolith but I have to coordinate with everyone else who lives in that monolith to get my module into production, it's leaving half the gains on the table.
regularfry··on What Even Are Microservices?
One team owning many services is fine, as long as they can deploy them independently of any other team. Less critical (but still useful) is being able to deploy their own services independently of each other.
regularfry··on Claude Opus 5
This isn't mimicking what a human would do, because it is an exceedingly rare human who would write their own renderer for a task needing a CAD file generating. The human who is capable of doing both is rare, let alone the human who can do it in any sort of reasonable time. Saying "but it's in the training data" is a cop-out: it's in Google too, would I do it? No. I would not.
regularfry··on Why Software Factories Fail (or: harness engineering is not enough)
The whole point is that if you don't maintain high quality, any software will degrade to the point where the factory breaks. "Human-quality" here isn't quite the right measure because human software factories also bog down in exactly this way.
regularfry··on Laguna S 2.1
I got usable token rates (10-20tps from memory, so marginal) with the Qwen A10B a while back, well before all the new speculative speedups landed in llama.cpp. There's an unmerged branch which allegedly supports this, but I don't know how well yet. Looks worth investigating but I might give it a few days to see what bugs get shaken loose.
regularfry··on Laguna S 2.1
What's your hardware? I've got 64GB RAM and a 4090 here, wondering if it's worth a play.
regularfry··on Gemini last models: temperature, top_p, and top_k are deprecated and ignored
Couple of other more businessy reasons:

- SynthID hides the watermark in the sampling RNG. No randomness -> no watermark.

- If you want to distil on the model outputs, you want temp=0 outputs. No temp=0 -> worse distillation.

← PreviousPage 3 of 34Next →