HNHacker News
TopNewBestAskShowJobs

Zambyte

4,905 karma · joined December 1, 2020

website: https://robbyzambito.me

email: contact at robbyzambito.me

bksy: @zambyte.robbyzambito.me

zambyte.at.hn

submissionscomments
Zambyte··on You said no MCP
Exactly. It's confusing to talk about MCP vs API when "API" describes the whole class of possible values, of which "MCP" is one of them.
Zambyte··on You said no MCP
It's common when talking about web services, where the kind of API is generally unambiguous. The problem here is we're discussing different kinds of APIs (shell commands, MCP, REST / JSON RPC) and so calling one of those kinds of APIs "API" is very confusing.
Zambyte··on You said no MCP
... Are both of you using "API" as a synonym for "REST"?
Zambyte··on Calling the AI bluff: Adding "Do not guess" cut made-up claims from 71% to 20%
> I've yet to successfully convince one to tell me when it knows something.

This is a near daily experience for me when using a coding harness. I will ask it for some favts about the environment, and it will continuously explore the environment until it exhausts reasonable exploration, or it finds the facts.

Zambyte··on Feds Target AI Critics as "Foreign Agents"
As much as I hate appeal-to-authority based arguments, it's basically the only way I can see how to answer this without just totally rehashing the things he says in the video. He was an AI forecaster at OpenAI for several years. (I think 2019? to 2024). He quit on the grounds that OpenAI (as well as the rest of the industry) were / are throttling towards RSI irresponsibly, and even refused to sign an anti disparagement clause on exit at the risk of losing 2 million dollars (80% of his family's net worth) so he could talk freely about OpenAI. Now he runs a non-profit that does AI forecasting and advocacy.

https://www.aifutures.org/

https://ai-2027.com/

Zambyte··on Feds Target AI Critics as "Foreign Agents"
You said:

> I haven't seen one convincing model of AI existential risk. Can anyone offer one?

I apologize for answering. This is the only episode of this podcast I have ever watched, and I have watched it four times now. The person being interviewed knows what they're talking about. You would learn from them.

Zambyte··on Feds Target AI Critics as "Foreign Agents"
It's long, but I recommend watching this video[0]. I believe the person being interviewed has one of the most lucid understanding of the potential futures regarding if AI is allowed to continue or not.

[0] https://www.youtube.com/watch?v=_g4l7YkDQwA

Zambyte··on Tokens Too Cheap to Meter
Not really related to the central point, by but I couldn't help but get caught up by

> Generally, models intended to be run locally will be much smaller, such as Muse Glimmer or Qwen3 Coder.

That is such an interesting set of models to use as examples here. One being essentially obsolete on release a month ago, and the other being completely ancient in LLM time. I really wonder how they landed on those two.

Zambyte··on Jev in 25 Lines of Python
Are OpenJev, openjev-sglang, and OpenJev on DiffusionGemma using Qwen3-0.6B-Q8_0.gguf, or did you just want to emphasize the part that was unrelated to your previous response?
Zambyte··on Jev in 25 Lines of Python
I wonder if Qwen 3.0 0.6B q8 would have noticed

> note: this is a parody blog post

Zambyte··on Jev in 25 Lines of Python
> note: this is a parody blog post
Zambyte··on Claude Opus 5.5
> To me it clearly means "releasing frontier models at any pace less than as fast as possible".

It also seems to misimply that the "frontier" that they release is the same as the frontier behind closed doors. Who's to say they are not throttling full speed towards RSI privately while pacing their public releases?

Zambyte··on Claude Opus 5.5
Not the same person but... nothing. Haiku just hasn't been an interesting model for a long time. If you want cheap and fast, there are lots of options that are simultaneously cheaper, faster, and capable than Haiku.
Zambyte··on MiMo v2.6
American models are on the frontier of capability. Chinese models are on the frontier of efficiency. The problem for American labs is that Chinese models are more than capable enough for the vast majority of applications that people care about at this point, so efficiency is more interesting for people.
Zambyte··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
> especially with access to frontier models

... so not made in house.

Zambyte··on Grok 4.7
I wonder how much of this is due to reliance on Twitter data. Or even just RLHF from humans that have a preference for Twitter style information.
Zambyte··on Qwen-Image-2.1: Compact, efficient, and unified image creation
You could say the same thing about "license-washing" the model. It seems like you're just going through a guaranteed expensive process to have roughly the same risk as just using the model and potentially getting hit with legal fees.
Zambyte··on Qwen Image 2.1
Why would you even do that? Just... use it? There hasn't been any legal precedent on if models can even be copyright restricted. Labs just keep publishing license documents as if they matter.
Zambyte··on Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
How? I'm running 4bit with a q8 kv on a 24gb card, and I'm not able to get 100k out of it. I use a context size of 90k.
Zambyte··on Qwen 3.8 Omni Flash
How are you sure about that? Undercutting your competition (at cost even) to minimize their power is a common strategy.
Zambyte··on How to Write with an LLM
How about Anthropic? I've seen people say it's surprising that their blog isn't written with AI, but I think it's far more likely that it is written with AI; they are just good at reviewing and getting a nice prose out of their models.
Zambyte··on Rate limits on GitLab.com are changing
It's not, it just requires creativity. One example I can think of is limiting connectivity by distance between nodes. Something like Meshtastic seems unlikely to ever be commercialized in the same way that the Internet has been. Sure, you would lose some useful applications of the Internet, but you would gain other things.
Zambyte··on Ask HN: How many runs before you trust a coding agent's result?
Until it says it's done. One of the more surprising advancements for me in recent AI is that models don't just indefinitely nit pick issues like I would expect. If you have them write something, and then have it review it in a fresh context, and then have it edit based on the review, and then review it again, eventually they do say "it looks good, I would not recommend any changes" or something like that.
Zambyte··on Nvidia announces native GPU programming in Rust
... doesn't one of those imply the other?
Zambyte··on How much of F-Droid is LLM generated?
The other tool I wanted to throw out there is voxtype (plus wtype). I've been looking for a good, global, local dictation solution for wayland for a while now, and finally landed on this one. It's great for rambling details that pi+qwen can then convert into concrete openspec specifications.
Zambyte··on Claude Cowork and chat are now one Claude
LLMs are chaotic pure functions. Their input is usually randomized.
Zambyte··on Why I'm still bearish on LLMs after Navier-Stokes
It's at least important from philosophical perspective regarding whether LLMs have free will or not. It's hard to say whether or not a human thought or action is driven by free will, people have debated it for centuries. We can provide the same input to an LLM and get the same output, because they are deterministic. Surely it is obvious that a pure function does not have free will, no matter how expensive it is to compute.
Zambyte··on Why I'm still bearish on LLMs after Navier-Stokes
Traditional non-chaotic systems*

LLMs are deterministic. They are chaotic, which people confuse for non-deterministic.

Zambyte··on How much of F-Droid is LLM generated?
No problem. I think it's less of people withholding information to have an advantage, and more people still not having settled on a workflow they like. Herdr is the most recent addition in my workflow as of only a few days ago, but it directly solves problems I have been having (juggling tons of terminals, even with my tiling wm has been a little unwieldy). The rest I've pretty much settled into for a while now.
Zambyte··on How much of F-Droid is LLM generated?
Up until the last couple of months, I have treated LLMs as a supercharged stackoverflow. I would ask it questions on how to do something in a general sense, and then adapt the answer to my use case.

Now, my entire programming flow does not even include an editor. The tools I use are: pi.dev to write and implement openspec specifications, herdr to manage many pi instances, and ollama to run qwen 3.8 27b on my single 7900 XTX.

Writing good specifications is the key detail here. I will often iterate on a spec for hours until I am happy with it all of the details. Once I am happy with the spec, I can be quite confident that when I tell pi to apply the spec, the changes that I want will be done, and done how I want them, when I come back to check when it reports itself as done.

The landscale is fundamentally different from what it was. Feel free to ignore it, but you can absolutely generate high quality code if you know what you're doing.

Page 1 of 34Next →