HNHacker News
TopNewBestAskShowJobs

m11a

656 karma · joined November 2, 2018

twitter: https://x.com/themandeepc
submissionscomments
m11a··on Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
It sounds like Jev is not a generative model. If true, I don’t understand these clones which fine-tune a generative model, which appears to not be what Jev is doing.
m11a··on Two-tier encryption in the UK
This news is about the UK. Across the channel, the EU has effectively outlawed anonymous currency. I believe the UK FCA also has such rules in force. So your post is rather moot.
m11a··on Muse – Meta’s personal AI agent
Yes it's definitely not SOTA, and is worse than all the chatbots I use. But it's correct frequently enough that the convenience + speed does it for me.
m11a··on On the Navier–Stokes Millennium Prize Problem
Although there are companies trying to work around that too, from PhysicsX to some of the world model co’s.
m11a··on Muse – Meta’s personal AI agent
tbh I also use google AI tools frequently. It’s convenient (search bar is top of my window), and extremely fast at generation. For quick Qs, opening a chat bot and changing the default reasoning from my coding settings, switching from Work to Chat, etc etc, is too much of a faff.
m11a··on GPT-6 Astra
Many VCs also dislike these examples, I believe. I'm doubtful this is what they're being pitched.

As for public releases: I wonder if it's because these examples are easy to relate to. Many websites are just a long tail of industry or use-case specific stuff. What's valuable to me probably means nothing to you. This is unlikely to resonate with people-wit-large (and LLMs are marketed broadly) or requires the reader to think (and marketing that requires thinking is bad these days).

Second, it's arguably a good litmus test. If it still can't do the worn out examples of plane tickets and shopping, which would be a good assumption since we've been demo'd these use-cases for 2 years at this point, then ...

m11a··on GPT-6 Astra
I recall them saying they use models to write CUDA kernels and whatnot. Makes sense, and unsurprising that models are good at writing code.

But I think calling this “automating AI research” is misleading. I’m not sure there’s evidence yet that they do creative research work. Even in mathematics, but they are finding counter-examples by intelligent brute-forcing. Not to downplay the results, as they are incredible, but this is one very specific kind of proof and not the most creative type, which arguably requires generalisation.

m11a··on Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?
I did, at least until Fable 5.1. Grok’s models are excellent, amazing price-performance and speed too.
m11a··on Stacked PRs are now live on GitHub
In essence. But if you have a chain PR1 > PR2 > PR3, and PR1 gets merged, all the others (ime) seem to not cleanly rebase on main. They end up with conflicts that require manual fixing. I've not really figured out why, tbh. It'd also be nice to see a coherent link in the UI between PR1, PR2, PR3.
m11a··on Stacked PRs are now live on GitHub
GitHub's PR workflow doesn't nicely support being able to review individual commits, realise which comments are associated with which commits, etc. Or shipping individual commits to main, while working on some others (unless you allow cherry-picking and direct push to main). Or amending a certain commit with respect to feedback and seeing a diff from the previous patch of that commit to the next.

If GitHub's unit of change were a diff, and not a branch, then that would work pretty well.

m11a··on Stacked PRs are now live on GitHub
I think it's telling how long it took GitHub to release a v1 of this feature. Folks have wanted this for a long time. Graphite came along and did it years ago (and I'm sure they pondered whether GitHub would do this).

And the v1 is also a bit... basic, and buggy. And I'm surprised there's not clear documentation for agents (given using GitHub stacked PRs CLI won't be in models' training data yet).

It does feel like GitHub hasn't been great at shipping new features for a few years now. Nonetheless, I'm glad to see this rolling out. Once polished, it's going to be exciting to use.

m11a··on Open-weight AI is having its Kubernetes moment
I presume such US legislation isn't going to try claim worldwide jurisdiction to block all persons worldwide from using Chinese models. In which case, the French guy wouldn't be violating American law.

As for the American company, it's pretty difficult to check the provedance of open weights. It's even difficult to check the provedance of open source code, because chains of attribution aren't always clear. I posted elsewhere that Anthropic's MCP Python SDK is a fork of an open source project with the attribution removed. We saw the same with Cursor's Composer model, which didn't attribute its Chinese base. It's very hard to claim an American company should be liable for using a purportedly European model with attribution removed.

m11a··on Open-weight AI is having its Kubernetes moment
Because it's not losing money on each token? Aside from most global people using American inference providers to run the models, I suspect the cloud inference products of the Chinese labs are profitable, at least on the inference costs (ie: not including model training, salaries, etc).
m11a··on Open-weight AI is having its Kubernetes moment
Probably just the US. But the US could do what EU has done with e.g. GDPR, Digital Services Act, USB-C regs, where they force any company trading in their region to follow those regulations for domestic customers.

And basically any AI company has to sell to US companies or consumers. That'd probs be sufficient to force them to use US models.

m11a··on Open-weight AI is having its Kubernetes moment
Step 1: Chinese company publishes open weights on HF

Step 2: European company distills or just adjusts the model slightly, and publishes its model on HF

Step 3: American company uses model from step 2. Has to testify under oath where they got it from. "We got it from these French guys"

m11a··on Ben Bernanke Joins Anthropic Oversight Trust
How long did it take from the first DBMS to get to Postgres? The first OS to get to Linux? The first compiler to get to LLVM? For Postgres and Linux and LLVM to become mature enough to hold the revered reputation they have now?

The jury's still out on AI, but coding agents have only really worked for about 6 months now. It's not exactly a fair statement to make. Obviously good things take time and thought. And understanding the full implications of technological advancement also takes time and thought.

m11a··on Ben Bernanke Joins Anthropic Oversight Trust
I had conversations with Ben a few years ago about his economics research (incidentally related to some research I was doing at the time). I wouldn't have guessed he was close to 70. He's a sharp guy.

(I also think his economics research is excellent, and the comments elsewhere about his track record aren't particularly fair.)

m11a··on Cloudflare Meerkat - Globally distributed consensus
Right. But given that the entire point revolves around QuePaxa, it's strange to see no discussion on it. If that weren't the point, the article would be "Why Cloudflare implemented and deployed Paxos".

Which would also be a good read, but this article also isn't that. It doesn't discuss their experience deploying the protocol, aside from the following statement:

> Meerkat is not deployed to production, but we have run multiple proofs-of-concept with up to 50 replicas distributed around the world, to great success. Leaders in our proof-of-concept clusters constantly fail, and the cluster keeps operating with no increase in error-rate.

I think it would've been more interesting to read why Cloudflare chose the specific algorithm they did, see an example of a pathological but common situation Cloudflare sees at their scale that makes other protocols unsuitable for them, therefore they made X choice and this led to Y gains in production (or on dummy workloads, or whatever). As it stands, there's nothing here actually specific to Cloudflare's workload or deployment. It doesn't even state their use-case beyond "small pieces of control plane state (e.g., leadership for replicated databases)"

m11a··on Cloudflare Meerkat - Globally distributed consensus
This article is a bit hard for me to grasp the main ideas of because, given Cloudflare's requirements (e.g. no strong leaders), it immediately seems like they should be comparing to leaderless protocols like Paxos-class algorithms. Comparing to Raft and saying it's better because Meerkat is leaderless is confusing, because Raft is an adjustment to Paxos to specifically have strong leaders. So I'm 3/4 the way into the article and I don't see what's unique here.

I think the unique idea here is supposed to be QuePaxa's idea of avoiding timeouts for ensuring liveness. The actual discussion of QuePaxa is limited to one paragraph at the end, and tbh only a couple sentences of that paragraph.

I feel like the article could've been titled "Consensus protocols and linearizability: a brief explainer", or "Paxos vs Raft", or similar. It just doesn't feel like it communicates what it claims to communicate, and is a bit confused on who its audience is, just IMO.

m11a··on Knowledge Should Not Be Gated
It depends on the model provider. OpenAI's is very limited and precisely written.

Plus, open-source models hosted on SaaS inference providers tend to come with a strong ZDR agreement too.

m11a··on Knowledge Should Not Be Gated
Most corporations likely have zero data retention agreements with LLM providers, at least for API usage.

(Sure, you could be sceptical on whether the LLM provider is upholding that, but I personally do trust them. The trust betrayal if ZDR wasn't actually ZDR would be too great and commercially damaging for them to lie.)

m11a··on Meta is adding rate limits and soft paywall to smart glasses
Says a lot about just how much money they’re raking in from core business.
m11a··on Claude Code is steganographically marking requests
I mean, you’d resort to an obfuscated approach if you thought the ‘malicious’ users would remove your direct telemetry. The other users could be perfectly happy about it, but if you announced the change, obviously the malicious users would hear about it too and disable it.

This isn’t a comment on whether I agree with the change. Just that your analogies aren’t applicable here.

m11a··on U.S. government will decide who gets to use GPT-5.6
I mean, after the US just signed an export control ordering Fable access blocked to non-US users (including European nationals), I doubt European and "US-aligned markets" are eager to ban Chinese models against their own interests.
m11a··on How to convert between wealth and income tax
California already has very high taxes. I think marginal tax rates are higher in California than for UK tax residents, certainly for CGT, and roughly similar for income tax.

I'd say the fact that California remains the epicenter of tech despite its high taxes suggests concentration of talent matters far more than tax rates.

m11a··on Incident Report: Railway Blocked by Google Cloud [resolved]
If Railway's account was suspended due to an error, not a TOS violation, I doubt they'll pull such a card. If Google were sued, such a blatant lie would be found in discovery pretty quickly and I doubt a court would look upon it favourably. They'd also be starting a public PR battle that they'd almost certainly lose (assuming there was indeed no real TOS violation) which would make them look completely untrustworthy. So I doubt Google will do that.

bastawhiz is probably right in that Google will offer some credits and an apology, and Railway will reduce their dependence (probably), rather than a lawsuit. And I doubt Railway wants to be on the bad side of one of the few big cloud providers. But I'd be surprised if Railway didn't have a good argument to make for compensation or a lawsuit.

m11a··on Curl and Mythos
I’ve learned that in AI, until you can get your hands on something and try it yourself, assume all claims are false.
m11a··on Cal.com is going closed source
I'm not sure security through obscurity is a great practice?

Not to mention, I presume the core bits of Cal.com's source code are already in place and aren't going to change significantly?

Like, this feels like a business decision and not a security decision

m11a··on Welcome to FastMCP
Cloudflare's Code Mode is conceptually the same as Anthropic's Code Mode (https://www.anthropic.com/engineering/code-execution-with-mc...), or the various open source implementations that predate and postdate those blog posts.

tbh, that companies tried to make something proprietary of this concept is probably why its adoption has been weak and why we have "MCP vs CLI/Skills/etc" debates in the first place. In contrast, CLI tools only require a general a bash shell (potentially in a sandbox environment), which is very standardised.

m11a··on Welcome to FastMCP
If I recall correctly, the ‘official’ Python one is a fork of FastMCP v1 (which then removed the attribution, arguably in violation of the original software’s license)
Page 1 of 8Next →