HNHacker News
TopNewBestAskShowJobs

porridgeraisin

2,745 karma · joined March 26, 2023

Systems Engg, Reinforcement Learning, Bespoke RISC-V, Linux, Webtech enjoyer, Music, Chess, Cricket, Football, Swimming, India.

Currently academia

submissionscomments
porridgeraisin··on LRU is harder to beat than the KV-cache papers suggest
https://x.com/badlogicgames/status/2098868942685012405
porridgeraisin··on Will There Be a 7G?
There was full team that found out that did the test that 26ghz doesn't work IRL. It didn't even work across a street corner. But they were ignored in favour of some well, career pushers.

In india, jio is using 26ghz for wifi now instead. It's a service called airfiber. This don't even use any 5G. It's a weird custom hybrid solution where all the frames are wifi. Density here means that it's feasible to put the receiver on the rooftop of apartment buildings wherever you can get LoS. So in some sense it's been salvaged and those bands are actually being used. It's a fairly popular service.

URLL failed because it's main immediate usecase was supposed to be vehicle to vehicle comms. And auto didn't want telcos middlemanning there. But there is still hope here.

porridgeraisin··on Will There Be a 7G?
Yeah. In india jio was SA from the start. Airtel began NSA then I think now they have mostly moved to SA.
porridgeraisin··on ElevenLabs Music v2.5
A lot of streaming platform listen minutes has always been this type of filler music. Until now, you could not really disambiguate intentional listening and filler listening. Now you can and the filler listeners wont even notice[1]. A huge chunk of the listen minutes whose value flowed to the artists (from a market perspective, technically this is inefficient), is now not going to go to them. Spotify will prompt optimise based on listening stats and perfect it. TBH until now, many artists were doing stream minutes optimisation of their pop music as well, this is kind of the logical conclusion of that. Human artists will have to stick to intentionally listened music, which is going to be a difficult competitive market. And most listeners actually don't care about this type of music to the same extent, measured in streaming minutes, so the market will be smaller as well. Popularity will end up mattering quite a bit more than it does today I suppose.

Example: This one is easily a replacement for people who listen to linkin park type music passively: https://elevenmusic.io/tracks/6aa1ca95740e34b282c33b0a Well except the vocal only pre-drop part, that was horrendous lyrical flow from the clanker.

[1] https://www.reddit.com/r/EDM/comments/1mq3s5p/just_noticed_t...

porridgeraisin··on So you want to use OpenRouter?
I think you're focusing only on the general coding agent aspect of LLMs.

> If openrouter is only for people who "make their own benchmark" in a mission to roll something that's usually made by scientists with large budgets, it should say so and thus fade into deserved obscurity.

That is the way LLMs have to be used for highest reliability. While the term stochastic parrot has been co-opted by unreasonable LLM skeptics, that is indeed what LLMs are. You have to have grounded evals that check outcomes if you want to use them reliably - or a human in the loop works too.

The more general your family of tasks, the less likely you can make automated evals. So for "general" coding agents, you need a human in the loop that can verify and it's not that easy to write an eval.

But if you have specific tasks, then you can spend the time to make a eval, and then you can optimise the way you use the LLM and get extremely good success rates. It's not like it's black magic. Nor does it need large budgets.

> that is indeed the basis of this massive corporations entire business plan.

No. Individual developers using codex (for extremely underspecified general engineering) needs human in the loop, is not amenable to evals but is only a fraction of all LLM usecases.

porridgeraisin··on So you want to use OpenRouter?
You'd be surprised. Not enough teams still have their own evals. India[1] and the US[2] is my experience. To get many of them to understand the benefit of putting a couple of people to do data labelling and write a couple of verifiers for just a few days every few months was so difficult. Many folks have understood it all wrong and made allotments like "big model for this task" "small model for this task". Small was sometimes parameters, sometimes brand version number, or sometimes because it has "mini" in its name.

Proper LLM adoption beyond fucking around with claude code or github copilot is so low and is only going to go up as people figure it out. I also think the new cloud agents thing might accelerate adoption among these companies, since it's a bit more plug and play. But they are more likely to be stingy about it, so I'm not sure about the high margin expectations of certain model families.

[1] not witch, but BFSI. Surprisingly parts of witch companies have it down to science already, they embedded openai or anthropic or startups like devrev more than a year back.

[2] again not high fly SF companies, BFSI.

porridgeraisin··on Recreating Minecraft Is Not a Benchmark
There is also the aspect of model generations being robust under a particular ctx mgmt strategy/tailored harness.

Only since this year are most models robust in this way. In many open models, the problem is that successive generations are not that robust yet. i.e if you have a working setup, next generation of the same model family will need way too much reworking. So it's difficult to upgrade the model. Gemini is stellar in this respect (probably because, and I suspect, their flash models are distillations of the same larger model). OpenAI/Anthropic are too, but higher cost of newly released models is quite the blocker. Deepseek is especially hard to upgrade. Kimi doesnt target this segment, so people that can post train it do so. Qwen and llama are the stellar open ones, incredibly stable, although llama has stopped receiving upgrades in a while. It is still used a lot though.

porridgeraisin··on DeepSeek v4.1 Flash
There are thousands of such techniques across different parts of the system. In ML, there are way too many ideas, and lots of people knowingly and unknowingly restate the same ideas. It's a new field, so even common language is not there. For an extreme example, so many improvements are restatements of 1960 signal processing techniques - obviously very few ML people have done DSP beyond the undergrad course. It is also highly empirical and many parts of deep learning (heck, even non-NN ML) are not understood yet.

Thus, the reality is that most of these ideas become polished only when its actually deployed and it has to work outside of a PoC. Since LLMs are a high capex product, only very few people actually make non-PoCs. Deepseek is in the business of low cost, fast inference. So they are the ones actually polishing these efficiency-ish ideas and combining many of them (this one, then engram which is based on multiple previous ideas including google brain's ngrammer) to make a coherent system. Openai and anthropic's systems will also involve a polished combination of multiple ideas for each of their systems - Luna is likely a combination of a few efficiency-ish ideas. Shame they won't publish though.

As for microsoft, they don't really sell models, they sell azure. So there is no reason for them to do the high capex scale out of these types of bags of techniques. In a sense, it did benefit them, others developed the model and now many US customers can serve DS4.1 Flash on Azure datacenters.

If it is not clear, I am not understating anything. Combining these rough ideas and making them work actually involves real novel ideas on top and is what is much more difficult than the academic results that were built upon. This also does not mean that the academic results are useless, they are what give us useful priors at all in what is a highly empirical field.

porridgeraisin··on I trained a small transformer in 1.5hrs and it beats many LLMs
Yep. a multiplicative LSTM to be exact.
porridgeraisin··on DeepSeek v4.1 Flash
This is adapted from Microsoft research's YOCO. It was known for a while(2024!).

Yes, credit to Deepseek for actually scaling it up and releasing a frontier flash LLM.

Edit: the rest of this thread has become a US China infowar theory culture war. I am not of either of these countries and the above comment isnt meant to implicitly support either "side".

porridgeraisin··on What do Visa and Mastercard do? An intro to card networks
In india we use UPI. But people still use cards enough to be in decent enough terms with them for when you want to make a risky purchase and chargeback. In a sense, it's simply insurance. You pay extra 2% everywhere so you can dispute a txn at any point later.

Personally I make very few risky + expensive purchases, so my CC usage is non existent. I am comfortable enough to not really care if a random shady hobby electronics website fleeces me 500rs.

Another use is that sometimes you get CC offers on Amazon: "use $BANK $TIER CC to get extra 7k off" which are useful enough to justify paying extra everywhere else if you do your big shopping though Amazon festival deals. E.g you can get a 55k iphone for 45k.

porridgeraisin··on What do Visa and Mastercard do? An intro to card networks
This. My extended familys shop really benefitted from moving away from cash. Now the only source of cash is people who are paying from their black money stashes.
porridgeraisin··on Tao: Open math problems being non-renewably mined by AI
I phrased it badly just out of bed.

I meant what you're saying. That it's OK if it's slop if it serves a business function.

Edited

porridgeraisin··on Tao: Open math problems being non-renewably mined by AI
Because the problem has almost no value unto itself. The clay statement of navier stokes is not relevant to how CFD is done in practice.

It's about what is non verifiable versus verifiable. The same way it produces "slop" code (which, if you give it test cases, will be 100% correct), it also produces "slop" math.

Code that serves a business function, it's ok if its slop. Math that serves directly a business function also can be slop.

But most open problems are not directly for a particular usecase. People agree widely to attack it due to the perceived possibility of encountering useful mathematical objects along the way, that will then expand the world's mathematical toolset. This is not something that you can easily express in a verifier, and is thus something that is hard to force an LLM system to do.

You are right in that understanding it retrospectively is possible, but that is not going to be as useful as the desired "elegant" objects that expand and unify mathematics. You can't represent these concepts in verifiers.

Again, if you let AI rip at something like say "beat shannon capacity" and suppose it comes up with MIMO as paulraj did, great! It's useful and you can retrospectively understand it, say by expanding shannon to multiple dimensions, as foschini and telatar did. But most math problems are not in that category.

The question then is, if AI is really good at this type of math, how much of the existing mathematical community+process is necessary? I think it will still be necessary, just maybe in fewer cases. Wherever the primary purpose of the math is in a domain and that domain has a verifiable target, we can directly optimise it to that verifiable target in-domain rather than reach for the mathematical community. How well will this work? We'll see. It's not clear if it's even possible to represent most problems this way.

porridgeraisin··on Muse – Meta’s personal AI agent
It was behind the heavy 100$/mo+ plans. Clearly meant for enterprise and not consumer.
porridgeraisin··on Muse – Meta’s personal AI agent
Yep. But I suppose these are early forms of the product and who distributes the next, high PMF product the best is what matters. So I wouldn't write off others.
porridgeraisin··on Muse – Meta’s personal AI agent
Instinct exists with the same features. It is quite good. I was thinking they were gonna get acquired by meta 100%, but looks like meta has a competitor.
porridgeraisin··on Navier-Stokes – Tristan Buckmaster [pdf]
I agree on that, my post was not meant to be opposing this.
porridgeraisin··on Navier-Stokes – Tristan Buckmaster [pdf]
The for-case for this type of method is that this is economies of scale for mathematics.

We are basically mass manufacturing math. Just like you have just 100 designers for a product selling millions of units, you will now need 100 mathematicians to make millions of advancement. Yes you have factory workers, but if we are being realistic they have negative leverage in the world and the analogue of that is not something most of today's mathematicians would want to do. They would want to be in the 100.

Like Tao says, each advancement is now significantly less useful since it yields fewer usable objects. However, we will get many many advancements. Is the tower made with many worse bricks better or worse than the tower made with a few amazing bricks? Depends on the tower. And time will tell.

For some fields of math and some of it's usecases, economies of scale will be positive ROI overall. In others it won't. But we will know which is which only after it's been fully scaled up, which will take 10-15y in my estimate.

Some feel that in the majority of usecases it is negative ROI, some feel the other way, but that opinion is for practicing mathematicians like Tao to hold. Also, some opinions on either side are held in the context of a particular field or practice, and should not be interpreted generally.

porridgeraisin··on Working on Economics with Fable 5
It is not pseudoscience. Have you read about how it works?
porridgeraisin··on The Dataflow Model Revisited
There is a reason imperative programming is so common. It is more amenable to poorly designed, under specified, iterative development. Most real world software is in that category naturally. If you're meticulously designing and engineering it really well from the start, sure functional languages represent it well without leaving much room for misinterpretation and thus bugs. But no one does that.

Marginal cost of adding a feature has to be proportional to the revenue made by that feature. Then In Java or go you just add a ugly special case to appease the large customer and ignore the small ones' emails. Bugs getting shunted around instead of truly fixed at the root is also totally OK as long as they are not in the major revenue/cost centers of the product. No one has time to replace these piles of hacks and eventually you would have given up most of your languages benefits and your types now mean nothing there is probably a hundred flags making it a union effectively.

Rust is one language though where hacking around goes a long way without breaking too many of the guarantees, although it's not perfect, from my experience at a company where services spanned java go rust and ruby.

porridgeraisin··on Speculative Decoding in vLLM on AMD GPUs
> it must perform it's normal autoregressive decoding to know what is the correct token in order to have something to compare with

Correct except for the word "autoregressive". When you have to verify a sequence of tokens (which were autoregressively generated by the cheap model), you can do each token in parallel. This amortizes the cost of loading the weights from vram to the processors (the primary cost in LLM serving) across those tokens. Cost here is wall clock time, as well as power.

The autoregressive decoding that generates this batch of tokens is delegated to the cheaper model where the cost of loading the weights is lower and so not amortizing it is fine.

Verification means, how close is each token in this sequence to the one I would have output. You keep the longest prefix that is close enough for your liking.

porridgeraisin··on Is There I/O After Death? What Happens to Io_uring When a Process Dies
TBH, your username gives that part away atleast. My mind went to slavic when I saw the "ov" - it might be factually wrong, but thats waht happened.
porridgeraisin··on The moral panic over data centres is foolish
The noise is the main issue. I'm not american, but one of my relatives works in bernie sanders's office, and one of my friends (30ish) is involved somehow in the democratic party not sure how exactly. And what I hear from them is that the issue is mostly the noise only. The water aspect was there but that was mostly disruption during construction.

They are powering it with turbines right, so if its next to your house it sounds like you're permanently in on the airport tarmac. Not great. I saw a video (not on the internet, but from them). It was pretty annoying. But it makes sense, there is no way a power delivery system in a developed country adapts so fast to such intense energy needs. Here in india, many 250Mw datacenters are put in places where a proper grid connection is itself coming for the first time, so the same problem is not there. And frankly there are way more noisy industries in poorly zoned mixed-industrial-residential areas here for generator noise to even matter.

porridgeraisin··on Formalizing Fermat's Last Theorem
Literally many previous instance used one. Right from alphaevolve onwards.

I'm not some "LLM is just a next token predictor guy" (GP seems to have a thing against LLMs), but to use LLMs properly you genuinely do need a grounded verifier and a planner. Coding harnesses for example are exactly that.

For some plans, you can AR generate the search tree and that's what subagents being planned around by high level (LLM)agents and such are. Coding agents even with subagents are imperfect even on verifiable tasks only because of that. If you can put a human to simply guide it, it becomes a full system. This is what we all do today whenever we use codex. It's not something that is "never done before".

I also don't subscribe to the purist view which is taken by GP. I prefer to think in terms of concentration inequalities. P(failure rate > r) < epsilon. You get different levels of autonomy for different values of r for the planner and verifier each. If you have a good planner and a good verifier, r is very very small and it's super useful. Autonomy at a given r comes from how much of the planner and how much of the verifier is automated at that r. All levels of autonomy are economically useful. Many values of r are economically useful.

In this case of FLT, the verification was entirely automated using lean, and it is correct upto lean compiler bugs (so a very small r). The planner was essentially a maintained graph (afaik. Prove2me doesn't use A* or any heuristic/evolutionary methods to limit or prune the frontier), AND importantly - I'm not seeing anyone on HN mention this - some human nudges, literally, which nodes to open.

The way to make AI systems more useful is to build great verifiers and great planners, which is what many companies and startups are doing. LLMs are already really really good proposers due to excellent generalization (to be pedantic, multiple stacked specialisations), especially MoE models, making them amenable to proposing at every point in a vast search tree without any adaptation.

Yes, it is possible to do complex tasks purely AR, so long as you can AR simulate the search, which in the case of LLMs corresponds to verbalising the search tree[4]. This is trivially true. Can this be useful? Yes. Can a millenium prize problem be solved purely AR? Sure. It's a hard problem for humans, there is no reason it has to be difficult to reach in the conditional distributions of every future LLM. In the trivial limit, an LLM trained on the solution 100% you can sample it out. An LLM 2 generations behind that may have it at p=0.001, entirely reachable given a planner, but probably not AR. An LLM 1 generation behind may have it at p=0.05, plausibly reachable purely AR.

But the key question is: is `r` smaller or larger if you have a planner versus not? The answer there is obvious. Second, if you have a threshold `r` that decides usefulness, is the set of things you can autonomously do under that threshold higher with planners and verifiers? Again the answer is an obvious yes.

Copy pasting code from chatgpt repeatedly is worse than using a coding harness where it gets grounded feedback, LLM weights kept constant. Keeping the history of things and the overall plan that worked fixed and isolating LLMs to do subtasks is better than developing a whole database in one continuous context. In some cases, the overall plan "tree" can itself be entirely verbalised, but most commonly there is human modifications/steering.

Can pure-LLM coding harnesses with just verifiers one shot most e commerce sites including planning? Yes. But we want to do more with it than e commerce sites. Will it keep improving thus enabling us to do more and more complex things? No obvious reason for a fixed limit to exist in theory[1]. But at any point on the progress curve, using it with a harness always gives better results versus not. Concretely, with fable 5.1, using it without a harness could not prove FLT in reasonable token budgets [3]. However, it is possible for say, idk, GPT9, trained on this, to verbalise this whole proof tree, and also potentially generalize it to another open problem, purely AR, in a reasonable token budget[2]. This was how we got from gsm8k to FLT in the first place.

It's not a binary "AR is useless" "AR is all you need".

[1] the limits are mostly economic, and time is itself a limit, see https://news.ycombinator.com/item?id=49161078 Tl;dr diminishing returns of test time scaling. Noam brown also has a piece about this.

[2] if it's too many tokens that we run out of time or money literally, that is the limit described in [1]. It is not linear or constant scaling necessarily as described again in [1].

[3] [2] is why we have to add token budgets as another axis apart from r and the autonomy level.

[4] And, the distribution conditioned on that verbalisation must be amenable to sampling the verbalisation of the execution of the plan from. This is not a given, see https://arxiv.org/abs/2504.09762 and https://news.ycombinator.com/item?id=49277303

porridgeraisin··on The asteroid currently hitting front end web development
Yep. It's not that important unless you're really just relying on one shot responses.
porridgeraisin··on Ask HN: Who is using MCP in production?
Many of the other search APIs are better than in the inbuilt claude or codex ones. You can try this out yourself by e.g enabling exa in codex plugins.
porridgeraisin··on Nvidia to acquire Hugging Face
HF is an american company.
porridgeraisin··on Inside Google’s $200bn Wall Street finance machine for Anthropic
https://archive.is/h6ysi
porridgeraisin··on Gemini 3.8 Flash and 3.8 Flash Cyber
agy is good for those cases where you are willing to put the effort into the harness specifically for a task or family of tasks. The full suite, with evals, monitoring, hooks, custom tools, custom verifiers, etc,. It is not good if you want a "general coding assistant" like codex or claudecode.

The reality is that if you optimise a harness for a family of tasks[1], then most of these models give successful output. And there, gemini flash's speed shines.

For general coding assistant, you want it to be well, general, and you use a harness without too much customisation to something specific. Here you need deeply post trained coding assistants and implementors like codex/sol or claude/opus. Gemini flash in its current form will be too happy-go-lucky if you try using it the way we all use codex and is better used in a constrained setting.

tl;dr gemini flash for "LLM-aided workflows in production" is super good today. Cheap as well.

[1] Stuff like this: https://antigravity.google/blog/teamwork-when-ai-becomes-a-r...

https://hamel.dev/notes/llm/evals/

← PreviousPage 2 of 34Next →