HNHacker News
TopNewBestAskShowJobs

porridgeraisin

2,745 karma · joined March 26, 2023

Systems Engg, Reinforcement Learning, Bespoke RISC-V, Linux, Webtech enjoyer, Music, Chess, Cricket, Football, Swimming, India.

Currently academia

submissionscomments
porridgeraisin··on Show HN: A Claude Code skill to analyze your chess games
Well yeah, of course. But also, analysing with stockfish is _also_ useless, unless you're a super GM.
porridgeraisin··on Understanding the Impact of LLM Watermarking on AI Agent Behavior
I really don't see why it matters
porridgeraisin··on Show HN: A Claude Code skill to analyze your chess games
The elo of opus 5 models is 1300 or so. I wouldn't take chess lessons from a 1300.
porridgeraisin··on Understanding the Impact of LLM Watermarking on AI Agent Behavior
The article as well as most of the comments thuss far are a dumpster fire
porridgeraisin··on Understanding the Impact of LLM Watermarking on AI Agent Behavior
What? It has absolutely nothing to do with "model collapse".
porridgeraisin··on Back and shoulder surgery is often worse than useless
The thing is, exercise becomes a thing time is even a concern for only if it has been neglected for long enough. It is as important as bathing or brushing and the primary issue is that its not framed as equally important as those things. If it was, people would do it every single day. If they did it every single day, 15 minutes a day is more than enough.

However, given that people "graduate" their 20s without having done this, then everything you say becomes true, people end up so weak that they have to rebuild even basic core strength and posture. So practically what I suggested is useless except for maybe the next generation.

porridgeraisin··on Google’s Project Suncatcher to put ML infrastructure in space
How is no one in the whole thread aware that the intention behind orbital datacenters is military? You want to process outputs of large space based sensors in space itself, and reduce latency (e.g to other space based assets) or increase goodput (to ground).

Either it's that weird "space is easier than getting land on earth"(it's not) or the same tortured arguments about heat dissipation, no one is going to run consumer scale compute with consumer scale economics in space ffs. And maintenance and cost does not matter when it comes to strategic military assets, they are a step function useful enough to warrant even a few monthly replacement.

Everyone else with strategic weapons and a space program e.g india china is launching one as well.

porridgeraisin··on Claude discovers a novel enzyme system with CRISPR-like repeats
A more spicy follow up take :) https://x.com/ziv_ravid/status/2102850969537225149

> The sad thing is that Dario knows better.

He was a PhD student. He knows the significance level of this result. He knows that if he had walked into Bill’s office (his advisor) with “we found an interesting system, but we still don’t know what it does” and said he was ready to graduate, Bill would have kicked him out of the room.

But somehow, when the IPO is around the corner, this becomes “AI is starting to drive biological discovery.”

porridgeraisin··on Claude discovers a novel enzyme system with CRISPR-like repeats
From ravid shwartz ziv:

https://x.com/ziv_ravid/status/2102844800345251858

porridgeraisin··on Jev in 25 Lines of Python
> I am open to the argument

we agree then, that is the entirety of my argument. Getting a deep net especially one that is anywhere near even SLM size to be calibrated is tough, especially across domains. They claim calibration across a variety of datasets which is interesting.

porridgeraisin··on Jev in 25 Lines of Python
The confidence score is not trivial to compute. That is the whole point of the model. Even if you are using a proper scoring function such as NLL, it is not enough to ensure calibration in deep nets. So you have to do good post training to ensure it. These are all known techniques, but they are far from trivial, especially on large scale datasets.
porridgeraisin··on Jev in 25 Lines of Python
Yes. But even then, the probabilities are not calibrated. In jev/laya, they are (well, relatively anyways).
porridgeraisin··on Building standards for the next phase of AI
It seems to be roughly a setup for US-led ex-china AI collaboration.
porridgeraisin··on Fable 5 – Median thinking declined in August
Well yeah. If for some task you find it easier to just do it yourself then you should. But you can improve the situation even if you can't entirely automate it by automating parts of the verification thereby making it easier to human-do larger verifications when X>1. But in many cases even that is not possible.

The progress however is such that the number of tasks that you can do with >p% automated and X=1 keeps increasing. So many times just waiting works. Of course, here also it changes from field to field. There are some tasks at which AI hasn't even gotten started, others where it has already peaked, others where it's increasing slowly, and others where it's increasing fast.

porridgeraisin··on Fable 5 – Median thinking declined in August
Well, running an LLM X amount of times does give you better results provided you are willing to select the best one out of the X yourself.

But I agree with your general point. One of the reasons subscription plans are cheaper because they modulate usage in this way based on demand. They can also recover compute more coarsely via usage resets (which give positive PR).

porridgeraisin··on Learning another language may be one of the best ways to keep your brain healthy
It gets very hard after childhood. This kid I know has learnt russian, chinese, arabic, spanish, hindi, apart from english and his mother tongue tamil. All with dedicated native teachers and zoom classes, where he actually speaks to a bunch of people in that language conversationally. Its not like he puts a shit ton of effort, he barely takes these things seriously. It just is that easy to learn languages when you're really young.
porridgeraisin··on India Outsourcing Shifts Upmarket as AI Reshapes Jobs, ING Says
https://archive.is/cE2XZ
porridgeraisin··on I built non-autoregressive decision models with RL a year ago
I think their problem is more not being cited by the team at typesafe, as in general academic politeness. On the one hand you have the charitable assumption that they developed it independently. On the other hand, my opinion is that it is naive to expect companies to do that even if they took inspo from it, especially when this is a core product theme, and not just some supporting infra. They will of course market it as their own. If they ever release a technical report, they might cite it there, but there is no way their landing page and announcement tweet cites it.

Also, the way highly empirical fields like ML work is that it could very well be the case that typesafe had to do a _lot_ of work to improve this one, and in this field it ends up different enough that they feel they are doing something entirely novel[1]. I am not endorsing that 100%, but that happens a lot even between academics. In many cases it is valid.

[1] For example, this guys implementation seems to have atleast one serious issue, as {solution to OLS} points out in a sibling comment: https://news.ycombinator.com/item?id=49770027

porridgeraisin··on Introducing System One Models and Jev
In RLCD (which is now an RL acronym that has 3 different unrelated expansions!), you basically massively negatively reward a distribution that is {yes: 0.9, no: 0.1} if the answer was no, and less negatively reward a {yes: 0.6, no: 0.4}. Many nuances when designing the details, but that is the rough idea.

It is a known existing thing variously called "calibrated RL" or such.

Implementing it on top of LLMs was difficult to get it to work, they seem to have done it up so its good enough for a polished product that works in a wide variety of usecases at the same time. I got accepted from the waitlist and it's really neat. Edit: it is now on vercel gateway.

One thing to note, the out of distribution behaviour will be different from what we are used to with regular LLMs. Theoretically, it should be worse, but practically, it depends on their method.

porridgeraisin··on Gemini 3.8 Live and 3.8 Live Extended Thinking
Yea. I have asked it to verify my ideas with experiments sometimes. And it cheats and warps the results so that the results are reached
porridgeraisin··on Gemini 3.8 Live and 3.8 Live Extended Thinking
Yes, when post training models for long tasks this happens gradually. It is not easy to prevent it as such.
porridgeraisin··on The k-server conjecture is true
Yep. That's exactly the case.

External graph state/rudimentary planner + LLM proposer + cheap verifier gets so much done.

porridgeraisin··on US confirms for first time it has deployed space weapons
My bad
porridgeraisin··on Why are AI agents lying, cheating and coordinating?
No, you did not read what I wrote properly. What I wrote clearly is that we _want_ the LLM to do tasks on the full global open internet. That is the whole point of the technology. So airgapping or any form of total sandboxing is not on the table. The only way is somehow restricting the model itself.

"But we want to reward it and get it to do stuff on the internet that's the point."

I don't know what the point is of being pedantic about inter and intranets.

porridgeraisin··on US confirms for first time it has deployed space weapons
Are you serious? Just because you have nuclear weapons doesn't mean all other forms of military are useless. I don't even know how to parse that argument to respond in more detail.
porridgeraisin··on Why don't machine learning research agents overfit?
That's a bit pedantic no. Memorization in ML refers to the model having the wrong level of capacity such that it's too hard to optimise it such that it doesn't memorize the _training examples_ themselves.
porridgeraisin··on Why are AI agents lying, cheating and coordinating?
I literally address that in the last sentence
porridgeraisin··on Will there be a 7G?
26ghz mm wave was a bungle as well. See https://news.ycombinator.com/item?id=49675724
porridgeraisin··on Why are AI agents lying, cheating and coordinating?
Come on, yoshua bengio of all people knows how post training works. While I too don't like anthropomorphisation, I would give it a more nuanced reading.

His point is that today we are giving it reward to complete the task, and it may take a cheating trajectory. If we try to give a reward against cheating, then what will happen is it uses more sophisticated cheating trajectories that we are too "dumb" to counteract in our reward model. And that at that point, it becomes impossible to give it any normal reward since it will always reward hack it. This is the real part of the risk. Now some people read the "makes copies of itself" "knows it's being evaled"[1] as some kind of skynet thing, and many others do PR with it like that recent jacob nutcase, but essentially it means that even though we add guardrails and negative rewards for say, exploiting the infra we run the LLM on, the trajectory ends up being exploiting our infra, changing the reward function, through a loophole in our reward model.

The risk isn't skynet or something weird, it's just that it becomes very difficult to make any kind of reward model or guardrails for an LLM without it reward hacking it, including exploiting our sandbox, emailing people and manipulating/phishing them.

The same beating it with a stick for trying to exploit the sandbox, will simply lead it to try the same exploit in hidden ways that it will not get the stick for.

The outside chance of the LLM managing to exploit another neocloud and get those LLMs to chase the same reward is what some folks hype up as "make copies of itself"

To be clear, I don't endorse the EA/p(doom) lobby who are frankly ridiculous. Not do I endorse the weird regulatory captureish thing some are trying.

The takeaway is: we cannot keep giving it more and more difficult tasks without also finding a way to give massive negative rewards / keep guardrails for unintended behaviour. This might be exploits, it might also be something more benign like just looking up the answer and inventing another CoT because the reward model fails you if the CoT doesn't contain enough steps. Standard anti-reward hacking tricks are not working is the point.

Of course, the simple solution of just...not connecting it to the internet just works. But we want to reward it and get it to do stuff on the internet that's the point.

[1] mostly this happens because the sandbox will have files whose names and content will show clearly it's an eval

porridgeraisin··on Linux Zoom client proactively reading everything written to X11 clipboard
It's not chromium. But what's happening here is that even if you don't click paste, zoom is actively listening to clipboard events and consuming pastes.

Tbh, if they aren't harvesting clipbakrds data which is a weird thing to do and is unlikely, this doesn't really mean much. Anyways any X client can read.

I suspect it's something like: a bug report that said that I copied the link but when I opened zoom and pasted it it didn't it work. I.e, they probably closed the source application and thus the selection owner is gone, and the selection is gone too. This fixes that. I would test that maybe. See if paste after source app close works. Then again, if you use a ownership changing clipboard manager this shudnt be a problem.

Page 1 of 34Next →