HNHacker News
TopNewBestAskShowJobs

jeremysalwen

1,027 karma · joined November 25, 2011

submissionscomments
jeremysalwen··on Ukraine Can't Stop Russia's Jet-Powered Shahed Drone
It is not weird to call something by the name of something that is much more widely known, that it is closely related to. It is maybe a little bit weird to get really angry online about that, unless you are upset about the association being reinforced.
jeremysalwen··on Understanding the Impact of LLM Watermarking on AI Agent Behavior
It introduces statistical regularities, but all RNGs introduce statistical regularities. So the question is on average are these statistical regularities better or worse than those introduced by the alternative, and the answer is no, if implemented properly.
jeremysalwen··on Understanding the Impact of LLM Watermarking on AI Agent Behavior
You have a misunderstanding of how LLM generation works. Before any watermarking gets involved with these models there is ALWAYS a random seed used for generation. For any prompt, some seeds will give better answers, and some will give worse ones.

Let's say there are four billion possible seeds. There are four billion possible ways we could watermark the generation. We could say "we will choose seed 1, that way we will know exactly what output it produced", we could say "we will choose seed 2, that way we will know exactly what output it produced"... etc etc. Now, if we decide "not to watermark", we STILL must choose a seed. So we are actually still applying one of the watermarks, the only difference is we are not careful to remember which one. Could some seeds give a better or worse answer to some specific prompt? Yes. Could choosing a random "watermark" to apply be better or worse on average than choosing a random seed to apply? No. It's mathematically impossible.

This is like an open source project changing their seed from "12321" to "43", and saying that because we changed the seed, the quality is "necessarily lower".

jeremysalwen··on Coulomb's law remains tricky to test at home
I thought it was "obvious" based on the principle that two charges at the same location should have the same force as one combined charge at that location. Of course this immediately brings up the question of the self-force of a point charge...
jeremysalwen··on Best LLM for every budget, updated daily
It's missing Opus 5.5 which was released over a day ago (and also is clearly on the pareto frontier).
jeremysalwen··on A misalignment of AI in mathematics
I agree that the declaration doesn't focus on credit, but I think it's still at the root of the problem. Because ask yourself: if the AI generated proofs are not creating any new ideas or insight, just brute forcing a boolean true/false result, then why can't mathematicians simply ignore their results? Why does it matter if OpenAI or even amateurs with AI are "solving" these problems, without contributing to any deeper understanding?

I don't think "intellectual poisoning" is really the mechanism that harms the mathematics community.

The harm is if you have a community of mathematicians who are focused on expanding human understanding, then having instant access to a bunch of AI proved results muddies the water about who has contributed what. If someone could scoop any significant theorem at any time by pointing an AI at it, how do you really demonstrate that you have created new understanding? Or that your new understanding is about something important? How do you prove that the AI needed your new concepts to be able to solve it?

jeremysalwen··on A misalignment of AI in mathematics
To me it doesn't seem like what AI has destroyed is the ability for mathematicians to develop understanding and share it with each other, but rather it's destroyed the yardstick (solving open problems) that has traditionally been used to measure how much they have contributed to that understanding.

I do see how this is a problem in terms of assigning credit, but I think the cat is already out of the bag in terms of these models being capable. Even without AI labs spending millions of dollars to solve millennium prize problems, there are plenty of other people who will use them to pick low hanging fruit. I don't think any social solution is going to make things go back to the way they were, where you could share your progress towards a famous open problem without risking someone "scooping" you within a couple of days.

I think that the most likely outcomes are either mathematics becomes more secretive, or there is a more deliberative approach to assigning credit than who was "first" to solve some problem. In the former case, this may slow down progress, and in the latter case, this could mean that credit would become more subjective, and be a continual source of controversy.

jeremysalwen··on New type of dice guarantees no tie when deciding who goes first
Sorry, by harder I meant for the same number of faces. E.g. the article is about 5 60 sided dice, which was a breakthrough, but there are other known solution where all the dice are less than 60 sided, but they are unequally sized.
jeremysalwen··on New type of dice guarantees no tie when deciding who goes first
The key part is that they have to be fair when any subset of the dice is rolled together, not just when all five are rolled. Also if the dice are allowed to be different sizes then it's easier as well.
jeremysalwen··on Pacing model development in an era of cyber-critical capabilities
Just going to add my commentary as someone who has trained LLMs and understands them deeply: absolutely nothing he has posted in this thread suggests he lacks any relevant understanding of how LLMs work. If you disagree, you should point out something specifically that you think was wrong. (And honestly you should have made a specific point in your first post).
jeremysalwen··on Advanced AI Sycophancy
I had a similar thought when I did a file by file review of a large codebase for bugs, and Claude commented at the end "the worst file was foo.cpp", which happened to be one of only a few files I hadn't mostly written myself. It made me field proud for a second, and then uncomfortable thinking how it had access to the git history, and surely it could predict I might feel complimented...
jeremysalwen··on Trade, merchants, and the lost cities of the Bronze Age (2019) [pdf]
What is a meal though? Do you get the same volume of food from all vendors, and the same number of calories? I definitely order multiple items from restaurants that have small portion sizes, so I suspect some vendors would be better or worse deals even at a uniform price.
jeremysalwen··on The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" (2023)
I am surprised nobody links this blog post demonstrating that the paper's conclusion is not true (even for gpt3.5): https://andrewmayne.com/2023/11/14/is-the-reversal-curse-rea...

It seems like restrictions on the model talking about non famous people might have been responsible for the appearance of the models being unable to do this.

jeremysalwen··on Peopleless economy? Not technically impossible
Wages could go up, it's just in the form of trillionare paying another trillionare a trillion dollars a day. GDP would be looking rosy!
jeremysalwen··on Peopleless economy? Not technically impossible
I don't see why having ASI would make the top 0.001% less interested in using the energy, minerals, and land on earth. Just because they have no interest in your labor doesn't mean they have no interest your house or your energy supply. "Humans will be so rich they will fuck off to other parts of the world and leave gorilla habitat alone" hasn't really panned out so well.
jeremysalwen··on Robinhood now lets your AI agents trade stocks
That sounds like more training data that the human is just regurgitating. Nobody I know has ever had an original thought, just combined existing thoughts that were in their training data, in new combinations.
jeremysalwen··on Green card seekers must leave U.S. to apply, Trump administration says
Is the legal precedent they are ignoring only 36 years old? No? I guess that makes us talking about case law older than 36 years then. (As we all know, laws less than 40 years old are option to follow anyways).
jeremysalwen··on FusionCore: ROS 2 sensor fusion (IMU and GPS and encoders)
I don't like saying "oh this is AI generated" because of exactly this scenario, where maybe there is care behind it and you just used AI to help you write. The problem is that as a reader I have no way to tell, and aillI have a lot of experience reading AI generated writing where it's subtly incorrect or misleading.

I would really suggest you rewrite the README, to remove the cliches and make it clear how much effort you put in. You clearly are able to communicate well about this project on your own.

jeremysalwen··on FusionCore: ROS 2 sensor fusion (IMU and GPS and encoders)
I looked into open source sensor fusion libraries for open mower, it seemed like the WOLF framework was very promising as a pluggable sensor fusion system. http://mobile_robotics.pages.iri.upc-csic.es/wolf_projects/w.... I only used it briefly though.

I hate to say this, but this submissions readme seems obviously AI generated.

jeremysalwen··on Ask HN: Who is hiring? (April 2026)
Against Malaria Foundation | Senior Software Engineer | Remote (UK) | Full Time

The AMF works to fight malaria in an extremely cost effective way with insecticide treated bed-nets. Every year hundreds of millions of people will be infected with malaria, and half a million of them will die, with many more being debilitated or disabled. Independent charity evaluators, such as GiveWell, have ranked the AMF as one of the most cost-effective charities in the world for over a decade, due to our long and well-studied track record preventing hundreds of millions of cases of malaria.

Our tech team is just a couple people, and all-remote. Our software is essential to the charity work we do; it is one of the ways we stand out among other charities, by allowing us to automate, analyze, and validate our work in novel ways. Our tech stack includes C# .NET, Blazor, and Python.

If you are a senior SWE in the UK who is interested in making a difference, please reach out to me by email at my first name at againstmalaria.com

jeremysalwen··on Sodium-ion EV battery breakthrough delivers 11-min charging and 450 km range
Just because you state your opinion confidently, does not mean you are correct. For example, as of 2024, there are 30 billion kilograms of proven reserves of lithium, more than enough to replace every single one of the 1.5 billion ICE cars in the world with an electric car. Please focus more on getting the facts right, and less on speculating about the character of other commenters in an overemotional manner.
jeremysalwen··on Some silly Z3 scripts I wrote
I'm suspicious of the theorem proving example. I thought Z3 could fail to return sat or unsat, but he is assuming that if it's not sat the theorem must be proven
jeremysalwen··on Deal infinite damage for 4GRU, as long as the twin primes conjecture is true
Does anyone know of a board state that doesn't involve arbitrary actions, but a specific finite setup, which would nonetheless require the solution of an unsolved conjecture in order to resolve who wins?

I am thinking of an example like constructing two incredibly large numbers (of the same sort as grahams number) where it is not known which is larger, and then e.g. doing x damage to a creature with y health. Ideally this happens in a way that nobody would object to the construction if x and y, since they are simple to describe and construct.

I understand that you can construct a turing machine to perform an arbitrary computation, but that defeats the spirit of this question. The question is whether there is a simpler way to construct such a paradox where you might follow along happily until you get to the end.

jeremysalwen··on Tesla ending Models S and X production
I'm not sure I agree. I think just having wings that flex a bit is mechanically simpler than having an additional rotating propellor. After all, rotating axles are so hard to evolve they never almost never show up in nature at a macro scale. Sort of a perfect analogy to lidar actually. We create a new approach to solve the problem in a more efficient way, that evolution couldn't reach in billions of years
jeremysalwen··on My Tamagotchi is an RL agent playing Slither.io
As someone who implemented some RL algorithms and applied them to a real world game, (including all the ones mentioned in the article), I would be surprised if the implementation is not buggy. That is one of the most striking things about RL, the extent to which it is hard to find bugs, since they generally only degrade the performance instead of causing a crash or obviously wrong behavior. The fact that he doesn't mention a massive amount of time spent debugging, and the longish list of things that were tried that really should have worked but didn't, suggests to me it's probably still buggy. I suppose it is possible that LLMs could be particularly good at RL code since it's seen it repeated so many times... But I would be skeptical without hard evidence.
jeremysalwen··on Accounting for Computer Scientists (2011)
Double entry book keeping is just recording the "edge" in two places, once based on the source node, and once based on the target node. So you have a nice list of all edges coming from each node and a nice list of all edges going to each node. This was important before computers, since the process of looking up all edges going to/from a node would take real time and effort.

For your example of the landlord and the tenant, think, what if the landlord wanted a list of all payments that went into a specific bank account, what if the tenant wanted a list of all rent payments, etc. It's basically a database index to speed up those queries, but for a written database that is updates by hand. The fact that there is redundancy is just a bonus because you can now notice if the two places a piece of information are written down don't match.

jeremysalwen··on I don't think Lindley's paradox supports p-circling
Admittedly not a statistician, but I think the article is missing the point. The reason why people circle the P values is because nobody actually cares about the thing the p-value is measuring. What they actually care about is whether the null hypothesis is true or some other hypothesis is true. You can wave your hands around about how actually when you said it was significant what you were really saying was something technical about a hypothetical world where the null hypothesis is factually true, and so it's unfair to circle your p value because technically your statement about this hypothetical world is still true. This is not a good argument against p value circling, but rather it merely demonstrates that the technical definition of a p value is not relevant to the real world.

The fact remains that for things which are claimed to be true but turn out to not be true later, the p values that were provided in the paper are very often near the significance threshold. Not so much for things which are obviously and strongly true. This is direct evidence of something that we already know, which is thst nobody cares about p values per se, they only use them to communicate information about something being true or false in the real world, and the technical claim of "well maybe x or y is true, but when I said p=0.49 I was only talking about a hypothetical world where x is true, and my statement about that world still holds true" is no solace.

jeremysalwen··on AI Is Destroying the University and Learning Itself
Did anyone else think there were several transitions that seemed like pure GPTisms?

> This isn’t innovation—it’s institutional auto-cannibalism. The new mission statement? Optimization.

jeremysalwen··on Zero knowlege proof of compositeness
To be honest I feel like I have seen much better expositions of zero knowledge proofs. The playing cards example is nice in some ways, but people are often exposed to trickery regarding playing cards. The recipient of the proof needs to verify that the deck of cards is a normal deck of cards, that no cards have been swapped out or altered, etc. These are all precisely the things that magicians are regularly able to fool people about. So really you have to make an additional assumption of "no funny business", which distracts from the mathematical core of what you are trying to demonstrate.

Likewise, the example of compositeness is a bit off because even though there is knowledge about the composite number that the proof does not reveal, that knowledge is in fact not known the to person constructing the proof either! The proof is not really zero knowledge either, since it gives the reader knowledge of a specific witness to its compositeness.

Even the wikipedia example of going into the cave (which used to be featured more prominently in the article) I think is terrible. Why wouldn't you just walk a loop to prove you know the way through the secret door? Also, it's clearly not zero knowledge, as it reveals some information about how quickly they can pass through the gate.

In general I think avoiding physical examples is necessary, since reality is complicated, and in the real world some information always leaks.

I think the best example for teaching about ZKPs is the graph isomorphism problem: Given two large graphs, you can prove that you know a isomorphism between two graphs by generating a new randomly labeled graph that is isomorphic to both of them and showing it to the provee, who can then ask you to demonstrate that this new graph is isomorphic to either graph A or graph B. Since you don't know ahead of time which one they will ask for, the only way you could consistently pass this test is if you actually do have a graph that was isomorphic to both A and B simultaneously. But since you only reveal one of the isomorphisms, it really is zero knowledge.

jeremysalwen··on Fizz Buzz without conditionals or booleans
Make it throw an exception with an index out of bounds to terminate the loop.
Page 1 of 10Next →