HNHacker News
TopNewBestAskShowJobs

sweezyjeezy

2,032 karma · joined December 30, 2014

submissionscomments
sweezyjeezy··on Gemini 2.5 Pro Preview
Again - I'd argue that the extraordinary success of LLMs, in a relatively short amount of time, using a fairly unsophisticated training approach, is strong evidence that coding models are going to get a lot better than they are right now. Will it definitely surpass every human? I don't know, but I wouldn't say we're lacking extraordinary evidence for that claim either.

The way you've framed it seems like the only evidence you will accept is after it's actually happened.

sweezyjeezy··on Gemini 2.5 Pro Preview
I would argue that what LLMs are capable of doing right now is already pretty extraordinary, and would fulfil your extraordinary evidence request. To turn it on its head - given the rather astonishing success of the recent LLM training approaches, what evidence do you have that these models are going to plateau short of your own abilities?
sweezyjeezy··on Ask HN: Share your AI prompt that stumps every model
o4-mini got this right 4 times out of 4.
sweezyjeezy··on A puzzle of two unreliable sensors
There are two iid Uniform noise variables, the Us in A~x + U and B~x or U are independent.
sweezyjeezy··on A puzzle of two unreliable sensors
Let U~Uniform(0,1) Let sensor target measurement be x, so A ~ (x + U), B ~ x or U with probability 0.5. We draw a from A and b from B, we want the estimator to minimise mean absolute error - the bayes optimal rule is the posterior median of x over the likelihood function of L(x | a, b).

Note that if a = 0 and b = 1 -> we KNOW b!=x because a is too small - there is no u with (u + 1) / 2 = 0. I'll skip the full calculation here, but basically if b could feasibly be correct its "atomic weight" ends up being as least as large as 0.5, so it is the posterior median, otherwise we know b is just noise, and the median is just a. So our estimator is

b if a in range [b/2, (b+1)/2]; a otherwise

This appears to do better than OPs solution running an experiment of 1M trials (MAE ~ 0.104 vs 0.116, I verify OPs numbers). The estimator to minimise the mean squared error (the maximum likelihood estimator) is more interesting - on the range a in [b/2, (b+1)/2] it becomes a nonlinear function of a of the form 1 / (1 + piecewise_linear(a)).

sweezyjeezy··on CVE program faces swift end after DHS fails to renew contract
haha, sage advice
sweezyjeezy··on CVE program faces swift end after DHS fails to renew contract [updated]
"for"? You realise this is a homeland security matter for the US as well as the EU?
sweezyjeezy··on CVE program faces swift end after DHS fails to renew contract [updated]
Your comments feel a bit incoherent - just extend your reasoning for why you think Europe should want to fund this back to the US again.
sweezyjeezy··on Googler... ex-Googler
I agree that it's a little hard to care about this author's situation as much as other stories I've heard in the past couple of years. But that said, losing a job like this is never a nice place to be, and I don't hate this person for having those emotions. People are allowed to feel things, shaming them for that is not nice.

But I would caution people against writing public statements like this when they are still in shock, you might regret them later, better to try and regain some balance first.

sweezyjeezy··on Why do we need modules at all? (2011)
I agree, but also agree with the author's statement "It's very difficult to decide which module to put an individual function in".

Quite often coders optimise for searchability, so like there will be a constants file, a dataclasses file, a "reader"s file, a "writer"s file etc etc. This is great if you are trying to hunt down a single module or line of code quickly. But it can become absolute misery to actually read the 'flow' of the codebase, because every file has a million dependencies, and the logic jumps in and out of each file for a few lines at a time. I'm a big fan of the "proximity principle" [1] for this reason - don't divide code to optimise 'searchability', put things together that actually depend on each other, as they will also need to be read / modified together.

[1] https://kula.blog/posts/proximity_principle/

sweezyjeezy··on Powers of 2 with all even digits
Base‑10 is just our chosen way of writing numbers, it doesn’t need to have any deep relationship with the arithmetic properties of sequences like the powers of 2. For most series (Fibonacci numbers, factorials etc), the digits for large members will be essentially random, their digits don't obey any pattern - it's just two unconnected things. It seems extremely likely that 2048 is the highest, but there might not be a good reason that could lead to a proof - it's just that larger and larger random numbers have less and less chance of satisfying the condition (with a tiny probability that they do, meaning we can't prove it).

Interestingly, there are results in the other kind of direction. Fields medalist James Maynard had an amazing result that there are infinitely many primes that have no 7s (or any other digit) in their decimal expansion. This actually _exploits_ the fact that there is no strong interaction between digits and primes - to show that they must exist with some density. That kind of approach can't work for finiteness though.

sweezyjeezy··on The Einstein AI Model
> but it also possible to train AIs to approach them without going through the same process as the human scientists

With chess the answer was more or less completely brute force the problem space, but will that work with math / science? Is there a way to widely explore the problem space with AI, especially in a way that goes above or even against the contents of it's training data? I don't know the answer, but that seems to be the crucial question here.

sweezyjeezy··on Natural occurring molecule rivals Ozempic in weight loss, sidesteps side effects
Now we're arguing semantics. "Naturally occuring" in English means not synthetically produced / found naturally in the environment outside of human influences.
sweezyjeezy··on Natural occurring molecule rivals Ozempic in weight loss, sidesteps side effects
They did specify "molecule". Certainly not every pharmaceutical occurs naturally, and other materials such as Teflon don't either.
sweezyjeezy··on I struggled with Git, so I'm making a game to spare others the pain
> Every explanation of git I have seen starts by explaining that git is cool because branches are just pointers and then talks about the index/staging area.

Which is precisely _not_ a "just memorise these commands" approach to git right?

Also re: your general snark - try a bit more empathy? I'm sure you are an experienced dev but we are talking about people LEARNING git, they don't have the same points of reference as you do today.

sweezyjeezy··on Math That Matters: The Case for Probability over Polynomials
> Statistical significance is bullshit. Learning about it is as useful as learning about phlogiston.

Ok, that's where I draw the line - statistical significance is not "bullshit" - however as you say, leaning on it too hard can cause things to break quite badly. Scientists misusing it do not negate all the medical advances we have made from moving to a significance-based system. It is an absolutely essential tool for people using statistics to understand, but its limitations must be emphasised when taught, and it must be understood that it is a tool, not a conclusion. Also other alternatives should be taught more widely (e.g. Bayesian inference).

sweezyjeezy··on I struggled with Git, so I'm making a game to spare others the pain
I actually think git is a bad example of "just memorise these commands" - unless you are working with projects with a small number of users / branches - or you're fine to just delete and reclone if things get really hairy. I think a lot of my struggles with git came down to not grokking what it was actually DOING at first. Examples:

- not understanding branch pointers / staging / committing corrently. E.g. [add file] [modify file again] [commit] - what just happened? (IMO these things could have been named better). Also reset vs revert vs restore - easier to use these if you've internalised branch pointers etc

- git pull fails because it says it would overwrite a file you've never heard of - how is that file on your local? Is it ok to delete it?

- times when you (or your colleagues) need to rewrite history (rebase / squashing etc) - require a pretty good mental model of what is going on to both diagnose issues and to fix them

sweezyjeezy··on Terence Tao – Machine-Assisted Proofs [video]
I was posting questions from the last week, so not in o3 training data.
sweezyjeezy··on Terence Tao – Machine-Assisted Proofs [video]
As an ex-mathematician I was really interested to see how well o3 etc handles difficult, unseen math questions, so I tried giving it some hard-ish questions from mathoverflow [1] (mainly non-trivial questions on graduate+ level topics). It definitely isn't great and may even be more harmful than useful currently. The main issue it will never say "I'm not sure how to do this", it will almost always give a complete answer from start to finish, with 'bugs' along the way that can be very subtle.

But I found it genuinely shocking some of the steps it manages to take successfully, and it definitely doesn't feel like we're a million years from something could replace big parts of researchers' work. I honestly found some things it could do extremely unsettling as a thought-worker.

[1] https://mathoverflow.net/

sweezyjeezy··on Is Ketamine Neurotoxic?
Definitely on the high side, but I'm guessing they mean "chronic abusers" rather than just "user". The serious ketamine addicts I knew when I was younger took the drug every single day, and multiple times during the day. I think that figure seems right, since a single large line can be north of 200mg.

I know: it's insane that someone would do that much ever, let alone every day, but I assure you that it happens.

sweezyjeezy··on After 20 years, math couple solves major group theory problem
Yeah agreed - there are actually many, smaller flashes of insight, but most of them don't lead to anything. I once joked that you could probably compress all the time I was actually going in the right direction in my PhD down to about a month or two. That's a bit glib, often seeing why an approach fails gives you a much better idea of what a proof 'has to look like' or 'has to be able to overcome'. But many months of my PhD were working on complete dead-ends, and I certainly had a few very dark days because of that. Research math takes a lot of perseverance.
sweezyjeezy··on Is ChatGPT Down?
Marked resolved, still not working in London
sweezyjeezy··on Lines of code that beat A/B testing (2012)
Agreed, and articles like this don't help. That's the only point I was trying to make really.
sweezyjeezy··on Lines of code that beat A/B testing (2012)
Yes but here's a exaggerated version - say were to sample for a week at 50/50 when the base conversion rate was at 4%, then we sample at 25/75 for a week with the base conversion rate bumped up to 8% due to a sale.

The average base rate for the first variant is 5.3%, the second is 6.4%. Generally the favoured variant's average will shift faster because we are sampling it more.

sweezyjeezy··on Lines of code that will beat A/B testing every time (2012)
Sampling 50/50 will always give you the best chance of picking the best ultimate 'winner' in a fixed time horizon, at the cost of only sampling the winning variant 50% of the time. That's true if the reward rates are fixed or not. But some changes in reward rates will also cause MAB aggregate statistics to skew in a way that they shouldn't for a 50/50 split yeah.
sweezyjeezy··on Lines of code that beat A/B testing (2012)
One of the assumptions of vanilla multi-armed bandits is that the underlying reward rates are fixed. It's not valid to assume that in a lot of cases, including e-commerce. The author is dismissive and hard wavy about this and having worked in in e-commerce SaaS I'd be a bit more cautious.

Imagine that you are running MAB on an website with a control/treatment variant. After a bit you end up sampling the treatment a little more, say 60/40. You now start running a sale - and the conversion rate for both sides goes up equally. But since you are now sampling more from the treatment variant, its aggregate conversion rate goes up faster than the control - you start weighting even more towards that variant.

Fluctuating reward rates are everywhere in e-commerce, and tend to destabilise MAB proportions, even on two identical variants, they can even cause it to lean towards the wrong one. There are more sophisticated MAB approaches that try to remove the identical reward-rate assumption - they have to model a lot more uncertainty, and so optimise more conservatively.

sweezyjeezy··on A Puzzle about a Calculator
Nice - the n1 + n3 = n2 + n4 equality is only necessary (mod 11) e.g. 9020 works - this is because 99...99 with even # of 9s is divisible by 11 and with odd # 9s is divisible by 11 if we subtract 9 (or add 2) so then is = -2 mod 11. So then for example with 4 digits

  1000a + 100b + 10c + d = [a + b + c + d] + [999a + 99b + 9c]
                         = [a + b + c + d] - 2a - 2c (mod 11)
                         = (b + d) - (a + c) (mod 11)
sweezyjeezy··on Can LLMs write better code if you keep asking them to “write better code”?
for 10^5, to get the same collision probability (~2 * exp(-10)), you would just need to compute the 10 maximum/minimum candidates and check against those.
sweezyjeezy··on iTerm2 critical security release
Yes I use them together, iterm has a great tmux integration. Tmux vanilla does not have great UX (in my opinion).
sweezyjeezy··on iTerm2 critical security release
I use it primarily for its split pane functionality. Invaluable if you need to see multiple things happening on the same machine at once. I work in data science and often have several long running jobs on a single server, a notebook server, htop/iotop, nvidia-smi, or simply just having different panes cd'd to different directories - with iterm you can organise to a single terminal tab for each machine (including local), or group tabs across machines if they are for related work.
← PreviousPage 3 of 17Next →