HNHacker News
TopNewBestAskShowJobs

thatguysaguy

453 karma · joined May 23, 2023

submissionscomments
thatguysaguy··on Kolibri: A Sovereign Open-Weight Model
> 3B active

> built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace

Cool to be doing more independent model lineages, but I sure hope no one actually uses this as part of any aerospace engineering...

thatguysaguy··on GPT-6 Astra
part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.
thatguysaguy··on GPT-6 Astra
presumably that's a safety evaluation not a training setting
thatguysaguy··on Chip Off The Old Block
Scott's parenting posts are some of his best
thatguysaguy··on Show HN: TownSquare, a tiny presence layer for websites
The contrast between the example screenshots and the standard internet behavior in the live demo is hilarious
thatguysaguy··on Python JIT project was asked to pause development
The python overhead of launching big ML jobs is nontrivial, so I think speeding that up would be meaningful. (I mean the initial tracing and other setup, not things once the GPUs are actually doing the work).
thatguysaguy··on xAI Is Reportedly Using Just 11% of Its 550k Nvidia GPUs
It looks like people in this thread are confusing fleet utilization and MFU. If they're doing a lot of RL, it's really not surprising to see such low numbers.
thatguysaguy··on Google plans to invest up to $40B in Anthropic
That's b/c the people working on Gemini serving are in GDM.
thatguysaguy··on Our eighth generation TPUs: two chips for the agentic era
What sort of workloads are you thinking of?
thatguysaguy··on Lean proved this program correct; then I found a bug
What is up with people saying you cannot prove a negative? Of course you can! (At least in formal settings)

For example it's extremely easy to prove there is no square with diagonals of different lengths. I'm the hard end, Andrew Wiles proved Fermat's Last Theorem which expresses a negative.

That's just a nit though, you're right about the infinite regress problem.

thatguysaguy··on Can you reverse engineer our neural network?
Ah dang. When I did this I also thought the length bug was intentional but I didn't figure it out before I started my new job, so I dropped the puzzle.
thatguysaguy··on Large-Scale Online Deanonymization with LLMs
Maybe I missed something, but I see little evidence that there is a concerning ability to deanonymize. Many people post under a pseudonym but then link to their GitHub etc. In fact by construction the HN dataset _only_ consists of people who are comfortable with their real identity being linked to it.

The real question is whether someone who is pseudonymous and actually attempting to remain so can be deanonymized.

thatguysaguy··on Gemini 3 Deep Think drew me a good SVG of a pelican riding a bicycle
You can just try other svgs, I got some pretty good ones.

(*Disclaimer: I work for Google, but also I have zero idea about what they trained deepthink on)

thatguysaguy··on TPUs vs. GPUs and why Google is positioned to win AI race in the long term
TPUs predate LLMs by a long time. They were already being used for all the other internal ML work needed for search, youtube, etc.
thatguysaguy··on AI has a deep understanding of how this code works
I'm actually not talking about whether the PR works or was tested. Let's just assume it was bug-free and worked as advertised. I would say that even in that situation, they should not accept the PR. The reason is that no one is the owner of that code. None of the maintainers will want to dedicate some of their volunteer time to owning your code/the AIs code, and the AI itself can't become the owner of the code in any meaningful way. (At least not without some very involved engineering work on building a harness, and since that's still a research-level project, it's clearly something which should be discussed at the project level, not just assumed).
thatguysaguy··on AI has a deep understanding of how this code works
A big part of software engineering is maintenance not just adding features. When you drop a 22,000 line PR without any discussion or previous work on the project, people will (probably correctly) assume that you aren't there for the long haul to take care of it.

On top of that, there's a huge asymmetry when people use AI to spit out huge PRs and expect thorough review from project maintainers. Of course they're not going to review your PR!

thatguysaguy··on FFmpeg dealing with a security researcher
It's a volunteer run project... Saying that they have a duty to do anything other than what they want is quite strange.
thatguysaguy··on Updated practice for review articles and position papers in ArXiv CS category
Verification via LLM tends to break under quite small optimization pressure. For example I did RL to improve <insert aspect> against one of the sota models from one generation ago, and the (quite weak) learner model found out that it could emit a few nonsense words to get the max score.

That's without even being able to backprop through the annotator, and also with me actively trying to avoid reward hacking. If arxiv used an open model for review, it would be trivial for people to insert a few grammatical mistakes which cause them to receive max points.

thatguysaguy··on Meta is axing 600 roles across its AI division
FAIR is not older AI... They've been publishing a bunch on generative models.
thatguysaguy··on BERT is just a single text diffusion step
Back when BERT came out, everyone was trying to get it to generate text. These attempts generally didn't work, here's one for reference though: https://arxiv.org/abs/1902.04094

This doesn't have an explicit diffusion tie in, but Savinov et al. at DeepMind figured out that doing two steps at training time and randomizing the masking probability is enough to get it to work reasonably well.

thatguysaguy··on [dead]
I would recommend going and reading what the BlueSky leadership actually wrote, rather than this post's summary of it.
thatguysaguy··on Are OpenAI and Anthropic losing money on inference?
Why would you think that deepseek is more efficient than gpt-5/Claude 4 though? There's been enough time to integrate the lessons from deepseek.
thatguysaguy··on Are OpenAI and Anthropic losing money on inference?
37 billion bytes per token?

Edit: Oh assuming this is an estimate based on the model weights moving fromm HBM to SRAM, that's not how transformers are applied to input tokens. You only have to do move the weights for every token during generation, not during "prefill". (And actually during generation you can use speculative decoding to do better than this roofline anyways).

thatguysaguy··on What are the real numbers, really? (2024)
Joel's blog in general is an extremely great read. I highly recommend subscribing.
thatguysaguy··on A.I. researchers are negotiating $250M pay packages
At least part of is is that the capex for LLM training is so high. It used to be that compute was extremely cheap compared to staff, but that's no longer the case for large model training.
thatguysaguy··on Ask HN: What's with the repeated job posts on "Who's hiring"?
I both got a job through such a thread, and have now seen the other side of the applicant pipeline. The average applicant (in general, idk about HN in particular) is not very strong! Especially true when you consider the alternative of preserving runway and being patient.
thatguysaguy··on I Want No One Else to Succeed
> Do you think the students in that poll had really thought about the credibility of their university when voting?

That's fair, and I'm not sure of course. I guess a more interesting question would be what if there was another option based on what I described. I do know that this is a conversation that we explicitly had in my department many times. It's was an open secret that cheating was rampant, and a degree from that CS department isn't very prestigious. Those two things aren't unrelated.

To your point about a single class giving all A's not damaging it, you're right of course. My point is that this is a classic tragedy of the commons. One plane flight, extra datacenter etc. isn't moving the needle on climate change, but all put together it does.

> If that were the case, the people protesting student Lian forgiveness should be at tge forefront of demanding increased coverage for Medicare and better social security systems in general

I agree, but my response is simple: I don't think either major party in the US has a principled stance on economic issues. There is wild fluctuation in behavior on an issue-to-issue basis. The fact that most students/political parties/people in the universe don't have a coherent set of princples shouldn't stop us from trying to have them!

thatguysaguy··on I Want No One Else to Succeed
I think the author doesn't understand the example correctly, although to be fair I don't think the professor put the most important option on there either.

Imagine there are two schools, one gives all students a 95% in all their classes, one grades normally. Which school do you interview people from? When a teacher gives out free A's (or when students cheat), it's not a victimless change, it degrades a shared resource (the credibility of the school).

The loan forgiveness thing is not a free action, it is a handout of money to a specific demographic (college graduates), and in particular one that is much more affluent than the people who need government subsidies the most.The government handing out money is not a free action!

thatguysaguy··on The "AI 2027" Scenario: How realistic is it?
Yeah I wouldn't make a deal like this with someone who is operating in bad faith... The cases I've seen of this are between public intellectuals with relatively modest amounts of money.
thatguysaguy··on The "AI 2027" Scenario: How realistic is it?
Some people do actually have end of the world bets out but you have to structure it differently. What you do is the person who thinks the world will end is paid cash right now, and then in N years when the world hasn't ended they have to pay back some multiple of the amount the original amount.
Page 1 of 4Next →