HNHacker News
TopNewBestAskShowJobs

davesque

7,116 karma · joined February 17, 2012

Opinions are my own, blah blah
submissionscomments
davesque··on America.gov
I'm a bit concerned that a government chatbot like this would include a user-feedback signal like thumbs up/down. I think this could be gamed by bad actors to promote/demote specific types of content in the training data, assuming they use those signals in some sort of post training. Hopefully they're accounting for that.
davesque··on I don't want to read what you didn't write
AI-generated writing does not bother me, unless it is intended as prose writing, where the flow communicates as much as the specific facts that are described. For technical writing, I usually just scan for important information anyway, and in that sense the only thing that bothers me about AI-generated writing is if it's too verbose and obscures the important information by leaning ineffectively on jargon and high-minded verbiage. I find that Claude Opus is very guilty of this, whereas GPT 5.6 variants are not.
davesque··on Pion, an agent designed to run any company autonomously
Can we please make it so that AI CEOs are smart enough not to optimize everything according to the focus group that is their shareholders? Seems like we could solve the enshittification problem with this if we play our cards right.
davesque··on David Sacks: OpenAI and Anthropic Don't Need Regulations to Pace Frontier Models
Unless humanity can figure out a way to make a binding international agreement to pace development, it doesn't seem practical to expect US frontier labs to do it.
davesque··on OpenAI’s Navier-Stokes release included a Lean 4 formal proof
Regarding automatic formalization of proofs using AI, how do we know the formalization doesn't contain errors?
davesque··on The Navier–Stokes Millennium Prize Problem
Isn't it also a simple idea that a model designed to recall relevant information from its training data, which is also known to have been trained on data from user transcripts, would, in fact, reproduce directly relevant work by leading experts in the field? Seems like Occam's razor would apply to that situation as well.

We know that LLMs are trained to recall relevant info. We know AI vendors are using user transcripts to train models. Two plus two equals four, right? I mean, an LLM that failed to recall the transcripts of those researchers would be a bad model.

davesque··on Navier-Stokes – Tristan Buckmaster [pdf]
Honestly this whole thing is so fucking weird. I feel like there's an argument that absolutely no one involved in the final crossing of the finish line to the proof actually did any work (other than just intelligently directing an LLM) and deserves any credit. As the author of this doc mentions, the mathematicians who did the actual work that led to the formulation of this approach (without the use of LLMs; just good ole' fashioned human intellect) are the ones who deserve the credit.

Imagine that a no name janitor used their time in the evenings to go spelunking through the literature to push an LLM to this result. No one would care because that person isn't an anointed expert. So why would the expert deserve any more credit? Because they sort of understand the result, even if they couldn't have achieved it on their own? The whole issue of credit for AI-assisted discoveries seems like it's going to run into a brick wall pretty soon.

davesque··on 'Lake America' makes one thing clear: We can't trust U.S. tech companies
And then he wanted New Mexico to be New America. Capitulation is not just about what they're asking for now, but also what they'll be emboldened to ask for in the future.
davesque··on Political meddling at the Census Bureau damages the US statistical system
Corruption in the US Census Bureau? No fuckin' way!
davesque··on The American Worker vs. the Most Qualified
The article doesn't address what I think are the main complaints about H-1B: that it is mostly used to cut costs and hire people who have very little freedom to switch jobs. That and all the trickery with ghost jobs meant to only give the appearance of hiring stateside. The example given is also not great. Von Braun was one remarkable foreign expert in a leadership position. On the other hand, in my experience, many H-1B hires tend to occupy low-level jobs; ironically the very ones that the author says will be filled by motivated Americans. I mean...just try working at any American FAANG. The evidence is totally in your face.
davesque··on Canada will match US tariffs 'dollar for dollar' as trade talks break down
In a way, it's nice that Trump is giving such a good, practical demonstration of why being a bully never works out the way you think it will. A good education for those who need it. Just sucks that we all have to pay for it out of pocket.
davesque··on The Trumps' Crypto Project Just Got One Step Closer to Becoming a Bank
The voting public is not the same as the general public.
davesque··on The Trumps' Crypto Project Just Got One Step Closer to Becoming a Bank
Who's clapping?? I think the man should be in jail. A minimum of half the country has always hated this guy's guts.
davesque··on US Military's cyber command unit grapples with cluster of deaths by suicide
That's not saying much. The real distribution is probably heavily skewed by risk factors such as poverty, age, drug abuse, etc.
davesque··on LLMs reward expertise
I actually don't feel like Tao's recently published conversation is the best example of this idea. As intelligent as Dr. Tao is, and surely more so than me, I got the feeling that he wasn't running up against failure states of the model, which I'm not sure you could attribute entirely to his expertise. I honestly think it was more a matter of luck that the model apparently had so much training data on the topic or that it was architecturally so well suited for it. On the other hand, I've had really surprising moments where Claude was just failing terribly to execute simple dev ops tasks having to do with log processing. And I'd be so bold to say that I don't think it could have been explained by a lack of expertise on my part, or even a misuse of the model.

So yeah, sometimes LLMs reward expertise, sometimes they don't. I guess either way it helps to have it.

davesque··on Startup founders urge U.S. government not to shut off Chinese open weight AI
Everything he does benefits his adversaries because they're all playing him like a fiddle because he's basically the dumbest person alive.
davesque··on Terence Tao's ChatGPT conversation about the Jacobian Conjecture counterexample
I disagree. In fact, I think the field of dev ops gives a clean analogy with mathematical proofs. My point was that my work often requires that I figure out a consistent way to prove to myself what the condition of a system is by asking the right probing questions about it. What I've seen is that even the best LLMs lack a good intuition about what questions they should be asking and instead reach for the quickest and most obvious checks that leave edge cases uncovered. Maybe it is something about the domain of the problem; I don't know. But it makes it hard for me to imagine that an LLM wouldn't make similar errors in other cases, especially when generating mathematical proofs that will soon be too dense for humans to review.
davesque··on Show HN: Claude-thermos keeps your Claude session warm for you
Exactly. I really wish people wouldn't use this. If this becomes popular, anthropic will just modify their cache policy to be much less fair. It's not like they have infinite cache.
davesque··on Terence Tao's ChatGPT conversation about the Jacobian Conjecture counterexample
Without any more context, "keep going" seems to be doing a lot of work. The user is placing a lot of faith in the LLM to not make subtle logic mistakes and to take good approaches to each problem. In my experience, even frontier models (such as Fable) are quite capable of getting confused during even simple technical work I've done in the dev ops world. For example:

LLM: This package hasn't made it to production.

ME: are you sure? i see it right here!

LLM: You're right to push back. I inferred that based on weak data. I see now that the package has been deployed!

If the above conversation is typical for me, how could one expect to achieve a sound result by repeatedly prompting an LLM to simply "keep going" in dense mathematical proofs? Perhaps the user in this case had actually checked the LLM's work before issuing the prompt, but I think you see my point anyway.

davesque··on House Votes for Permanent Daylight Saving Time
Not sure I follow. I can't imagine humanity has evolved much since 1974.
davesque··on Professor denounces mass AI fraud on an exam at Brown
We're not talking about ambulances though?
davesque··on Professor denounces mass AI fraud on an exam at Brown
What if taking longer leads to a better result? Doesn't faster imply less thought?
davesque··on Professor denounces mass AI fraud on an exam at Brown
Care to explain?
davesque··on U.S. allows Anthropic to release Mythos AI to ‘trusted’ US organizations
Also known as, the tu quoque fallacy. Just because politicians in both parties have been doing this for decades doesn't mean that this administration is not especially hypocritical for doing it after whinging so much about free speech and free markets.
davesque··on U.S. government will decide who gets to use GPT-5.6
I'm finding it extremely hard not to have a cynical perspective on all of this. There's an idea that I've been mapping onto this whole this, which could be called something like effective knowledge. Regular old knowledge is just information, or access to information. Effective knowledge is the integration of all that information into an understanding that can be acted upon. That requires things like time, money, and involves the usual socioeconomic hurdles that have separated people into groups like "laborers" and "knowledge workers". Sure, in theory "anyone" could read textbooks and learn, but only a select few have the time, money, mentors in their lives, and so forth to really do that.

The rise in capability of LLMs over the past year has basically removed a lot of these boundaries for people. Learning, building, and experimenting is a lot easier when you have a capable partner like Claude to help you along the way. Claude doesn't always get everything right, and you have to be a skeptic, but it's a lot better than nothing.

When I see the government restricting access to LLMs (or Anthropic as they were doing with Mythos before the whole Fable debacle), I basically just see the same old pattern of the ruling class moving to protect their advantage by keeping the great masses in ignorance. Broadening access to LLMs (i.e. effective knowledge) would put everyone on a more level playing field. But we can't have that, because politics, nations, the economy, blah blah reasons reasons. Guess utopia will just have to wait a bit longer.

But then again, this feels a lot like cryptography export controls. Those controls are in place, but I doubt anyone really thinks they work or make much of a difference. Software is not like nuclear weapons, and a data center is a much smaller lift than a Uranium enrichment facility. So maybe this is just a temporary roadblock. But let me tell you, I sure am ready for it to feel more like the government is working for (not against) the people.

davesque··on Feds freaked over Fable 5 after 'fix this code', not jailbreak, say researchers
Kind of highlights how ridiculous their notion of safety is in this case. By this measure, I guess making the model "safe" means making it play dumb and intentionally ignore security bugs that it notices in the code? And what will the eventual legality of this look like? "Yes, your honor, we allege that this AI system that was sold to us willingly and knowingly ignored a critical security vulnerability in our software system, thereby leading us to be hacked and causing our business to fold."

It's exactly the same problem as backdoors in crypto systems. Criminals will find the crypto that isn't broken and use it regardless (or make it for themselves), while the rest of us losers are stuck with the broken version that we're allowed to use.

On this issue of cyber security, it seems better if authorities just start acting like the cat is out of the bag instead of pretending like it isn't. ASI is basically here now, so what are we going to do about it? Let's not bother pretending otherwise.

On another note, I doubt this was anything other than a vindictive administration enacting revenge on a party that refused them. We all know the Trump admin's priorities.

davesque··on Statement on US government directive to suspend access to Fable 5 and Mythos 5
And the US gov could pull the rug out from under our business at any time? That's confidence inspiring?
davesque··on Statement on US government directive to suspend access to Fable 5 and Mythos 5
I like to think that the long arc of history bends towards greater access to knowledge and intelligence. I mean, isn't that what we all want? To be collectively less ignorant and more aware of how the world works? But I guess that's not what the US gov wants. Crazy times, truly. The mask is really coming off lately.
davesque··on If Claude Fable stops helping you, you'll never know
And they probably don't enforce those restrictions within their own company would be my guess.
davesque··on Only 17% of all 64-bit Integers are products of two 32-bit integers
Honestly, this is a larger portion than I would have expected.
Page 1 of 34Next →