HNHacker News
TopNewBestAskShowJobs

gr_norm

860 karma · joined May 10, 2023

submissionscomments
gr_norm··on Pacing the frontier
Sorry, I should've used clearer language. I just mean intelligence in the sense humans possess it. Computers have always been able to exceed limited elements of human intelligence (say, at arithmetic), but the problem is to simultaneously match or exceed all of them. I specifically think that the important parts they currently lack make them a no-go for supposed existential risk.
gr_norm··on Pacing the frontier
Moving the goalposts, so to speak, is how science works! We must of course update the things we think based on an improved understanding of how the world works. Only the deeply incurious could consistently demand that one adhere dogmatically to a prior way of understanding the world in the face of changing knowledge. It's important not to treat science as a game to be won (of which 'goalposts' are evocative), but as the collaborative effort that it is.

Note well that I have not said LLMs are not useful, only that no amount of interaction with them has convinced me they remotely approach general intelligence. Their actual function of heuristically regurgitating all information ever known to mankind remains extremely useful in lots of domains.

And by the way, it's not obvious an LLM could fool me (or the average person) over a sufficiently long period of time, exactly because they are heuristic machines. That they lack thought guided by underlying cognitive models seems to always leak out, in the end. I suspect pass rates by LLMs on long-range Turing tests (if there have been any conducted) might drop as people become acclimatized to them.

gr_norm··on Our position on open-weights models
Not only is the DNA sequence for smallpox available, but since 2018 there's been a well-documented end-to-end synthesis procedure for the very closely related horsepox virus [1]! No LLMs needed. Caused quite a stir in the synthetic biology community back then.

The world has not come to an end, of course, because even with peer-reviewed and experience-driven (rather than hallucinated and therefore dangerous) instructions detailing obstacles encountered during synthesis and how to overcome them, actually going out and acquiring the materials and ability to use them sufficiently skillfully is another matter entirely. Biosecurity is an important topic, to be sure, but what the AI labs have to say about it (or anything) at this point does not necessarily survive contact with reality.

[1] https://journals.plos.org/plosone/article?id=10.1371/journal...

gr_norm··on Our position on open-weights models
The effective altruism/rationalism/AI xrisk people have always had a shockingly poor grasp on subjects outside computer science, despite their attempts to speak on them. I don't blame the actual biologists and chemists working at the frontier labs for wanting to skim a few bucks off all the money flying around, though! I know a couple who've had not-so-kind words to say about their employers' intelligence.

I suspect there's at least some "telling the bosses what they want to hear" going on. A massive financial incentive exists to exaggerate and fearmonger even internally to the company, because it makes you and your job seem more important.

gr_norm··on Our position on open-weights models
The emptiness of AI companies' waxing poetic about the future of humankind is laid bare by simply looking at what they actually do, and who they do business with. Actions speak louder than words, and they've driven the worth of their words into the dirt many times over.
gr_norm··on Our position on open-weights models
It's baffling that they thought the mental gymnastics in this blog post would make them look better. I'd rather they simply fall silent on the issue; I would respect them more (or at all) for it. Open models obviously threaten fierce competition, if not outright destruction of their bottom line. But no, they needed to try and argue that they have the moral high ground for attempting to singularly consolidate power over all human labor.
gr_norm··on Our position on open-weights models
> Distillation does not allow the CCP to obtain equivalent or superior AI capabilities to the US, but it can bring the Chinese frontier to within a few months of the US frontier

A message to their investors, it would seem. "They caught up just because they distilled! Obviously they couldn't actually be as good as us!" Really funny thing to say right after an OpenAI higher-up stated point-blank that the performance of K3 can't be chalked up to mere distillation of American models.

gr_norm··on Our position on open-weights models
I don't really understand how they can argue the security angle with a straight face. It's not like GLM 5.2 is a slouch. I've seen it do things like exploit an IDOR issue when I was experimenting with a quick-and-dirty web automation task. I simply fixed it, as one does. Open models make the world better to a far greater degree than they set it aflame.

Their position is analogous to trying to, say, ensure digital privacy for everyone not by making encryption freely available (because that would let the bad guys use it!), but by making it so you can't use general purpose communications devices that can listen to transmissions not intended for you. Do they hear how moronic that sounds?

Each passing frontier-level open model release makes Anthropic's patronizing rhetoric a little more insufferable, because it becomes clearer how unmoored from reality they've become in pursuit of profit.

gr_norm··on Our position on open-weights models
> Open-weights models that don’t have dangerous capabilities are a public good

Note the hedging against 'dangerous capabilities'. Undoubtedly, all the useful ones trigger this condition in Anthropic's eyes. The rest of the post is filled with similar weasel-wording. Make no mistake, this absolutely confirms that Anthropic is against open models in the sense that any reasonable person understands them.

The way the rest of the post unabashedly appeals to the current US administration's China hysteria is hilarious, and not at all subtle.

I guess we'll see about all the doomsaying here, won't we? Kimi K3 is frontier-level, and there's no stopping it now. As far as the world is concerned, anyway. If the US wants to kneecap itself that's another matter.

gr_norm··on We have proof automation now
> I’m not even going to get into how you could provably transform brute force propositional logic into efficient algorithms.

> The entire problem is showing that an efficient/reliable program actually implements those rules.

Reliability is a standard matter of correctness and captured (partly) by specifications. Efficiency tends to be easy to empirically test, but it is also possible to capture at the specification level [1]. Mind that specifications need not be all-consuming.

> But I did spend a grad class with rocq (coq at the time) and a decade working with “systems engineers” and am not convinced that this is a realistic expectation.

Agree! But this stuff just got massively more accessible, and the tooling around it is growing quickly. I think we'll end up growing specification systems specific to various domains which will be palatable to those "systems engineers", but I err on the side of optimism here. There's definitely a lot left to do for practicality.

[1] See the work of https://cs.nyu.edu/~shw8119 for the case of provably-efficient parallelism and garbage collectors

gr_norm··on We have proof automation now
It is true that finding the correct specification is a formidable task; knowing what correctness even means is arguably most of the difficulty of programming. However! "Moving bugs up from programs to types" isn't how this shakes out in practice, at all. Another commenter already noted that it's often much easier to communicate your intent through specifications, because you can essentially always say what a computation should do much more simply than you can say exactly how to do it.

I think it's also important not to miss the forest for the trees: even relatively simple specifications like "the compress and decompress functions must be inverses for all inputs" rules out vast classes of bugs in a compression library. This is not a complete specification; for instance, it does not speak about how the decompressor behaves on malicious input. But in my experience, even partial specifications carry the promise of hitting warp speed with LLMs in a way that I haven't seen anywhere else. After a certain level of specification, you have decent guarantees of being able to whole-heartedly forget about the implementation details of the synthesized program. And you get a better-built, more robust program out of it at the end!

The comment at the end of the article about having LLMs directly generate assembly against specifications and letting them rip with finding custom optimizations is the sort of crazy stuff this enables. I really think we're only seeing the tip of the iceberg here. People keep asking what we can do with LLMs that we couldn't before; this is the answer.

gr_norm··on Nvidia, Microsoft, Meta warn against overregulating open-weight models
To maximize personal influence and wealth, of course. These days, many with power don't seem to care much about uplifting society so long as they get theirs.

Hopefully Anthropic eats crow here, lest their wish is fulfilled that we all become slaves to the anointed few who work there. Never getting another dollar from me after pulling these stunts.

gr_norm··on Be skeptical of OpenAI's rogue hacker agent story
Many who rush to the defense of AI companies' marketing departments seem to take criticism personally, as if not buying into it all hook, line, and sinker is an affront to them. A foreseeable consequence of becoming cognitively dependent on LLMs.

And for what it's worth, I'm not an AI skeptic. I fully believe that the frontier models are capable of exploiting (chains of) vulnerabilities, having seen GLM and now Kimi do it myself. What I find no reason to accept is the sci-fi existential risk subtext peddled by the salesmanship around it. We will reach a new equilibrium with more secure software, and LLMs (by finding vulnerabilities, generating proofs, etc) will help us along.

gr_norm··on Startup founders urge U.S. government not to shut off Chinese open weight AI
Login-walled for me. Is there a Reddit proxy like Nitter?
gr_norm··on OpenAI’s accidental attack against Hugging Face is science fiction that happened
I could fully see them thinking the incident disclosed yesterday would have made them look good ("wow, OpenAI's models are so capable!"). That it didn't occur to them to discuss specific preventative measures to be taken in the future (airgapping as a foolproof one already familiar to the CTF world, anyone?) indicates to me they're not taking their job seriously; they are the ones treating this as a marketing charade.

It's very difficult for me to reconcile belief in the existential risk business with what they actually did. So I agree with you that this makes OpenAI look badly incompetent; but their communication on this makes me think they don't realize it.

For what it's worth I don't agree with the xrisk-ness of these models; they're dangerous, but almost certainly only temporarily while a new equilibrium is reached via more secure software. Open models are probably an essential part of the recipe (as you noted) for doing so. I also have a personal suspicion that LM-accelerated formal verification will have no small role to play here, sidestepping the cat-and-mouse game of bug finding-and-fixing.

gr_norm··on Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
Yeah, open weights hasn't fully caught up yet, but it's getting very close. And it's certainly passed the point of being reasonably interchangeable with the frontier for (programming) work. Add in the benefits of not being rug-pulled by the frontier labs silently messing with, the knobs on their models or outright denying you the ability to do certain kinds of work (c.f. the HuggingFace fiasco), and they probably come out ahead in several respects.
gr_norm··on Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
The idea that the Chinese labs cannot make progress except by copying superior American products is just prejudice against the former and exceptionalism of the latter at play. Even the OpenAI top brass have admitted otherwise [1]. China is an equal match in every respect, and we'd better admit this to ourselves sooner rather than later so as to see the game clearly.

[1] https://xcancel.com/deanwball/status/2078133895766114412

gr_norm··on OpenAI and Hugging Face address security incident during model evaluation
HF need not be party to it at all, beyond being the victim. I suspect the hack is real; I have observed GLM 5.2 being able to discover similar vulnerabilities in web applications I'm hosting (which I've then fixed!). At the same time, it seems very neatly timed at an inflection point in the conversation around open models, and there's questions around the incompetent isolation under which the hacking benchmark appears to have been run.

Remember that there is generational wealth on the line for most OpenAI employees, and consider what people might do to obtain it.

gr_norm··on OpenAI and Hugging Face address security incident during model evaluation
The timing after the release of GLM 5.2 and Kimi K3 is quite convenient, too, as an angle for regulatory quashing of open-weights models just as they're entering the mainstream conversation around usurping the American frontier labs. I accept my thinking here is conspiratorial, but there's also a hell of a lot of money on the line to encourage the unscrupulous.
gr_norm··on OpenAI and Hugging Face address security incident during model evaluation
They've been saying so from the beginning, and yet did not take the basic precaution of airgapping their off-the-leash model while it's been instructed to succeed at a hacking benchmark by any means necessary. So which is it? I _want_ to believe them, I do, but there's always these gaps between what they say and their actions on display that give me reason to think otherwise.
gr_norm··on OpenAI and Hugging Face address security incident during model evaluation
Yeah, agree on all counts. I'd give them leeway if they were still scrappy startups, but they have entire countries' worth of resources at their disposal and the best of the best on their payroll. No excuses at this point for oopses like this, I would think.
gr_norm··on OpenAI and Hugging Face address security incident during model evaluation
Why is a machine running these sorts of hacking benchmarks not airgapped? That seems a basic precaution, if OpenAI believes what they're selling. I mean, stuff like this is done for CTFs played by humans, too, to rule out collateral damage; it's not some new concept. So this is either thorough incompetence by OpenAI, a marketing piece, or both.
gr_norm··on Advertise in ChatGPT
This is already detestable given how much they're charging, but how long before LLM ads become difficult to distinguish from the main output? Remember how Google ads started out? And for all their faults, Google _after_ 20 years of maximizing-shareholder-value disease is a far more trustworthy company than OpenAI _today_. How long do you trust the ethical instincts of the OpenAI leadership to leave money on the table?
gr_norm··on Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Yeah. It's something I can do myself in a couple seconds if I want, also on more varied SVG scenes. If this is going to be a benchmark people turn to I'd like to see more effort put into it than just a one-sentence prompt.
gr_norm··on Who's afraid of Chinese models?
Yup, fair's fair. Anything else stinks of 'rules for thee but not for me' (a maxim the frontier labs seem worryingly happy to apply, on several counts).
gr_norm··on NSF slashes research programs to support new tech initiative, insiders say
Many words expended here to avoid asking an answerable question. Which complaints? Must I recapitulate the last 70 years of history of politics around science in the United States? You have fixated on a singular line from the article rather than reply substantively to anything I've said (which indeed responds to such complaints inasmuch as this is possible without hearing a specific one), because you have little interest in reaching a shared understanding of its subject material (you do not remotely care about any such historical complaints, or I would've heard one by now) relative to your desire to advance the culture war, as laid bare by your other comments. This sort of low-trust behavior, exacerbated by pretending otherwise, is typical of your unfortunate mindset.

The tactics of discourse you employ are tired and show no signs of improving. You've unfortunately lost the courtesy of a further response from me.

gr_norm··on NSF slashes research programs to support new tech initiative, insiders say
Sorry, I'm not in the business of responding thoughtfully to low effort questions slung rapid-fire over the fence, which moreover take nothing I've said into account. Case in point, this comment; it seems innocuous but would take an extreme amount of effort on my part to reply to what is, I suspect, something you've made up your mind about.

I think it's clear from your other comments cited here that you're exactly who I'm talking about: a gullible soldier of the culture war who has perfected the art of wasting endless quantities of time arguing online.

gr_norm··on NSF slashes research programs to support new tech initiative, insiders say
That is an enormous budget cut; it is exactly an evisceration. And the general trend is to cut both science and its application. Trials for life-saving treatments _by biotech companies attempting to commercialize them_ are being halted or cancelled due to the science cuts and general corruption.

Many people have many complaints; many should be ignored. There's lots of money (and indeed, huge market incentive) to commercialize potentially successful science, and as such it has been done consistently. Curiosity-driven science must feed it, and the idea that it does not unambiguously benefit society at large, with extreme and breathtaking return on investment, is a fantasy perpetuated by those susceptible to the idiotic culture war.

gr_norm··on NSF slashes research programs to support new tech initiative, insiders say
You and people you know will lead worse and less fulfilling lives due to this. You, or someone you love, will likely die of causes that would have been preventable without this destruction of domestic science. Academic culture undoubtedly needed reform, but this evisceration bent on shortsighted retribution will help absolutely nobody.
gr_norm··on Is Meta destroying its engineering organization?
I wouldn't normally reply again, but please try and take my arguments in the good faith and appeal to the heart and humanity that they were given. I think you're quite capable of reading what I've written and responding to it coherently; consequently, I also think you must be aware that you haven't done this in the heat of the disagreement. It's only if we're charitable, not cynical, about this stuff that the world may improve. Thanks. :)
← PreviousPage 3 of 5Next →