860 karma · joined May 10, 2023
Note well that I have not said LLMs are not useful, only that no amount of interaction with them has convinced me they remotely approach general intelligence. Their actual function of heuristically regurgitating all information ever known to mankind remains extremely useful in lots of domains.
And by the way, it's not obvious an LLM could fool me (or the average person) over a sufficiently long period of time, exactly because they are heuristic machines. That they lack thought guided by underlying cognitive models seems to always leak out, in the end. I suspect pass rates by LLMs on long-range Turing tests (if there have been any conducted) might drop as people become acclimatized to them.
The world has not come to an end, of course, because even with peer-reviewed and experience-driven (rather than hallucinated and therefore dangerous) instructions detailing obstacles encountered during synthesis and how to overcome them, actually going out and acquiring the materials and ability to use them sufficiently skillfully is another matter entirely. Biosecurity is an important topic, to be sure, but what the AI labs have to say about it (or anything) at this point does not necessarily survive contact with reality.
[1] https://journals.plos.org/plosone/article?id=10.1371/journal...
I suspect there's at least some "telling the bosses what they want to hear" going on. A massive financial incentive exists to exaggerate and fearmonger even internally to the company, because it makes you and your job seem more important.
A message to their investors, it would seem. "They caught up just because they distilled! Obviously they couldn't actually be as good as us!" Really funny thing to say right after an OpenAI higher-up stated point-blank that the performance of K3 can't be chalked up to mere distillation of American models.
Their position is analogous to trying to, say, ensure digital privacy for everyone not by making encryption freely available (because that would let the bad guys use it!), but by making it so you can't use general purpose communications devices that can listen to transmissions not intended for you. Do they hear how moronic that sounds?
Each passing frontier-level open model release makes Anthropic's patronizing rhetoric a little more insufferable, because it becomes clearer how unmoored from reality they've become in pursuit of profit.
Note the hedging against 'dangerous capabilities'. Undoubtedly, all the useful ones trigger this condition in Anthropic's eyes. The rest of the post is filled with similar weasel-wording. Make no mistake, this absolutely confirms that Anthropic is against open models in the sense that any reasonable person understands them.
The way the rest of the post unabashedly appeals to the current US administration's China hysteria is hilarious, and not at all subtle.
I guess we'll see about all the doomsaying here, won't we? Kimi K3 is frontier-level, and there's no stopping it now. As far as the world is concerned, anyway. If the US wants to kneecap itself that's another matter.
> The entire problem is showing that an efficient/reliable program actually implements those rules.
Reliability is a standard matter of correctness and captured (partly) by specifications. Efficiency tends to be easy to empirically test, but it is also possible to capture at the specification level [1]. Mind that specifications need not be all-consuming.
> But I did spend a grad class with rocq (coq at the time) and a decade working with “systems engineers” and am not convinced that this is a realistic expectation.
Agree! But this stuff just got massively more accessible, and the tooling around it is growing quickly. I think we'll end up growing specification systems specific to various domains which will be palatable to those "systems engineers", but I err on the side of optimism here. There's definitely a lot left to do for practicality.
[1] See the work of https://cs.nyu.edu/~shw8119 for the case of provably-efficient parallelism and garbage collectors
I think it's also important not to miss the forest for the trees: even relatively simple specifications like "the compress and decompress functions must be inverses for all inputs" rules out vast classes of bugs in a compression library. This is not a complete specification; for instance, it does not speak about how the decompressor behaves on malicious input. But in my experience, even partial specifications carry the promise of hitting warp speed with LLMs in a way that I haven't seen anywhere else. After a certain level of specification, you have decent guarantees of being able to whole-heartedly forget about the implementation details of the synthesized program. And you get a better-built, more robust program out of it at the end!
The comment at the end of the article about having LLMs directly generate assembly against specifications and letting them rip with finding custom optimizations is the sort of crazy stuff this enables. I really think we're only seeing the tip of the iceberg here. People keep asking what we can do with LLMs that we couldn't before; this is the answer.
Hopefully Anthropic eats crow here, lest their wish is fulfilled that we all become slaves to the anointed few who work there. Never getting another dollar from me after pulling these stunts.
And for what it's worth, I'm not an AI skeptic. I fully believe that the frontier models are capable of exploiting (chains of) vulnerabilities, having seen GLM and now Kimi do it myself. What I find no reason to accept is the sci-fi existential risk subtext peddled by the salesmanship around it. We will reach a new equilibrium with more secure software, and LLMs (by finding vulnerabilities, generating proofs, etc) will help us along.
It's very difficult for me to reconcile belief in the existential risk business with what they actually did. So I agree with you that this makes OpenAI look badly incompetent; but their communication on this makes me think they don't realize it.
For what it's worth I don't agree with the xrisk-ness of these models; they're dangerous, but almost certainly only temporarily while a new equilibrium is reached via more secure software. Open models are probably an essential part of the recipe (as you noted) for doing so. I also have a personal suspicion that LM-accelerated formal verification will have no small role to play here, sidestepping the cat-and-mouse game of bug finding-and-fixing.
[1] https://xcancel.com/deanwball/status/2078133895766114412
Remember that there is generational wealth on the line for most OpenAI employees, and consider what people might do to obtain it.
The tactics of discourse you employ are tired and show no signs of improving. You've unfortunately lost the courtesy of a further response from me.
I think it's clear from your other comments cited here that you're exactly who I'm talking about: a gullible soldier of the culture war who has perfected the art of wasting endless quantities of time arguing online.
Many people have many complaints; many should be ignored. There's lots of money (and indeed, huge market incentive) to commercialize potentially successful science, and as such it has been done consistently. Curiosity-driven science must feed it, and the idea that it does not unambiguously benefit society at large, with extreme and breathtaking return on investment, is a fantasy perpetuated by those susceptible to the idiotic culture war.