Indeed. How many scientists understand how a compiler works? Or know about branch prediction on the CPU? If they do, do they lose sleep because it's non-deterministic? I don't. Seems fine.
> You don't study mathematics or computer science and information theory to produce commodities. You study them to transform your mind.
As a taxpayer I do not fund CS and math research because it "transforms minds." I fund it because it produces commodities. If you want to transform your mind, whether through math or meditation retreats, you are free to do so, but don't expect research funding to do it.
Really needs a comparison to the "megaprompt" itself (i.e. "here is a tweet, rate it as ironic or not, considering the following properties; explain your reasoning then output your final answer at the end"). I bet that would get you very far towards the logistic classifier, and would generalize much better out of distribution.
But we don't pour billions of dollars of research funding into mountain climbing because we think it's going to lead to wider breakthroughs in science and technology. And when we need to get people on top of a mountain for an important purpose -- like a military or search and rescue operation, for example -- we absolutely do airdrop them right on the top.
So that raises the question: is mathematics simply a pursuit of passion? Are problems solved "because they're there"? If so, then mathematics can join the ranks of things like mountain climbing, cycling, and weight lifting. But if we are trying to accomplish something important (design better airplanes, find theoretical guarantees about cryptography, factor matrices faster), mathematics needs to become more like a military or search and rescue operation, using the best technology available to secure the outcome we need. Given that the NSF pours billions into scientific research every year, it sure seems like mathematicians want to think of themselves as being in the latter category.
Indeed, I wonder how a similar letter by Uber drivers would be received -- "navigation is an intrinsically human domain, personal relationships are critical for passengers and drivers to progress in the world, etc etc." Or doctors, for that matter.
We are all going to have to come to terms with entities more capable than we are, and in many cases, letting the real work be done by the AIs will be the right thing to do. For all the huffing and puffing about the "human touch" in medicine, it will eventually become downright irresponsible to consult only with a human doctor. I am not sure if this is the case in mathematics or not, but if it isn't, that suggests math will be relegated to more of a hobby than a cutting edge scientific discipline.
It's been pretty obvious to me that the Chinese labs are operating mostly on a fast-follow strategy. The distillation attacks are well-documented, and there is good reason to believe they are able to copy architectural innovations as well. If US labs stagnate I would expect Chinese labs to stagnate as well. Their engineering is great, but in terms of frontier innovation (which requires heavy compute to search for new strategies that work at frontier scale) they are very far behind.
A lot of the low level stuff is outsourced to biochemistry: the physical properties of proteins, and the various self-regulating biochemical systems of an animal, can "encode" a lot of intelligence, easing up on the computational demands of the brain proper.
The reason this is a thing is that other people can bid on your product (Amazon has an entire sponsored product category for this reason). So the reason for Seth Godin to bid on ads for "Seth Godin The Knot" is to try to crowd out other people bidding on the same term (eg another book on the same topic). Can be smart but you don't have to play this game if you don't want to, and the ACOS will tell you if it's worth playing. Even if you don't bid on those ads your own product will be the first result. It's just whether you want to bid up the price for competitors -- if you don't bid on your own product someone else can come in and buy the ad spots cheaply.
Crazy how a smart person like this fails to understand the gumbel softmax technique. It does not affect writing quality at all, provably. The very fact that there is generally no "best next token" with 100% certainty is precisely why the trick works (you cannot watermark a response to "respond with the To be or not to be soliloquy from the first folio Hamlet", for precisely this reason).
The default is set for the marginal new user, which at this point is probably not someone like you (who benefits a lot from manual mode) -- it's someone who's more "code-naive" and might get anxious about approving random bash script commands they don't recognize. Safely getting the user from prompt --> first vibe-coded app is the "user journey" now, and since auto mode seems pretty good at not letting Claude rm -rf'ing the home directory, this is 100% the right business move. For people who know what they're doing (like you), manual mode is just a shift-tab away
Surely the two leading candidates must be (a) the model is just not that good, or (b) it is misaligned in a pretty obvious way that can't be swept under the rug.
They already quietly agreed to "all lawful use" with the Pentagon, to no real fanfare. Gemini Slaughterbot Edition, coming soon to a DHS facility near you?
I don't think so. I've used Claude to generate 3D animations in python purely by making and manipulating raw meshes, and I've had great results. A more likely explanation is that 3D graphics is just pretty straightforward matrix algebra, and models (or Claude specifically) has that down cold.
Why not? Slowing down / halting biological weapons research mostly worked. Yes there have probably been modest sized defections here and there, but the pace of bioweapon development is a crawl compared with (a) what is possible with science already, and also with (b) the pace of bioweapons development from 1910 to 1970.
Indeed, I suspect the failure rate of, say, new jet engine designs is rather high as well -- those failures just never get reported in a federal repository, unlike RCTs, since they never make it out of the simulator or the prototyping lab. And we have, comparatively, much better computational models of how airplanes fly than how cancer cells mutate. FWIW this clinical stat is far better than Edison's supposed lightbulb-idea failure rate!
I'm actually surprised it isn't going up over time. That is naively what you would expect as the low-hanging fruit is plucked. So the fact that it's been stable is probably a sign that scientific advances are roughly keeping pace with the (presumably) increasing challenge of finding ever more targets for drugs.
Indeed, the B2B / no-data-retention market is still going to provide plenty of business for American companies even if every hobbyist uses open-weight models.
(1) What does it score on the private test set?
(2) Does this approach generalize to, e.g., Atari or NES games, or is it just hard-coding priors about the games into the model (as Chollet specifically warned was a chronic problem in benchmarks in the original Arc-AGI paper)