45,352 karma · joined July 8, 2009
Previously: Physical Intelligence (robots), Meta, Google, Microsoft
My blog: https://james.darpinian.com/blog/
https://x.com/Darpinian
Creator of See A Satellite Tonight: https://james.darpinian.com/satellites/
Based in Palo Alto
The only convincing argument here is that these things are battle tested (literally in most cases I would guess), with tons of research that never gets published because it's unsuccessful. A whole lot of human effort has gone into trying to break these things. A lot more than went into any of the math problems AI has solved so far. It's going to take a while before LLMs can equal and surpass that amount of human effort. And they might have to surpass it by many, many times to actually break these, if it is even possible, which is not certain.
Of course I told the interviewer I'd heard it before and then gave the correct answer. In my case we just ended up chatting about previous experience instead of doing another brainteaser and I ultimately passed the interview, I think. But afterward the recruiter strung me along for weeks telling me they wanted to make an offer but not giving me one, and I ended up going to Microsoft instead.
Not the worst interview experience I've had, though. That would be the time I interviewed for a full time position after an internship and a group of guys who knew me and had worked with me all summer asked me a pointless brainteaser as the only interview question. I crashed and burned for a full 40 minutes in front of them. Humiliating, and pretty much a pointless hazing ritual as they offered me the job anyway. Luckily I got a better offer and was able to turn them down.
Here are some other brainteasers I've been asked in interviews. I actually think these physics based ones are fun (probably because I had no trouble solving them), but they're still terrible interview questions:
You're in a boat on a lake with a bowling ball. After you drop the ball overboard and it sinks to the bottom, is the lake water level higher or lower or the same?
Three balls are on three downward sloping tracks. One track is a straight line down to the end, the second is the same except for a small hill in the middle, and the third is the same except instead of a hill it has a small dip. All tracks start at the same height and end at the same height and cover the same horizontal distance. The balls are released at the same time and roll to the end without leaving their tracks. Which one gets there first?
Sure Tau is teleoperated, but teleoperating an agile movement like that is actually really hard to get right and still involves AI to keep the robot balanced. Tau is a lot closer to real deployment than Google, and when it performs useful tasks it is simultaneously collecting the data to eventually automate those tasks.
Edit: Anthropic clearly intended this statement to deflect criticism, but in order to achieve that goal they stretched too far and made a statement which is false. Furthermore, I argue that "open weights" implies an ability to modify model behavior, just as "open source" implies an ability to modify software. If for example some mechanism was found to share floating point numbers that are encrypted in some way so as to allow running a model but disallow behavior modification, that model would not be "open weights", in the same way that releasing obfuscated source code that can be compiled but is designed to resist modification would not qualify as an "open source" release. So I don't really see how any capable model could ever be both "open weights" and "safe" under Anthropic's preferred testing regime, regardless of future research progress.
What happens if a model fails the test? Surely one can use Kimi K3 for evil, somehow or other. What now?
"Mandatory safety testing" implies consequences for failing, yet Dario has nothing to say about what the consequences should be. He says he doesn't advocate a ban but it's hard to imagine what his alternative would be if he won't say it.