HNHacker News
TopNewBestAskShowJobs

modeless

45,352 karma · joined July 8, 2009

modeless@gmail.com

Previously: Physical Intelligence (robots), Meta, Google, Microsoft

My blog: https://james.darpinian.com/blog/

https://x.com/Darpinian

Creator of See A Satellite Tonight: https://james.darpinian.com/satellites/

Based in Palo Alto

submissionscomments
modeless··on Learning more about Claude's mathematical capabilities
I wonder if Jarred (the Bun guy) just got lucky here, or if he made progress before all the actual mathematicians at Anthropic because they aren't prompting Claude as ambitiously as he is.
modeless··on Study links GLP-1 drugs to bigger jump in women's employment than a degree
The study compares people who actually started GLP-1 treatment to people who say they want to but didn't. It sure seems to me like people who say they want something but don't take action to get it would also be less likely to get jobs, independent of what the thing they want is.
modeless··on Should you stop cracking your knuckles?
This is a bit weird, but I started getting mild wrist pain from computer use in college, and I credit two things with reversing it. One was rock climbing which strengthened my forearms and wrists, and the other was developing a habit of cracking my wrist joints. I stopped rock climbing many years ago but I continue to crack my wrists to this day. The habit causes me to flex and stretch my wrists every so often while I'm working and I really think that has helped prevent the development of repetitive strain injuries from mouse and keyboard use.
modeless··on DeepSeek V4 Flash 0731
OK? Maybe DeepSeek set their price that way, but the majority of these OpenRouter inference providers are in the US, not China, and have no reason to care about CNY at all.
modeless··on What happens if an entire class of workers loses faith in their careers
Many occupations have been made obsolete throughout history, and this article examines none of them. I'm not sympathetic to "this time it's different", I think people are just too lazy to try to find out what actually happened in the past.
modeless··on DeepSeek V4 Flash 0731
This would be more convincing if those providers had converged on a number that was not the exact pricing of DeepSeek themselves. Clearly DeepSeek is setting the price here and without them holding it down I expect increases.
modeless··on DeepSeek V4 Flash 0731
DeepSeek has announced an upcoming "significant increase" in price, so this line may have to move to the right soon. https://api-docs.deepseek.com/quick_start/pricing/
modeless··on Scientists discover Kelvin-Helmholtz Instability on the surface of the Sun
Obviously not based on DNA/RNA. I think it is pretty lacking in imagination to believe that DNA/RNA in liquid water is the only possible configuration of atoms that could make a self-replicating and evolving organism.
modeless··on Scientists discover Kelvin-Helmholtz Instability on the surface of the Sun
I wonder if we will ever discover life inside stars. There's definitely complex stuff going on in there.
modeless··on I'll be stepping back from leading product for X
No, the login wall was first added by pre-Elon management https://news.ycombinator.com/item?id=28289263
modeless··on I'll be stepping back from leading product for X
The login wall predates Elon BTW https://news.ycombinator.com/item?id=28289263
modeless··on I'll be stepping back from leading product for X
Twitter was famously unreliable for practically its entire existence. Don't remember the "fail whale?"
modeless··on I'll be stepping back from leading product for X
X has the best tools for managing your feed of any social network. "Mute words" is a godsend and can practically eliminate any topic you hate (politics of course). The Following tab can give you the exact reverse chronological view of only your friends' posts that everyone claims they want. The "Not interested in this post" button works really well to modify the algorithm to your taste. You have the option to pay to completely remove ads. There's really no excuse to complain about the algorithmic feed these days IMO.
modeless··on LLMs won't break symmetric crypto
Agreed, we have probably seen only the tip of the iceberg on that. I wouldn't want to be holding niche crypto coins right now.
modeless··on LLMs won't break symmetric crypto
I don't really find the "because it's difficult" arguments convincing at all. Especially the one claiming it's hard because it requires designing and running a large number of tests and reasoning about the results of each one. That kind of tedious grinding is exactly where LLMs should shine vs humans!

The only convincing argument here is that these things are battle tested (literally in most cases I would guess), with tons of research that never gets published because it's unsuccessful. A whole lot of human effort has gone into trying to break these things. A lot more than went into any of the math problems AI has solved so far. It's going to take a while before LLMs can equal and surpass that amount of human effort. And they might have to surpass it by many, many times to actually break these, if it is even possible, which is not certain.

modeless··on That time when I failed the Microsoft interview
Yeah I can see people using these as proxies for age discrimination
modeless··on Apple says more ex-employees may have taken confidential data to OpenAI
Apple leaks like a sieve. The extreme secrecy culture is a pointless drag on productivity, maintained long past its usefulness for the sole benefit of the execs practicing their Steve Jobs "one more thing" keynote cargo cult.
modeless··on That time when I failed the Microsoft interview
Uhh, strong reject
modeless··on That time when I failed the Microsoft interview
Oh yeah the estimation problems are terrible. The probability ones are at least notionally trying to test your math skills. But still pretty far from being relevant in almost all cases.
modeless··on That time when I failed the Microsoft interview
That's why the problem statement specifies that the bowling ball sinks.
modeless··on That time when I failed the Microsoft interview
Most things that can fairly be described as a "small dip" should decrease travel time. When I was presented with the problem it came with illustrations to clarify. Surprisingly hard to make it watertight with text alone.
modeless··on That time when I failed the Microsoft interview
Once long ago when I interviewed at Apple I was asked the classic "fork in the road, two guys, one always tells the truth and one always lies" riddle, with complete earnestness as far as I could tell. Possibly the single worst interview question I've ever been asked.

Of course I told the interviewer I'd heard it before and then gave the correct answer. In my case we just ended up chatting about previous experience instead of doing another brainteaser and I ultimately passed the interview, I think. But afterward the recruiter strung me along for weeks telling me they wanted to make an offer but not giving me one, and I ended up going to Microsoft instead.

Not the worst interview experience I've had, though. That would be the time I interviewed for a full time position after an internship and a group of guys who knew me and had worked with me all summer asked me a pointless brainteaser as the only interview question. I crashed and burned for a full 40 minutes in front of them. Humiliating, and pretty much a pointless hazing ritual as they offered me the job anyway. Luckily I got a better offer and was able to turn them down.

Here are some other brainteasers I've been asked in interviews. I actually think these physics based ones are fun (probably because I had no trouble solving them), but they're still terrible interview questions:

You're in a boat on a lake with a bowling ball. After you drop the ball overboard and it sinks to the bottom, is the lake water level higher or lower or the same?

Three balls are on three downward sloping tracks. One track is a straight line down to the end, the second is the same except for a small hill in the middle, and the third is the same except instead of a hill it has a small dip. All tracks start at the same height and end at the same height and cover the same horizontal distance. The balls are released at the same time and roll to the end without leaving their tracks. Which one gets there first?

modeless··on So you want to use plants to reduce CO₂
This guy actually built an algae farm to test how much you need to breathe. https://youtu.be/AAbyUaLN2QA
modeless··on Gemini Robotics 2 brings whole body intelligence to robots
Cost, weight, durability. For cameras, bandwidth.
modeless··on Gemini Robotics 2 brings whole body intelligence to robots
I actually like Tau's videos a lot better than Google's. An agile robot getting out of a car carrying a tote is actually novel and difficult and useful. In contrast, what Google is showing here is mostly glorified pick and place with low success rates on a needlessly complex robot and a voice LLM slapped on top to pre-announce its moves like an anime character, with fancy video editing and bubbly music to distract you from how slow it is.

Sure Tau is teleoperated, but teleoperating an agile movement like that is actually really hard to get right and still involves AI to keep the robot balanced. Tau is a lot closer to real deployment than Google, and when it performs useful tasks it is simultaneously collecting the data to eventually automate those tasks.

modeless··on Our position on open-weights models
I do not. A ban on capable open-weight models for an indefinite period of time falls into the category of bans on open-weight models. If you wanted Anthropic's statement to be true you would need to qualify "ban" or "open-weight models" in the statement, e.g. "permanent ban" or "safe open-weight models".

Edit: Anthropic clearly intended this statement to deflect criticism, but in order to achieve that goal they stretched too far and made a statement which is false. Furthermore, I argue that "open weights" implies an ability to modify model behavior, just as "open source" implies an ability to modify software. If for example some mechanism was found to share floating point numbers that are encrypted in some way so as to allow running a model but disallow behavior modification, that model would not be "open weights", in the same way that releasing obfuscated source code that can be compiled but is designed to resist modification would not qualify as an "open source" release. So I don't really see how any capable model could ever be both "open weights" and "safe" under Anthropic's preferred testing regime, regardless of future research progress.

modeless··on Our position on open-weights models
My point is that advocating a de facto ban on capable open source models is inconsistent with Dario's statement here that "Anthropic has never advocated for a ban on open-weights models." Call a spade a spade.
modeless··on Our position on open-weights models
OK that is a position they could take but my point is that's inconsistent with "Anthropic has never advocated for a ban on open-weights models". What you're describing is a ban on capable open-weights models until some future time.
modeless··on Our position on open-weights models
Any sufficiently capable open weights model would fail "safety" testing though, as any "safeguards" of the sort Anthropic likes could be removed. That's just another way of saying they want a ban on capable open source models which would contradict their earlier statement, or at least make it very misleading. It's hard to see how this post can be internally consistent without some hint from Dario about what he believes should happen to models that fail safety testing and/or how capable open weights models could possibly pass a safety test of the kind he proposes.
modeless··on Our position on open-weights models
> All sufficiently capable models, open and closed, should go through mandatory safety testing

What happens if a model fails the test? Surely one can use Kimi K3 for evil, somehow or other. What now?

"Mandatory safety testing" implies consequences for failing, yet Dario has nothing to say about what the consequences should be. He says he doesn't advocate a ban but it's hard to imagine what his alternative would be if he won't say it.

← PreviousPage 5 of 34Next →