HNHacker News
TopNewBestAskShowJobs

modeless

45,359 karma · joined July 8, 2009

modeless@gmail.com

Previously: Physical Intelligence (robots), Meta, Google, Microsoft

My blog: https://james.darpinian.com/blog/

https://x.com/Darpinian

Creator of See A Satellite Tonight: https://james.darpinian.com/satellites/

Based in Palo Alto

submissionscomments
modeless··on Our position on open-weights models
Any sufficiently capable open weights model would fail "safety" testing though, as any "safeguards" of the sort Anthropic likes could be removed. That's just another way of saying they want a ban on capable open source models which would contradict their earlier statement, or at least make it very misleading. It's hard to see how this post can be internally consistent without some hint from Dario about what he believes should happen to models that fail safety testing and/or how capable open weights models could possibly pass a safety test of the kind he proposes.
modeless··on Our position on open-weights models
> All sufficiently capable models, open and closed, should go through mandatory safety testing

What happens if a model fails the test? Surely one can use Kimi K3 for evil, somehow or other. What now?

"Mandatory safety testing" implies consequences for failing, yet Dario has nothing to say about what the consequences should be. He says he doesn't advocate a ban but it's hard to imagine what his alternative would be if he won't say it.

modeless··on Ray tracing massive amounts of animated geometry using tetrahedral cages
Massively simplifying here but when doing one rasterization pass you process a flat list of triangles, in order, once. When raytracing a camera view you query a database of triangles, at least once for each pixel. To avoid scanning the whole list of triangles for every pixel, the database needs a fancy 3D spatial index whose structure depends on the exact position of each vertex. When you modify the position of vertices you have to fix up or rebuild the fancy spatial index too.
modeless··on SpaceX Starship Flight 13 livestream [video]
I imagine they built these a long time ago (anticipating fewer test flight problems) and have probably improved the design since then making these obsolete. They are going to build thousands so they can't be too precious.
modeless··on ARC-AGI Leaderboard
They state the puzzle is "Witness-like" which I assume means that it follows the rules from the well-known puzzle game "The Witness" which Opus definitely knows.
modeless··on SpaceX Starship Flight 13 livestream [video]
Also they plan to catch the ship next time! Excitement guaranteed!
modeless··on SpaceX Starship Flight 13 livestream [video]
Lots of missing tiles still. Best case is it's all infant mortality and replacing just those tiles will result in a durable heat shield for many more flights. Worst case is every flight loses that many tiles and the ablative underneath can't be reflown until replacement.

I wouldn't be surprised if we end up seeing a heat shield refurbishment attachment to the launch tower that allows technicians to access the side of the ship to patch up between flights without leaving the pad. They're going to need some kind of solution for the many back-to-back HLS refueling flights.

modeless··on If coding has been solved, why does software keep getting worse?
Personally I hope to use a lot less software in the future. Anything I want to know or do online I'll just ask an AI and it can wade through the sea of cookie popups or UI redesigns or whatever to get it done.
modeless··on SpaceX Starship Flight 13 livestream [video]
Successful Raptor on-orbit relight demo! That should clear the next flight to go to orbit and deploy operational satellites! Big milestone for the program.

And the ship landing was so soft it didn't even explode! That's never happened before. And it's still connected to Starlink too. The ship's cameras are still working underwater! I wonder if they will be able to recover it intact?

We're getting some great views of the heat shield post-reentry. That's great data for the heat shield team. The heat shield is the biggest remaining question mark in the whole program. Reflight without heat shield refurbishment will be required for the whole architecture to work, and it's never been done before. None of the heat shields so far looked reusable without refurbishment, including this one.

Edit: Spashdown video https://x.com/SpaceX/status/2080804682918764619

modeless··on Fil-C: Garbage In, Memory Safety Out [video]
Fully agreed. At some point you have to draw a line and say that the rest is the responsibility of the kernel and hardware and user, and I think Fil-C drew that line in the right place.
modeless··on Fil-C: Garbage In, Memory Safety Out [video]
I'm not sure about the speed of ASAN, it may be comparable but doesn't guarantee memory safety. It's only for catching mistakes and not secure against an adversary. Valgrind is dramatically slower.
modeless··on Fil-C: Garbage In, Memory Safety Out [video]
The whole point of Fil-C is that it is fast enough to consider using in production for some applications while still guaranteeing memory safety. We already have ASAN and Valgrind and other tools for development purposes, that's not what Fil-C is targeting.
modeless··on Be skeptical of OpenAI's rogue hacker agent story
From what I read the actual escape was through a proxy that allows downloading Python packages from the internet. It's not supposed to allow general internet access but the AI found a previously unknown vulnerability in it. That is hardly "standard and well documented script kiddie methods", nor does it seem like criminally negligent sandbox design, though clearly they will need to reduce their attack surface in the future. I hope they are working on a physical air gap and faraday cage because it seems like it won't be long before it is legitimately required.
modeless··on Fil-C: Garbage In, Memory Safety Out [video]
Adding Fil-C-like runtime checks to Rust is definitely an interesting direction. As I mentioned upthread. It's not just the availability of the safe API that's interesting, though, but also the prohibition on using the unsafe API in the entire program and all dependencies. Which Rust could also do in theory but not yet in practice AFAIK.
modeless··on Fil-C: Garbage In, Memory Safety Out [video]
Better include Rowhammer too. Maybe Fil-C should run a test and refuse to start on any system with bad RAM or unpatched CPU errata. It could also monitor the voltage to protect against undervolting attacks. And you'll need some cosmic ray shielding too.

Of course I'm kidding, but it is absolutely the case that if you truly care about safety you need to consider more than the program source code and binary, but also the environment, including kernel and hardware. The user doesn't care if their web browser got hacked via stack buffer underflow or /dev/exynos-mem or Rowhammer; the result is the same.

modeless··on Fil-C: Garbage In, Memory Safety Out [video]
Fil-C has access to a memory safe API called mmap with a lot of the capabilities of the mmap system call, while Rust's safe subset does not. Rust allows you to use the mmap system call unsafely, while Fil-C does not. I feel these are both big advantages for Fil-C. You'll never have to audit a Fil-C codebase to determine whether its calls to mmap are memory safe, or indeed any other calls to any API or dependency, even those written in C.
modeless··on Fil-C: Garbage In, Memory Safety Out [video]
Rust doesn't runtime validate that your usage of syscalls is memory safe, while Fil-C does. For example you can call mmap in Fil-C and it is still guaranteed to be memory safe, while in Rust you can easily violate memory safety by calling mmap. This seems like an unambiguous improvement to me. This is as memory safe as it is possible to be on a system with a kernel that is not memory safe.

Adding Fil-C-like runtime checks to Rust is definitely an interesting direction.

modeless··on Claude Opus 5
It will continue to be valuable as a cost and speed benchmark long after it is saturated at the high end. And they are already working on ARC-AGI 4 and thinking about going even farther.
modeless··on Claude Opus 5
Opus is cheaper than Fable. They could probably replace Fable with Opus but why? They would be churning customers to different models for no reason. Even if a model scores better on benchmarks it can always regress in your specific use case, and customers don't like that. Customers want to be able to continue using their current model until they decide to upgrade themselves.
modeless··on Claude Opus 5
Yes, I think it indicates real progress in fluid intelligence. Clearly these models are making huge strides in usefulness which are well correlated with their ARC-AGI scores.

I don't think this is benchmaxxing. These companies are locked in a competition to produce the best software engineer, and falling behind is an existential risk. I doubt they are wasting time benchmaxxing ARC-AGI.

modeless··on Claude Opus 5
Wow, 30% on ARC-AGI-3 for $20k total. Huge jump from GPT-5.6's 7.8% at $20k per task. I continue to believe ARC-AGI measures something different and important compared to other benchmarks.
modeless··on Couple pay >$800k for a gene-editing therapy for their daughter. She died.
Removing the "evil" profit motive has always been an argument for socialism. If it would work for healthcare, why not everything?
modeless··on Couple pay >$800k for a gene-editing therapy for their daughter. She died.
Some people thought it was a good idea to cap [administration cost + profit] of health insurance companies as a percentage of premiums. As a result, health insurance administrators have two ways to increase their compensation: they can compete to acquire more customers, or they can artificially inflate costs per customer. Guess which one is easier! Now you don't have to wonder why medication is half the cost when you buy it without insurance...

Love to hear all the time how the "free market" is ruining healthcare. Of course everyone knows a "free market" is when you have profit caps, supply restrictions (residency cap), government-run marketplaces, byzantine tax subsidies tying insurance to employment for no reason, etc etc

modeless··on Launching Health in ChatGPT
I'm sure plenty of people will be aghast at the concept, citing privacy or accuracy concerns. So far I am still alive and well, and have not experienced any issues despite the total destruction of my sacred medical privacy implied by providing my genetic code to a chatbot. I'll keep HN updated should any terrible consequences befall me.
modeless··on Launching Health in ChatGPT
My experience with this was pretty bad initially. It's clear Epic and other players don't want to be disintermediated and they're going to make things as inconvenient and broken as possible. Par for the course in healthcare. I had much better results simply downloading test results and appointment notes and asking ChatGPT about them in regular chats.

I also had a good experience providing my full genetic sequence (raw data downloaded from sequencing.com) to Claude Code and having it do analysis for me. It downloaded a bunch of bioinformatics tools and did a good job answering my questions.

modeless··on Passkeys were invented by engineers with zero understanding of consumer brain
Those complaints aren't Google arbitrarily or maliciously locking accounts. They are mostly people who forgot their password and also simultaneously lost their sim card and phone number and recovery email and backup codes etc etc. Google provides excellent features to prevent this from happening and I have already taken advantage so that I won't simultaneously lose access to all my recovery options. You also have no way of knowing how many of those complaints are from scammers trying to compromise accounts, and if you think that isn't a gigantic problem in support forums you are fooling yourself.

You don't have any accurate data to support your suspicion that Google is worse than others, only anecdotes and vibes. The true source of your fear is a generalized mistrust of big tech relative to other institutions, which while common and popular these days is not a sentiment I share.

modeless··on Passkeys were invented by engineers with zero understanding of consumer brain
> Basing your security decisions on "it'll never happen to me" because there are billions of other users who will get burnt first

That's not my argument. My argument is that your evidence that Google is uniquely bad at locking people out of their accounts is not good evidence. A few stories in the news represents a beyond negligible fraction of billions of user accounts.

modeless··on Passkeys were invented by engineers with zero understanding of consumer brain
This is very far from the top of the list of risks I should be concerned about. Basing your security decisions on the frequency with which you hear about something bad happening in the news is not a wise strategy. You have to consider that the user base of Google accounts is in the billions.
modeless··on Terence Tao's ChatGPT conversation about the Jacobian Conjecture counterexample
"it's enough of partial results. let's finish with a complete unconditional counterexample"

"Worked for 88m 24s... >"

"<h1>Complete finite counterexample</h1>"

...

modeless··on Passkeys were invented by engineers with zero understanding of consumer brain
I find Google Password Manager makes passkeys pretty easy to use. As long as you don't accidentally create a passkey some other way. Hopefully websites will adapt to the reality of how people use passkeys in practice and some of the UX weirdness around them will disappear over time.

One annoying thing though is that while they recently added password sharing, they don't allow sharing passkeys. Basic passkey sharing would be nice, but it also seems possible to implement fancy sharing features that wouldn't be possible with password sharing. Things like sharing one time use passkeys or time limited passkeys or limited access passkeys or secure revocation of shared passkeys. I hope people are thinking about this.

← PreviousPage 6 of 34Next →