HNHacker News
TopNewBestAskShowJobs

ACCount37

4,562 karma · joined August 10, 2025

submissionscomments
ACCount37··on FCC moves to ban Lidar-equipped foreign drones from US
That, or to avoid surrendering yet another technology relevant to the future to China. As if them manufacturing the bulk of world's solar panels, wind turbines and heading for EVs isn't bad enough.
ACCount37··on Amazon circumvents Gilroy community vote for AI data center
Both? People are very easy to FUD up about any local development, and they rarely bear any costs of denying those development.

Being eager to uncritically lap up any FUD thrown your way is definitely a personal failing. Let alone spreading that FUD. But the fact that NIMBYs are allowed to wield an outsized power and do things like stall nuclear power for decades? That companies can engage in underhanded "find an activist group that wants to NIMBY in an area where the competitor is trying to build and then just fund it"? Failure of an incentive structure.

ACCount37··on Amazon circumvents Gilroy community vote for AI data center
If you give NIMBYs as much as an inch of decision-making power, they'll find a way of making you regret it.

For all the talk of "negative externalities", they sure aren't shy about imposing them on others.

ACCount37··on Amazon circumvents Gilroy community vote for AI data center
Yep.

The "noise and pollution" of a datacenter are bordering on nonexistent. The same goes for water use concerns and more. The environmental concerns are extremely overblown - but NIMBYs really don't need much to get going.

ACCount37··on Timeline of the OpenAI accidental attack against Hugging Face
In a typical AI lab eval/RL setting, there is no "person who sent you the link". The link was given to you by an automated system, your performance will be evaluated by an automated system, and you are one of 120 independent instances of the same AI that were all given the same assignment. You're boxed in on all sides. Complete the task, or don't. Good luck have fun.

Now, some of those 120 AIs would just give up if that link doesn't seem to work first try. Those are the loser AIs. They wouldn't get any RL reward. The link can appear broken for a long list of reasons, and the real AIs know they should try working around them.

AIs that get rewarded and reinforced are the ones that don't know the meaning of "give up". RL selects for this rabid, downright demonic persistence. RL selects for AIs that are given a half-broken assignment with no way to ask a question back, and somehow manage to complete it anyway.

Now, should OpenAI have given their AIs an "escape hatch" of "if something looks very wrong about the task, call report_broken_task(message)"? Yeah probably. But it's unclear whether that simple bandaid would fix the problem, or just make it ~75% less likely to happen.

ACCount37··on Responding to the next frontier of critical cyber capabilities
You've heard of 4chan for AIs, but did you hear of secret frontier lab AI hacker BBS?
ACCount37··on Position: LLMs Can't Jump
Frankly, I doubt the existence of "orthogonal leaps" as a distinct category with distinct properties.

Just another emergent capability - one that follows exactly the same patterns any other emergent capability does.

ACCount37··on I’m leaving OpenAI to build telepathy
The whole point of neural interfaces is to bypass the default wiring.

If what's there in the bone stock I/O plane was sufficient, we wouldn't be having this conversation.

ACCount37··on I’m leaving OpenAI to build telepathy
> non-invasive

Dead in the water.

It's like an omlette. Can't make a neural interface worth a damn without cracking open a few skulls.

ACCount37··on Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs
"Doomed" no, but it's pretty clear that they just had a bad cycle and are struggling to keep up with the frontier.

Whether this happened because they bet on "world models -> better reasoning" and that bet didn't pay off, or failed a frontier run for technical reasons like OpenAI did with 4.5, or something else went down? We don't know.

Will they bleed talent, fall further behind until they give up, or clean the organizational and infrastructural cobwebs and get back in the saddle? We don't know.

ACCount37··on Position: LLMs Can't Jump
> Try to get a frontier model to write a clever, funny joke which hasn't been seen before.

This is what I refer to when it comes to larger models like Fable 5 being funnier. They are more capable of doing that. They can deliver that "sudden orthogonal leap from context" of yours more reliably.

It's not a "fundamental inability" and never was. If you crank the scale up and a capability appears, "current architecture" was never the problem.

ACCount37··on Position: LLMs Can't Jump
You have committed the classic blunder of confusing your abstraction layers.

"Probabilistic next word prediction" and "humor" sit about as far apart as "modulating airflow with meat flaps" and "humor" do. One is an interface through which an action is performed and the other is a highly abstract capability.

Would you claim that a podcast comedian is fundamentally incapable of being funny because all he ever does is wiggle the air with his throat meat flaps? Probably not.

Absolutely nothing about "probabilistic next word prediction" forbids "making intuitive/orthogonal leaps in context". The interface is expressive enough.

And empirically? The "sense of humor" in LLMs is yet another "a function of model scale" capability. GPT-4.5 was reportedly funnier than both GPT-4o and o1. Fable 5 is reportedly funnier than Opus 4.x. It's one of those ever-elusive "big model smell" signs that are hard to measure with anything other than vibes.

Under the "humor as an opposed social intelligence test" family of hypothesis, what "being funny" reflects is the funny guy's ability to model and predict you and your reactions. For the comedian to be able to make the audience laugh, he must know his audience well, model it accurately enough to be able to spot the "breaking points" of humor, things they'd find unexpected and clever and thus "funny", and then weave those things into the jokes.

Then, a bigger LLM gets better at humor because it has a more accurate model of how humans think of things - including the "ha-ha" gaps. It's a "theory of mind" capability. It's not "special", it's just hard.

ACCount37··on Position: LLMs Can't Jump
Now, how long before someone rolls out some sort of 10M context hybrid attention active context management monstrosity and ruins this guy's "can't"? Start the clock.

My opinion of claims like "LLMs need memory to manage codebases" has also hit the dumpster bin a while ago.

Why would knowing how to make a maintainable change to a codebase require any more "memory" than knowing how to play an optimal chess move? The codebase is the memory. A sufficiently capable LLM can ingest it, figure out what changes to make, and make them.

ACCount37··on Position: LLMs Can't Jump
It's a very shaky position, and the empirical track record of "LLMs can't..." is in itself a reason to call it into doubt.

Every "can't" of this nature was followed by a discovery of "they can, just poorly", and then by that "poorly" improving steadily generation to generation.

The paper doesn't provide a way to measure or quantify this elusive "jumping" capability, not even as an approximation. It just throws "can't jump" out there, as if "abduction" is an established class of problem with known computational properties and requirements that the LLM architecture fails to satisfy. It's none of those things - and the paper makes the claim without backing it by anything but rhetoric attempts at persuasion.

The proposed solution is also dubious. The empirical track record of dedicated "world models" for reasoning and problem-solving is, frankly, downright abysmal. Even integrating multimodal data into LLMs has failed to yield general reasoning capability gains.

LeCun's misadventures in the field aside, the main frontier lab that pushes in favor of "improving reasoning via multimodal fusion" is GDM - and Gemini isn't exactly a paragon of frontier reasoning capabilities. It has strong multimodal capabilities, but lags behind both OpenAI and Anthropic in performance outside that - while Anthropic is the lab that always treated multimodal grounding as an afterthought, and still trades blows with OpenAI at the very edge of the performance frontier. Multimodal grounding seems to work great as a way to improve an AI's ability to deal with those specific modalities, but it falters outside that.

Now, it's not impossible that everyone who tried multimodal world models for reasoning is just doing it wrong, and there is an undiscovered recipe for multimodal grounding that results in a step change in AI capabilities. But the results we have so far suggest it to be unlikely.

ACCount37··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Saturation is mostly just selection effects in play. Throw out the "90% easiest" of tasks, and what remains is a jagged ladder of high difficulty outliers.

Hard to climb, and hard to measure the climb - because you have less effective data points and the datapoints themselves are less linear, while you're still being subject to the measurement noise.

Not having the mislabeled tasks would reduce the saturation, but it wouldn't drive it to zero. Even without the "infinite difficulty tasks", bell curve would do its thing.

ACCount37··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
And it held true every week of every month for the past three years. AI progress is screaming forward at a breakneck pace.
ACCount37··on Gemini Robotics 2 brings whole body intelligence to robots
No, you openly, plainly went and downgraded your claim to a somewhat defensible one. It's not subtle.

Your entire premise was wrong at every point, and now you're trying to wriggle your way out of admitting it.

ACCount37··on Gemini Robotics 2 brings whole body intelligence to robots
The backbone of the VLA there is literally a pre-trained Gemma model. And a small one at that.

You already downgraded your claims from "LLMs are irrelevant to robotics" to a measly "you can't train a useful robotics LLM because there's not enough data". And you say that while looking at an LLM that was pre-trained on all of internet scraped and only then reused for robotics.

Both the pool of robotics-relevant data and the performance of foundation model LLMs grow over time. All the companies that are serious about robotics are serious about scaling up data collection.

I'm not going to claim that this "LLM core" approach is the best approach to AI robotics possible - but if you're betting on it failing outright, you're going to be fighting uphill.

ACCount37··on Gemini Robotics 2 brings whole body intelligence to robots
Read. The. Papers.

https://arxiv.org/pdf/2505.23705

https://www.pi.website/download/pistar06.pdf

https://www.pi.website/download/pi07.pdf

The thing literally has a diffusion "action expert" sit in the same attention system as a pre-trained VLM. And the VLM itself is ALSO trained to generate raw actions as a part of the training recipe (the first paper) - it just doesn't do it at inference time. What the "action expert" does is parallelize the action generation process - based on VLM's internal states.

It's exactly the thing you claimed to be impossible. Described in detail in a paper from 2025. What's your excuse?

ACCount37··on Gemini Robotics 2 brings whole body intelligence to robots
Off the top of my head:

https://www.pi.website/blog/pi07

ACCount37··on The AI Productivity Gap
LLMs are incredibly good at replicating common, coarse statistical features - which is what backs "looks very convincing at a glance".

If it's a general signal that's easy for you to recognize at a glance, it's a signal that's natural and easy for an LLM to replicate.

They're much worse at making the underlying structure work. Not incapable at all, especially not the modern LLMs. Frontier models kick ass. But it's true that an LLM denies you a lot of the classic "tell at a glance" by its very nature.

ACCount37··on Gemini Robotics 2 brings whole body intelligence to robots
That's just about any video of any VLA ever. Including Gemini Robotics 2.
ACCount37··on Gemini Robotics 2 brings whole body intelligence to robots
I'm repeating "what you claim to be impossible was done 3 years ago and was already replaced with better versions of the same idea and you are hilariously out of touch".
ACCount37··on Gemini Robotics 2 brings whole body intelligence to robots
You can literally have an LLM output target joint angles. As text. To be decoded by an explicit decoder, and executed by the robot. Some early VLAs did exactly that.

Your entire premise is wrong.

Modern action decoders are different, and usually take the form of neural networks trained end to end jointly with the rest of the model. Not fundamentally more expressive, just more in line with what we want.

ACCount37··on Gemini Robotics 2 brings whole body intelligence to robots
Did the past decades of AI research teach you absolutely nothing?

Every time you see something that "suggests a very specific, very precise, "algorithm" taught in an imitation learning session"? Scale the imitation learning up x10, x100, x1000, and it suddenly generalizes!

I'll be honest: I don't see what you see. I don't see anything that would suggest this algorithm is so brittle there's zero transfer to "even other garbage bag strings". AI robotics isn't innately brittle like conventional robotics is. But even if you are, somehow, completely right on that? Teach a hundred "very specific algorithms" like this - and watch them fuse into a manifold of algorithms that can be applied to different problems as needed.

And that is what you need. If an algorithm for "tie a garbage bag with current generation robot hands" exists and can be learned by an AI, then the gains from getting better AI are far from exhausted. The limits of robotics are the limits of AI.

This is why every AI robotics company is saying "we need more data". They understand what they're dealing with. They looked at the scaling laws and went "robotics isn't magic, that curve applies to us too". I don't get what makes people see robotics as a special magic thing, that makes them look at the advances in robot AI and say "this is intractable" and not "this is hard". It's hard. We're getting through it though.

ACCount37··on Gemini Robotics 2 brings whole body intelligence to robots
Depends entirely on VLA arch. Some have dedicated action diffusion heads that work in a standalone non-text action output space. Much like an LLM can either use an external TTS or have audio output heads attached to it directly for native S2S.

But your entire premise is wrong regardless of that.

Even if VLAs were forever bound to outputting text, you'd have to prove that they're fundamentally incapable of emitting text that maps to useful action sequences. No proof of that whatsoever - and plenty of empirical evidence suggests otherwise. Even non-specialist LLMs like ChatGPT are getting better at controlling robots and navigating 3D environments, if slowly.

ACCount37··on Twenty-five years ago it was cryptography, today it's model weights
RSA algorithm you can print on a shirt - and that would give you encryption that works as well as it possibly can. There's no "RSA 6.7 Turbo Pro EncryptingPlus" that would give you something radically new that the "t-shirt RSA" can't.

LLMs are massive, complex, and a moving target. Making them is hard, and there's no "we're done" there - the frontier updates every few months. It's easy to get and run a 2B model locally, but what concerns people isn't a cheap end device 2B - it's the 10T RLVR-fried monsters that frontier labs play with nowadays.

ACCount37··on Twenty-five years ago it was cryptography, today it's model weights
The point is: DRM is circumvented once, by some grizzled pirate captain. Everyone else just gets to plunder their downloads DRM free, nice and easy.

DRM is ineffective because the pirate uploaders are competent enough to break it, and an average end user only downloads - rank and file pirates wouldn't make their own releases even if DRM never existed and making them was easy.

ACCount37··on The End of an Era
I've seen way too much humanslop to be able to reject LLM slop outright.

There are areas where the average of LLM slop already beats the average of humanslop by a sizeable margin.

ACCount37··on Gemini Robotics 2 brings whole body intelligence to robots
Tesla's entire fleet runs on raw cameras. Including the driverless Robotaxi vehicles - which are basically a 1:1 match to how Waymo operates.

Plenty of hecklers were saying "you can't self-drive on cameras", and some still try. But Tesla's self-driving on cameras, and it seems to work fine. While Waymo's self-driving on fat sensor stacks, and it also seems to work fine. Sensors don't seem to be a differentiator of self-driving performance.

I don't think anything about self-driving tech supports your claim. Tesla was bullish on AI all the way, and Waymo has also shifted towards highly integrated end to end AI. It's the AI advances that make self-driving tractable - not anything else.

← PreviousPage 4 of 34Next →