380 karma · joined March 5, 2025
That's true for most advanced robotics projects those days. Every time you see an advanced robot designed to perform complex real world tasks, you bet your ass there's an LLM in it, used for high level decision-making.
If the best you can do is bring up this garbage, then you have nothing of value to say.
Benchmarks are a really good option to have.
They still kicked ass.
It seems like those AIs just have an awful lot of location familiarity. They've seen enough tagged photos to be able to pick up on the patterns, and generalize that to kicking ass at GeoGuessr.
Payment processors have ways of passing some of the chargeback risks onto the stores, and it's not like Steam itself is chargeback central. If you just want free games, pirating them is extremely easy, and trying to abuse chargebacks gets you banned.
You can learn a lot from textbooks... or you can use textbooks to give you the absolute bare minimum, and then just use the language itself repeatedly.
That might have been what they tested at IMO.
"Human preference" is incredibly fucking entangled, and we have no way to disentangle it and get rid of all the unwanted confounders. A lot of the recent "extreme LLM sycophancy" cases is downstream from that.
Zig changes a lot. So LLMs reference outdated data, or no data at all, and resort to making a lot of 50% confidence guesses.
Fossil fuel companies are damn good at PR, and they know well that they simply can't make themselves look good. The next best thing? Make someone else look worse.
If an Average Joe hears "a company that hurts the environment" and thinks OpenAI and not British Petroleum, that's a PR win.
I don't think it's easy to get that level of similarity between two humans. Twins? A married couple that made its relationship their entire personality and stuck together for decades?
2. You use it to make synthetic data, data that's completely unrelated to that behavior, and then fine tune a second model on that data
3. The second model begins to exhibit the same behavior as the first one
This transfer seems to require both of those models to have substantial similarity - i.e. to be based on the same exact base model.
Going public with that was a bold call - CIA put its reputation on the line. But Ukraine was more prepared because of it - and so were its allies.
A lot of Ukrainian officials didn't believe that the war was about to start up until the moment it did. Imagine how much worse the situation could have been without US beating the drum.
This "world model" is what you get to peek into through the car's screen. By now, it even has basic "object permanence". Nowhere near as good as a human yet. But AI is getting better, and an average driver isn't.
There's a small, sharp, high resolution color-enabled area in each eye - but the bulk of your vision field is monochrome, and mostly sensitive to motion.
You don't notice that, because your image data is stacked and post-processed to shit to make it presentable. Your brain has been doing computational photography before it was cool - 90% of what you see at any moment in time is effectively AI-generated.
A big part of a self-driving car's "safety edge" is that it isn't going to go 80 in a 40, doesn't fall asleep at the wheel, and isn't capable of DUI.
Self-driving cars still struggle in some situations most human drivers wouldn't find challenging - AI issues - but they don't make the worst, the most unforced and avoidable "human factor" mistakes.
Take any self-driving car crash where the self-driving car was found at fault. Dump the blackbox, extract the raw sensor data. What will you see?
You'll see that the car had all the sensory data it needed to make the right call, many times over. And it didn't make the right call. That's not a "sensors" problem. The sensors are good enough. The main bottleneck for self-driving is, and always was, in AI.
Which is why you get things like that Cruise car dragging a pedestrian despite being equipped with 360 cameras and a total of 5 overlapping LIDARs. It had the sensors. What it didn't have was object permanence.
Autopilot shuts down when it can't handle the situation it's in. This doesn't help it "avoid blame" at all. Because Tesla considers Autopilot implicated in any crash that happened within 5 seconds from Autopilot being disengaged.
> To ensure our statistics are conservative, we count any crash in which Autopilot was deactivated within 5 seconds before impact, and we count all crashes in which the incident alert indicated an airbag or other active restraint deployed.
NHSTA's reporting requirements are even more conservative:
> Level 2 ADAS: Entities named in the General Order must report a crash if Level 2 ADAS was in use at any time within 30 seconds of the crash and the crash involved a vulnerable road user being struck or resulted in a fatality, an air bag deployment, or any individual being transported to a hospital for medical treatment.
Do you seriously think that the main challenge Tesla is going to face when trying to scale Robotaxi up is that there isn't enough room on the roads for all the Teslas? In a world where there's currently a dozen Robotaxi Teslas per city?
If you're doomsday prepping, there's no reason not to have both. They're complimentary. Wikipedia is more reliable, but also much more narrow in its knowledge, and can't talk back. Just the "point someone who doesn't know what he's dealing with in a somewhat sensible direction" is an absolute killer feature that LLMs happen to have.
Neuralink N1 is a fully invasive BCI, with over 1000 recording channels that go down to neuron level. In practice, that's barely enough to provide a useful, reliable control interface. It's still SOTA - anything else that exists is straight up worse.
The pathway to better BCIs seems to be in more invasive interfaces with greater channel counts - not the other way around.