My complaint is that after this settlement nothing has materially changed except that the big labs now benefit from higher barriers to entry in their market. Authors don't make more money (other than a one time protection payment from Anthropic to publishers and some lawyers). Literally no one else benefits, except I guess used book marketplaces and book scanner vendors.
To be clear, this isn't a problem with the court process. Everything here appears perfectly in accordance with the law. It's just an absurd state to be in.
Great, so now instead of allowing anyone to train on already scanned books for free, we can have only the richest big labs buy all the books and scan them privately to train their proprietary models. And since they buy the books used, authors still don't get any money. But at least the books are destroyed afterwards! What an improvement!
Is Fort Ross really "Renowned" as listed? I've never heard of it and honestly it doesn't look very interesting. I'd rather visit Fort Knox which is a real fortification currently in use, world famous, and not listed here.
From the very beginning of self-driving people have been making this claim that supervised systems would cause complacency that would make them perversely less safe. A lot of people believe it, but it's always been just an opinion based on speculation, not data. It's actually an empirical question. The data is in and this claim is definitively disproven. There is no increase in crashes when people use FSD. Not back when it was actually bad, not recently when it was just OK, and still not today when it's actually quite good but not perfect.
Anthropic's "durable advantage" theory of US AI dominance is looking pretty silly. There's zero indication that it will be hard for China to keep pace as models improve and start contributing to their own training. Which pretty much invalidates their policy recommendations.
They can't even blame it on distillation this time, unless they want to claim that their own preferred security measures were ineffective in preventing Chinese access to Mythos.
It doesn't necessarily mean anything to reach 99% on the public set. All of the public set is known in advance, so it's possible to hardcode rules that make this easy for the models. ARC-AGI-3 is supposed to measure generalization to unseen games, so the only score that matters is the score on the held out private test set that nobody outside the ARC prize foundation has access to. Also, I believe the private set is significantly harder than the public set.
Sure, it still needs supervision. Today. But it has definitely passed a threshold where it is now safer to supervise it than to drive without it. And it continues to improve quickly. I expect it to work unsupervised within two years (and unlike Elon I have not been saying this every year for the past decade).
Yeah, I bought it in 2018 with full knowledge that it would be many years before it worked at all. Today I used it for more than an hour around town. It's amazing. I won't buy any car without an equivalent feature in the future. And today there's nothing equivalent in any other car you can buy.
This is awesome. I would like to see tests like this done at 60 Hz as well, and also with non-3D apps. I suspect the results might look different in those conditions. A 500 Hz monitor is not the common case. 2ms is a whole frame!
This genre of AI researchers begging third parties to stop them from destroying the world is getting really tiresome. Setting up regulatory bodies with ill-defined goals to counter hypothetical threats from science fiction is a recipe for disaster.
The premise is that government is too slow moving to be able to react once actual problems are discovered, and I reject that. Yes, government usually moves slowly in most circumstances, but given singular existential threats it definitely can move quickly. Instead of acting out randomly and tying ourselves down with speculative regulation that probably won't even address the real problems, we should wait until the problems are obvious and then act decisively with targeted fixes.
Whisper small/tiny/base are almost four years old (they were not updated for Whisper v2 or v3). Is there really nothing better to benchmark against by now?
The mechanism is that people with the shingles vaccine are less likely to visit the hospital (because they don't get shingles). Because they have fewer hospital visits they are less likely to receive an incidental diagnosis of dementia from a hospital.
Good guess. The actual mechanism is that people who don't get the vaccination are more likely to need to visit the hospital to treat their shingles, and because they visit the hospital more they have more chances to get a diagnosis of dementia in a hospital. See this presentation: https://youtu.be/qlTnnQytOJ0
The lesson is to be extremely suspicious of findings of causation based on observational studies.
The limitations of mobile devices are mostly due to regulation (power limits and spectrum allocation and other requirements) and not inherent to the technology of handheld devices. And despite this, you can increase performance almost arbitrarily if you are able to increase the size and power and number of the antennas on the other side.
Yes, the Gen3 Starlink Direct-to-Cell constellation won't replace cell towers in urban or suburban areas. But I believe it could replace them in rural areas.
The limitations on the current T-Satellite service have a lot to do with the spectrum being shared with terrestrial towers and the low number of satellites.
The new constellation will be physically closer, with much larger antennas and a much larger number of satellites with a much higher capacity per satellite. It will also use dedicated spectrum with no terrestrial interference. Coverage and speed will be improved tremendously.
This is dangerously naive and misguided. They claim to want to avoid centralization of control but propose a world police state of AI regulation. Governments exerting this much control will only end in war and tyranny.
The rendering is a bit weird. It's disorienting to have the trains simultaneously underneath the map and occluding it. It also seems like the z-buffer is not being used correctly when rendering the train cars.
It was obvious from the live demo that this thing still hasn't learned when to shut up. When it stops tacking on "I'm here when you need me" to every response that could have just been "ok" or simply silence, maybe they'll have something. I think voice remains OpenAI's most disappointing product.
What is the actual per token price? The benchmarks look similar to Grok 4.5 also released today and priced at $2/M input tokens and $6/M output tokens.
Requiring SpaceX to set up a token reseller just to meet the the ownership requirement does exactly nothing for "national security". It's pure graft for whoever gets handed ownership of the reseller and nothing more. In fact if blatant enough it could be illegal under the FCPA.
Meh. I have gigabit fiber and it's enough for me right now. I could pay for 5 gigabit, but why? The last mile is almost never the bottleneck in my connection. 99% of the time it's upstream somewhere, with the only real exception being Steam game downloads. And my home network is gigabit ethernet so to actually get any benefit I'd have to upgrade a ton of my own hardware, the router and the switches and the NICs and even all the way down to the cables.