81 karma · joined May 7, 2015
The 75% round-trip efficiency (for shorter time periods) quoted in other threads here is surprisingly high though.
METR currently simply runs out of tasks at 10-20h, and as a result you have a small N and lots of uncertainty there. (They fit a logistic to the discrete 0/1 results to get the thresholds you see in the graph.) They need new tasks, then we'll know better.
The point above is valid. I'd like to deconstruct the concept of intelligence even more. What humans are able to do is a relatively artificial collection of skills a physical and social organism needs. The so highly valued intelligence around math etc. is a corner case of those abilities.
There's no reason to think that human mathematical intelligence is unique by its structure, an isolated well-defined skill. Artificial systems are likely to be able to do much more, maybe not exactly the same peak ability, but adjacent ones, many of which will be superhuman and augmentative to what humans do. This will likely include "new math" in some sense too.
And then there's the part of models that is hard to measure. Opus has some sort of HAL-like smoothness I don't see in other models, but meanwhile, I haven't tried gpt-5.2 for coding yet. (Neither Gemini 3 Pro; I'm not claiming superiority of Opus, just that something in practical usability is hard to measure.)
Profilerating .md files need some attention though.
You can do this with any agentic harness, just plain prompting and "LLM management skills". I don't have Claude Code at work, but all this applies to Codex and GH Copilot agents as well.
And agreed, Opus 4.5 is next level.
https://bitchat.free now uses nostr for non-mesh contacts somehow, but I see no-one there either.
And it seems research is bottlenecked by computation.
Meanwhile, the whole idea of iNaturalist has evolved around voluntary reporting, community involvement, and open data, and I think some of that needs to stay. They can't turn fully commercial.
Especially selling identification services, which is related to keeping the models private, would make sense. Museums and various kinds of biodiversity monitoring schemes need mass identification, and having AI there to partially replace people would be a cost saving for the researchers and potential funding for iNaturalist. Offering such a service for free is neither practical nor justified.
(Meanwhile, I can imagine there to be lots of naturalist who hate the idea of their services being partially replaced by AI. It may lower the quality but the cost margin between a human and an iNat model is really wide.)
I think EU had a plan on using AI identification in some of their monitoring schemes. It could have been iNaturalist or someone else, anyway it demonstrates the need.
That's probably not a sustainable situation.
If they feel like keeping the models to themselves, I think it's a fair game. I give them observations, they gave me the id service for free. Maybe they even sell the models to fund their development efforts? I wouldn't mind... they need to fund their functions somehow anyway.
And remember, their observation databases are open. In fact my observations are automatically copied to the databases of a national biodiversity institution (which is open as well, except for some critical species).
Institutions need to maintain themselves and be able to pay their employees for them being able to feed their kids, etc.
And like said, the researchers themselves are on X, even Gary Marcus is there. ;)
As a partially separate issue, there are people trying to punish comments quoting AI by downvotes. You don't need to have a non-informative reply, just sourcing it to AI is enough. A random internet dude telling the same thing with less justification or detail is fine to them.
I have a list in X for AI; it's the best source of information overall on the subject, although some podcasts or RSS feeds directly from the long-form writers would be quite close. (If one is a researcher themselves, then of course it's a must to follow the paper feeds, not commentary or secondary references.)
I'd add https://epoch.ai to the list, on podcasts at least Dwarkesh Patel; on blogs Peter Wildeford (a superforecaster), @omarsar0 aka elvis from DAIR in X, also many researchers directly although some of them like roon or @tszzl are more entertaining than informative.
The point about polluted information environment resonates on me; in general but especially with AI. You get a very incomplete and strange understanding by following something like NYT who seem to concentrate more on politics than technology itself.
Of course there are adjacent areas of ML or AI where the sources would be completely different, say protein or genomics models, or weather models, or research on diffusion, image generation etc. The field is nowadays so large and active that it's hard to grasp everything that is happening on the surface level.
Do you _have_ to follow? Of course not, people over here are just typically curious and willing to follow groundbreaking technological advancements. In some cases like in software development I'd also say just skipping AI is destructive to the career in the long term, although there one can take a tools approach instead of trying to keep track of every announcement. (My work is such that I'm expected to keep track of the whole thing on a general level.)
Imagine a software developer who refuses to use IDEs, or any kind of editor beyond sed, or version control, or some other essential tool. AI is soon similar, except in rare niche cases.