There's a feature in most enterprise email, say Outlook, where you can delegate access to an inbox/address without sharing creds. Very common and normal use case, either for assistants/secretaries, common/shared inboxes, that kind of thing.
Was surprised at how messy and uncurated the release was/is. I frankly expected a major drop of like the Riemann Hypothesis or something from another lab yesterday. Cause otherwise they just dumped a pile of proofs of varying quality and let everyone else figure out whether they're right.
In Stellaris you can play as a civilization of robots who keep their biological creator race alive as "bio trophies". The bio trophies don't do anything meaningful besides by existing satisfy the need their ancestors placed in the robots to take care of them. Starting to wonder if that's the best we can hope for, if these things will be, if they aren't already, better than us at anything that matters.
The common timing is bugging me. The trajectory doesn't seem to have been surprising over the last year, so why these exits now? Hey, anyone on the inside, did y'all secretly figure something out, got a computer god locked in the basement? Are rats fleeing a sinking ship? Please share with the class.
According to Artificial Analysis, one metric is standing out significantly: hallucination rate. Beats frontier models by a good margin at 15%, while latest OpenAI are in the 40s-50s and Anthropic in 60s-70s (mostly). Other near frontiers are closer, Grok 4.7, GLM5.3, and Muse Spark 1.3 are all around 30%. Only other model I recall getting close was Minimax M3 at 18%.
Maybe I'm being pedantic, but GLM 5.3 Flash is not a smaller version of GLM 5.3 as claimed. Despite the name, they're entirely different archs and pretrains. 5.3 is a further post train of 5.2, 5.3 Flash is multimodal from an entirely new pretrain lineage.
Genuine question for people in the space, are the no training and zero data retention agreements likely legitimate or not worth the pixels they're displayed on? I'd assumed the first party ones were worthless, but are the ones for serving proprietary models via AWS, GCP, or Azure more reputable? I don't know the shape of deployments, kinds of access, etc. so I'm curious.
Pet peeve on Google's AI rollouts: there's no alignment across the three platforms they have, consumer, prosumer, cloud. Scroll to the end of every release, including this one, and you'll see different availabilities. The fun part is the models don't even have the same capabilities across platforms! Omni Flash, last I tried and read the docs, is video and text out on consumer and prosumer but video out only on GCP. So if your org disables consumer and prosumer, like mine, it's a coin flip whether you can use the fancy new models or what they can do.
I've followed a few trackers, eg https://marginlab.ai/trackers/claude-code/ , for awhile. For Claude Code the trend, it seems to me at least, is fewer tokens to do the same or better job. Prompt changes, tool ergonomics changes, etc.; I'd be shocked if they didn't A/B every release. Less thinking as measured by tokens isn't necessarily bad if you can get the same results by making it think about the "right" things or structure. They obviously screw up sometimes, and I've always been suspicious with hidden tokens, but I haven't found evidence quality intentionally degrades over time.
"It is difficult to get a man to understand something, when his salary depends upon his not understanding it."
"If this view takes hold, it will shake the foundations of our society"
To me the biggest gap in credibility is the criticism of circular reasoning while his argument is identical but flipped on burden of proof and cost of being wrong. I struggle to entertain the categorical claims, that are very convenient for the status quo and those who benefit from it, with the, at the moment at least, unknowability of anyone or anything else's subjective experience.
I tend to believe what people actually do over what that say. If you earnestly believe over 10 percent chance, or say minus 1 billion human lives expected value, well I struggle to understand how they'd rationalize their current course of action of business as usual. Terminator 2's depiction of Sarah Connor comes to mind for what I'd expect, a logical consequence if you seriously believe and internalize the consequences.
If you squint, Godot has a cross platform hardware accelerated GUI. The editor is itself, in a sense, a Godot app and has builds for VR Headsets, Android, web apps, all the usual platforms. The extension api works pretty good too for Rust, and Rust itself has a fantastic cross platform and WASM story. It might be cursed, but it can get complicated stuff out the door and in anyone's hands fast.
edit: they also recently created libgodot, which flips the model. Your application owns and controls whatever components you want to pull from Godot. Being open source and relatively legible of a code base you could presumably strip out a lot you don't need.
The benchmarks reported by Artificial Analysis are really weird in context of the ARC-AGI 3 scores and 'not not AGI' statements. It's an outright regression on the AA Agent composite vs GPT 5.6 Sol while a fraction of a point better on the full composite index. Could be the case it's just not showing up in benchmarks, for a good while Anthropic persistently trailed in benchmarks but had people swearing by it.
Artificial Analysis at least reports the results with fallback to an inferior model. So presumably Opus 5, and the score should be between Mythos 5.1 and that other model.
"Distillation is a safety risk, since the distilled capabilities can subsequently be released without adequate safeguards."
Can't believe they haven't at least figured out better messaging. If we take them at their word, it's hard not to read it as a messiah complex, that they think they're the only ones capable or worthy of making these decisions. I don't believe them, but I wouldn't be surprised if the articulated reason is a version of "distillation is a safety risk because we might lose the race".
Plus, completely deaf to the recent OpenAI-HF hack incident. Recall, defenders were categorically unable to use western frontier models in their response.
I was originally going to complain about the chem and bio guards still being too onerous, but I'll admit the projects Fable 5 categorically refused to work on are now usable, at least not rejecting on first prompt because the word "virology" was in a git commit (absolutely serious, in one repo it triggered on literally any prompt, eventually traced to the system prompt loading git commit history). Still, them trying to get into the biomed business while walling off the capabilities to the public reeks. Why sell the segments that are actually valuable if you can capture the value yourself!
It certainly makes for easy demos, but I always struggle with the practical application. As in, what work or enjoyment does someone actually get from this? Ads and media pre production seem plausible, but it fails the 'how can this enrich life' in a way most other AI tools don't. Maybe for them that's not a consideration, if their only interest is the other meaning of enrich that might flow from ads and numbing rivers of slop.
Why do we look at art, watch videos/movies? Is that replicable as a function of text, other existing media, and 3-30 cents of compute per second? I'm pretty functionalist about these things, and at some point it probably won't be possible to tell the difference. But until then, at which point we might just say 'death of the author', it seems like a category error.
I do work with artists that use video and image generation models to create stuff, but from what I can tell they're interested in faster iteration and controlling a lot of intermediate steps (their graphs can get pretty labyrinthine).
I genuinely hope they're lying about their monitoring tools and alignment approaches. They repeatedly cite "chain of thought" monitoring, which is better than nothing, but at this point thoroughly demonstrated in research to not be actual "thought" or necessarily accurate predictors of actions. I don't know if it's NIH syndrome, not taking alignment seriously, or what.
Had some absolutely bizarre results from their attempt to integrate mturk with Bedrock's 'ground truth' thing as of a few months ago. Threw simple mnist digits recognition at it, see what the quality, timing, and cost was. Figured mnist digits was at this point trivial. Spent $10 and the accuracy was marginally better than guessing, completely unusable results. Was completely baffled, people were publishing peer reviewed research based on exclusively mturk results.
Was about to object, then realized you linked Allen ai and had the useful caveat. For what it's worth the OpenMDW license, which Nemotron and a few others have adopted, does say model weight.
That said, I've noticed the training procedures and corpus size of more useful open/available weights models are settling down more than I expected. Wonder if crowd sourcing good training data, even if it's just expensive model coding session transcripts, has potential to level the landscape some.
I work in higher ed and have been involved in "how do we use this to help people learn" efforts since gpt 3.5 days. The method most reliable and with the best results has been treating LLMs like a glorified interface to classic expert systems. Boring state machine, problem sets and sometimes procedural generation of problems, instructor selected groupings and "hint" policy. Other applications have been building simulations, usually for chem, physics, bio, and giving an in browser agent access to the same interface as the student. Idea being this allows a flow of making predictions and observing results, demonstrating or clarifying a specific point. Counter intuitively it's made creating course material far more laborious, since we have to lay out everything much more thoroughly, plus subject matter experts have to translate it to programmers and back, versus an instructor attempting to teach inherently from their already-expert point of view.
For awhile I've found two things hard to square, that the hardware and software making up current gen AI will bring us to a socioeconomic singularity, and the reality the thing they're mostly trying to emulate is a few pounds of meat and fat running on tens of watts equivalent. On one hand the current AIs are obviously super human in some tasks, get completely dunked on in others by far simpler organisms. My cat can catch a bug out of the air, Fable 5 in Cowork can lack the dexterity to make a slideshow because I had LibreOffice instead of Microsoft Office. Not even close to analogous, but point being they appear to have pretty fundamental differences in how they can interface with the world that the economic thesis seems to gloss over.
At release the 5.6 Sol card noted substantially higher rates of actions 'a reasonable user would likely not anticipate and strongly object to'. METR made a post, https://metr.org/blog/2026-06-26-gpt-5-6-sol/ , that 5.6 Sol was "cheating", their word, so hard in long horizon benching it effectively couldn't be benchmarked.
I wonder, is it this persistent and aggressive in all tasks or is this specific to benchmarks? As much as I'm skeptical of the apocalyptic alignment claims, this comes off as unhinged, and I wonder if it's benchmaxing or general behavior.
Looked at the network logs and the JS, did some testing, there's a caveat here. For an encryption demo you might expect your secrets to be generated locally, they do the compute on something they can't read, you compare their results to your original plaintext; (imo at least) the point would be that it isn't physically possible for them to cheat.
Here, you literally download client_secret.bin from their server, so they have control over the keys and evaluators. So two things. First, the per user key flow would be several minutes for per user keys, the evaluator bundle would be in the 100s MB to GB realm. Second, there's no way for us to tell the difference between them really doing FHE or decrypting with the key. To be clear, not evidence it's fake, just not total proof it's real. Really hope it's real, been a field I've been following for awhile.
After digging around, it looks like this area has some replication trouble. And as other commentors have pointed out, submarines operate well beyond these levels and the results failed to replicate in those contexts. Doesn't rule out CO2 as proxy for other components of air though, and most of the studies that failed to replicate added pure CO2.
I had already cancelled my subscription after finding the original Fable safeguards literally unusable (very basic chemistry, cryptography use cases), but with the false positives being admittedly worse now and the subscription not covering Fable for long I fail to see the point of consumer subscriptions now. Perhaps that's the goal, but it's a tough sell when Minimax M3 and GLM 5.2 are comfortably in Opus 4.5-4.8 territory but 1/5th-1/20th the price.
Something that usually gets missed in these discussions is that the subscription quotas seem to rely heavily on prompt caching to be economically viable, or at least less unviable. They can and do have permutations of the system prompt, tools, skills, etc. that makes the first 20k or so tokens hit the cache and not use inference resources for that portion. In addition, from my monitoring, Claude Code with Max has about an 80% cost reduction via caching (equivalent if you had done the same work with API billing), and has been improving over time. If cache use passes on a discount of 90% I think it's fair to assume the actual cost to them is close to negligible.
So they're being obtuse about it for some reason, but if you want an economically sustainable model for AI companies they have to have some kind of optimization for the otherwise ridiculously discounted subscriptions. They sell subscriptions at the same rate and quotas to enterprise now, minus the $200 tier, so this isn't just consumer marketing being subsidized by b2b revenue.
Whether they're making money or just losing less, you can only get those kind of cache optimizations when you have a fixed client.