https://blogs.microsoft.com/blog/2025/10/28/the-next-chapter...
4,679 karma · joined February 25, 2007
https://blogs.microsoft.com/blog/2025/10/28/the-next-chapter...
It sounds like a reasonable concept, but then Principia Mathematica takes 300 pages to prove that 1+1 is 2.
https://dbeaver.com/docs/dbeaver/Database-driver-SQLite/#rem...
> The requirements set forth in Sections 2 and 3 do not apply to: (a) internal use of the Software, defined as any use that does not make the Software, its outputs, or its underlying capabilities available to third parties; or (b) any use of the Software accessed through Moonshot AI's official products or certified inference partners.
"distilling" from their own hardware would be an internal use of the software.
You should definitely consider supporting your favourite news sources directly. ft.com, economist.com, lwn.net, etc. Maybe your outlook might change regarding whether they are marginalising themselves or making sure their financial motivations are more correctly aligned with high quality output.
Very good point. However, if one reads the transcript of the speech that Xi Jinping gave to the World AI Conference on 17 July, we see that he he is very much in favour of AI safety.
https://english.www.gov.cn/news/202607/17/content_WS6a5a1172...
> Second, we should strengthen risk-awareness and ensure that AI is secure and controllable. AI should be a trusted tool for humanity. We should take seriously the various types of inherent and secondary risks that AI may trigger. We should put in place laws and regulations, technological monitoring, early warning and emergency response systems in order to strengthen the line of security, prevent abuses and malicious use, and ensure that AI is always under human control. In the meantime, we should jointly oppose overstretching the national security concept in the field of AI and placing one country's security over that of others.
Now, let's contrast another important part of safety here. Amodei puts the fact that this will need to be a global effort as a mere note that sure, China will need to help too:
> Note that to be effective, testing would need to be global, which means even the CCP would need to be on board. I think this may actually be possible: as I wrote in The Adolescence of Technology, limited cooperation around preventing AI biological weapons may be possible because it is in China’s interest too.
International collaboration is the focus of Jinping's speech. But one imagines "cooperation" Amodei has in mind if "do what I say" while Jinping has more collaboration in mind here. (Even if you think 'china bad' they deserve credit for collaboration for their open weight models).
The law that caused the cookie banners also says companies cannot block access to the site if the cookies are not required for the functioning of the site.
Some German news sites have broken this and have "accept or pay" and I think this leaked to news sites in other countries. Facebook even tried it.
So, sure, if DNT is true, try to make people pay. Fine by me.
There is one. It's a DNT header. Knucklehead websites ignore it.
[1]: https://www.documentcloud.org/documents/26073615-c-2025-5430...
Yes and no. I've used a few different harnesses with closed and open models and there is definitely something going on that makes some harnesses work better than others. Many of the differences are hard to pin down and some are things people don't care about. But I wouldn't say they are commodified just yet.
1. Memory use. I have colleagues complaining that Clause Code uses several GB of memory. Meanwhile I haven't heard about that regarding codex or goose, or even opencode for that matter.
2. Suitability for local models. When you use Anthropic models, you use Anthropic as a provider. They can have software between the model and your harness that will fix issues with the model. One notable thing that even the best open weights models struggle with is broken tool calls. There is a lot that a harness can do to fix broken tool calls when working with a straight up ollama running a raw GGUF file.
3. Ease of use with non mainstream models. OpenCode has GREAT coverage of models/providers. Goose, less so as it relies on people to set up their own anthropic or openai compatability settings. e.g. Zed doesn't let you use Z.ai (which, if you speak British English, sounds ironic because "zed ai" isn't directly supported by Zed the editor).
4. Worktree support. Opencode and probably all the TUI harnesses works in a local directory - so you need the terminal to be in the worktree. Zed, however, works centrally on your git repo and tracks the worktrees so you can bounce around your work in a single window.
Of these, '2' is maybe the most important one but also the hardest to pin down as a feature. '3' is a one time cost. Of course '1' could be a blocker for someone using a macbook air or neo.
oh, .. why?
Are being handling this at all? Is it no longer needed because it gets rolled into AGENTS.md?
As sibling comment says, AA-Omniscience Hallucination Rate Benchmark puts Gemini 3.0 as the best performing aside from Gemini 3.1 preview.
I haven't encountered this. Could you name some?
And from the comments:
> From my experience in social science, including some experience in managment studies specifically, researchers regularly belief things – and will even give policy advice based on those beliefs – that have not even been seriously tested, or have straight up been refuted.
Sometimes people use fewer than one non replicatable studies. They invent studies and use that! An example is the "Harvard Goal Study" that is often trotted out at self-review time at companies. The supposed study suggests that people who write down their goals are more likely to achieve them than people who do not. However, Harvard itself cannot find such a study existing:
Can we inform dictionaries and encyclopaedia that data is now a mass noun and it is considered archaic to use data as a plural of datum?